Text analysis device, text analysis method, and text analysis program

The text analysis device optimizes low-performance LMs by measuring accuracy for each processing unit, addressing the inefficiencies of conventional methods and enhancing after-coding accuracy and cost-effectiveness.

JP2026056380APending Publication Date: 2026-04-01MITSUBISHI ELECTRIC DIGITAL INNOVATION CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-19
Publication Date
2026-04-01

AI Technical Summary

Technical Problem

Conventional text analysis methods require extensive preparation of teacher data and classification models, which are costly with high-performance Large Language Models (LLMs) or inaccurate with low-performance LMs, necessitating a solution to effectively utilize LMs with relatively low performance.

Method used

A text analysis device that measures the accuracy of classification results using an inference model for each processing unit, allowing selection of appropriate units to improve accuracy in after-coding, even with low-performance LMs.

Benefits of technology

Enables effective utilization of low-performance LMs by measuring and optimizing processing units, achieving accurate after-coding results while reducing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026056380000001_ABST
    Figure 2026056380000001_ABST
Patent Text Reader

Abstract

I want to make effective use of LLM, which has relatively low performance in after-coding. [Solution] The text analysis device 100 utilizes an inference model trained to perform processing according to the content of the input language information for each processing unit, which is a unit obtained by dividing the entirety of each description in the free-form field indicated by the classification target information 192 and each category indicated by the category information 193 into one or more parts. The device identifies each classification of each description in the free-form field indicated by the classification target information 192 from each category indicated by the category information 193 to create a classification result, and by referring to the model answer of the result of identifying each classification of each description in the free-form field indicated by the classification target information 192 from each category indicated by the category information 193, it calculates the accuracy of the classification result 196 for each category indicated by the category information 193 for each processing unit, and includes a performance measurement unit 121 that creates a performance measurement result showing the calculated accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a text analysis apparatus, a text analysis method, and a text analysis program.

Background Art

[0002] As a conventional technique, there is a technique for automating "post-coding" that classifies free-form fields of questionnaires into categories, as disclosed in Patent Document 1. In the conventional technique, it was necessary to prepare teacher data for realizing the target classification and create a machine learning model based on the prepared teacher data.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the conventional technique, when performing post-coding with a large number of categories, it is necessary to prepare a large number of teacher data and create a large number of classification models. Therefore, in the conventional technique, it is a problem that a huge amount of time is required for preparation work before analysis. Furthermore, there is a technique that streamlines pre-analysis preparation by replacing the classification models prepared for each category using generative AI (Artificial Intelligence) such as LLM (Large Language Models), thereby eliminating the need to prepare training data and create classification models. However, when using LLMs with relatively high performance, the cost tends to be high and cost prediction is difficult because many LLMs are based on a pay-per-use model. On the other hand, when using LLMs with relatively low performance, there is a high risk that the required accuracy will not be met. Therefore, in order to effectively utilize LLMs with relatively low performance, it is necessary to understand the performance of the LLM in question beforehand. This disclosure aims to effectively utilize LLMs, which have relatively low performance in after-coding. [Means for solving the problem]

[0005] The text analysis device related to this disclosure is Using an inference model trained to perform processing according to the content of the input language information for each processing unit, which is a unit obtained by dividing the entirety of each description in the free-text field indicated by the information to be classified and each category indicated by the category information into one or more units, the classification of each description in the free-text field indicated by the information to be classified is identified from each category indicated by the category information to create a classification result. The performance measurement unit calculates the accuracy of the classification result for each category indicated by the category information for each processing unit, by referring to the model answer obtained by identifying each classification of each description in the free-form field indicated by the classification target information from each category indicated by the category information, and creates performance measurement results showing each calculated accuracy. It is equipped with. [Effects of the Invention]

[0006] According to this disclosure, the performance measurement unit measures the accuracy of the classification results using the inference model for each processing unit and creates performance measurement results. Here, the information to be classified may be information representing each description in the free-response field of a questionnaire, and the classification results may be after-coding results. In this case, when using an LLM with relatively low performance as the inference model, the accuracy of after-coding can be improved by selecting processing units by referring to the accuracy shown in the performance measurement results during after-coding. Therefore, by utilizing this disclosure, it is possible to effectively utilize LLMs, which have relatively low performance in after-coding. [Brief explanation of the drawing]

[0007] [Figure 1] A diagram illustrating the outline of the text analysis system 90 according to Embodiment 1. [Figure 2] A diagram illustrating the classification result 196 according to Embodiment 1. [Figure 3] A diagram showing an example configuration of the text analysis device 100 according to Embodiment 1. [Figure 4] A diagram illustrating the input information 191 according to Embodiment 1. [Figure 5] A diagram illustrating the classification target information 192 according to Embodiment 1. [Figure 6] A diagram illustrating the category setting screen 181 according to Embodiment 1. [Figure 7] This figure illustrates the category information 193 according to Embodiment 1, where (a) is a table showing the parent category and (b) is a table showing the child category. [Figure 8] A diagram illustrating the model answer registration screen 182 according to Embodiment 1. [Figure 9] A diagram illustrating the model answer information 194 related to Embodiment 1. [Figure 10] A diagram showing an example of the hardware configuration of the text analysis device 100 according to Embodiment 1. [Figure 11] A flowchart illustrating the operation of the preprocessing unit 110 according to Embodiment 1. [Figure 12]Flowchart showing the operation of the main processing unit 120 according to Embodiment 1. [Figure 13] Diagram explaining the performance measurement results 195 according to Embodiment 1. [Figure 14] Flowchart showing the operation of the performance measurement unit 121 according to Embodiment 1. [Figure 15] Diagram explaining the batch processing according to Embodiment 1. [Figure 16] Diagram explaining the batch processing according to Embodiment 1. [Figure 17] Diagram explaining the batch processing according to Embodiment 1. [Figure 18] Diagram explaining the block processing according to Embodiment 1. [Figure 19] Diagram explaining the block processing according to Embodiment 1. [Figure 20] Diagram explaining the individual processing according to Embodiment 1. [Figure 21] Diagram explaining the individual processing according to Embodiment 1. [Figure 22] Diagram showing a hardware configuration example of the text analysis device 100 according to a modification of Embodiment 1. [Figure 23] Diagram explaining the performance measurement results 195 according to Embodiment 2. [Figure 24] Diagram explaining the LLM master 291 according to Embodiment 2. [Figure 25] Diagram explaining the recommendation screen 281 according to Embodiment 2. [Figure 26] Diagram explaining the characteristics of each LLM. (a) is a table showing the characteristics of the high-performance LLM 51 and the low-performance LLM 52, and (b) is a table showing the characteristics of each processing unit. [Figure 27] Flowchart showing the operation of the performance measurement unit 121 according to Embodiment 2.

Embodiments for Carrying Out the Invention

[0008] In the description and drawings of the embodiments, the same elements and corresponding elements are denoted by the same reference numeral. The descriptions of elements denoted by the same reference numeral are omitted or simplified as appropriate. The arrows in the figures mainly indicate the flow of data or processing. Also, "part" may be read as "circuit," "device," "equipment," "process," "step," "procedure," "processing," or "circuitry" as appropriate. The functions of each part of each device may be realized by firmware, software, hardware, or a combination thereof.

[0009] Embodiment 1. This embodiment will now be described in detail with reference to the drawings.

[0010] ***Explanation of the structure*** Figure 1 is a diagram illustrating the overview of the text analysis system 90. The text analysis system 90 consists of a text analysis device 100, a high-performance LLM 51, and a low-performance LLM 52. The text analysis device 100 is implemented in an on-premise environment as a specific example. The text analysis system 90 is sometimes called a questionnaire analysis system. The text analysis device 100 is sometimes called a questionnaire analysis device. In the preprocessing of the text analysis device 100, the classification target information 192, category information 193, and model answer information 194 are registered in the text analysis device 100. In the main processing of the text analysis device 100, the text analysis device 100 analyzes the performance of the low-performance LLM (Large Language Model) 52 by processing unit and by category based on the classification target information 192, category information 193, and model response information 194. Also in the main processing, when the user of the text analysis device 100 inputs information indicating the free-response fields of the questionnaire, the text analysis device 100 performs after-coding on the free-response fields of the questionnaire and outputs the classification result 196. At this time, the text analysis device 100 appropriately queries the high-performance LLM 51 and the low-performance LLM 52 to obtain responses based on the performance measurement results of the low-performance LLM 52. High-performance LLM51 is an LLM with relatively high performance and relatively high usage fees. Low-performance LLM52 is an LLM with relatively low performance and relatively low usage fees. Note that high-performance LLM51 may be a paid LLM and low-performance LLM52 may be a free LLM. Text analysis device 100 may use multiple types of high-performance LLM51 and multiple types of low-performance LLM52. An example of classifying LLMs into two types, high-performance LLM51 and low-performance LLM52, is explained, but LLMs may be classified into three or more types. An LLM is a concrete example of an inference model that has been trained to perform processing according to the content of the input language information. An example of text analysis device 100 using an LLM as the inference model is explained, but text analysis device 100 may use other generative AI (Artificial Intelligence) or machine learning models instead of an LLM as the inference model. Hereinafter, "inference model" refers to the inference model in question. Various LLMs may also be referred to collectively as "LLM". Classification target information 192 is information that indicates each description that is subject to classification. Classification target information 192 may also be information that indicates the answers that each respondent wrote in the free-response section of the questionnaire, that is, information that indicates each description in the free-response section of the questionnaire. Category information 193 indicates the categories used in after-coding. Each category is determined appropriately according to the content and purpose of the survey. Each category indicated in category information 193 may consist of a parent category and child categories. Model response information 194 is data corresponding to classification target information 192 and category information 193, and represents the correct after-coding answer to the entered questionnaire.

[0011] After-coding is a technique for converting open-ended questions into multiple-choice questions by appropriately grouping similar open-ended responses. In after-coding, each category to which an open-ended response belongs is identified based on the content of the response. In this case, multiple categories may be identified as belonging to a single open-ended response. Because qualitative information is quantified through after-coding, it contributes to reducing the difficulty and burden of survey tabulation. Here, open-ended questions may also be referred to as free-response, free-form answers, or FA (Free Answer). Figure 2 shows an example of classification result 196. In this example, the free response corresponding to ID 1 is classified as "Not Applicable" for the parent category "Communication Area," and as "Company A Telephone" and "Company A Tablet" for the parent category "Device Used." Classification result 196 may also be an after-coding result. The following describes an example in which the text analysis device 100 performs after-coding on the free-response section of a questionnaire. However, the text analysis device 100 is also capable of performing classification on groups of texts that are typically categorized, i.e., performing processing similar to after-coding. These groups of texts may, for example, consist of questions and answers other than the free-response section of a questionnaire, or descriptions in general text fields such as posts on social networking services (SNS). Here, in this specification, general text fields are referred to as "free-response sections." In other words, the text analysis device 100 may target free-response sections of other text groups instead of "free-response sections of questionnaires," and may perform other classification processes instead of "after-coding."

[0012] Figure 3 shows an example configuration of the text analysis device 100 according to this embodiment. As shown in Figure 3, the text analysis device 100 comprises a pre-processing unit 110, a main processing unit 120, and a storage unit 190. The text analysis device 100 may store high-performance LLM 51 and low-performance LLM 52, or it may access external high-performance LLM 51 and low-performance LLM 52. The preprocessing unit 110 performs preprocessing. The preprocessing unit 110 includes an input data analysis unit 111, a classification target registration unit 112, a category setting unit 113, and a model answer registration unit 114. The main processing unit 120 executes the main processing. The main processing unit 120 includes a performance measurement unit 121 and a classification execution unit 122. The memory unit 190 stores various types of information as appropriate.

[0013] The input data analysis unit 111 extracts free-response fields or items equivalent to free-response fields from each answer to the questionnaire indicated by the input information 191, and links the extracted free-response fields with the classification target registration unit 112. The input data analysis unit 111 may also extract free-response fields or items equivalent to free-response fields from texts other than the questionnaire. The input data analysis unit 111 is sometimes referred to as the questionnaire analysis unit. Input information 191 is information indicating performance measurement questions for measuring the performance of the low-performance LLM52. Figure 4 shows specific examples of each answer to the questionnaire indicated by input information 191 as a performance measurement question. Note that the questionnaire may include items other than free-response fields. The name of the free-response field in the questionnaire may be something other than "free-response field". Note that input information 191 may also be information indicating a set of text other than the questionnaire.

[0014] The classification target registration unit 112 creates classification target information 192 representing each description in the free-response field extracted by the input data analysis unit 111, and registers the created classification target information 192 in the storage unit 190. The classification target registration unit 112 is sometimes called the questionnaire registration unit. Figure 5 is a diagram showing, in table format, the free-response fields extracted by the input data analysis unit 111 from the questionnaire shown in Figure 4, as a specific example of classification target information 192. A classification target information table may also be provided as classification target information 192. Classification target information 192 may also be information indicating each description in free-response fields other than the questionnaire. Classification target information 192 is sometimes referred to as questionnaire information.

[0015] The category setting unit 113 displays the category setting screen 181 to the user. The category setting unit 113 also creates category information 193 indicating each category set by the user through the category setting screen 181, and registers the created category information 193 in the storage unit 190.

[0016] The category setting screen 181 is a screen for users to set each category. Figure 6 shows a specific example of the category setting screen 181. In this example, the user can input multiple types of categories consisting of two levels: parent categories and child categories. The user can also set target accuracy in after-coding for each type of category. The user may also consider each category for analysis and set analysis requests based on the extracted free-text fields. "Input fields" are fields that the user can input. In this example, target accuracy can be set for the parent category, but target accuracy can also be set for other groups related to the category. The category hierarchy may have three or more levels. After each category is registered, the user may transition from the category setting screen 181 to the model answer registration screen 182. Figure 7 corresponds to Figure 6 and shows a concrete example of category information 193. Figure 7(a) is a table showing the parent categories indicated by category information 193. Figure 7(b) is a table showing the child categories indicated by category information 193. "Classification item" corresponds to a child category.

[0017] The model answer registration unit 114 displays a model answer registration screen 182 to the user, creates model answer information 194 showing the model answers registered by the user through the model answer registration screen 182, and registers the created model answer information 194 in the storage unit 190. At this time, the user registers model answers for each category for analysis in order to measure the performance of the low-performance LLM 52. Model answers are used to determine whether the answers prepared by the LLM are appropriate.

[0018] The model answer registration screen 182 is a screen for users to register model answers. Figure 8 shows a specific example of the model answer registration screen 182. In this example, the model answer registration screen 182 displays the free-form text description indicated by the classification target information 192 and each category indicated by the category information 193. On the model answer registration screen 182, the user enters the correct value for each category for each free-form text description. There is no particular limit on the number of lines in the model answer, and there is a tendency for accuracy to increase as the number of lines in the model answer increases. Figure 9 corresponds to Figure 8 and shows a concrete example of model answer information 194. Note that this example is an image of the model answer represented in a non-normalized way. The physical structure of the model answer may differ from the physical structure in this example.

[0019] The performance measurement unit 121, for each processing unit, uses an inference model to identify each classification of each description in the free-form text field indicated by the classification target information 192 from the categories indicated by the category information 193, and creates a classification result 196. Subsequently, the performance measurement unit 121 refers to the model answer obtained by identifying each classification of each description in the free-form text field indicated by the classification target information 192 from the categories indicated by the category information 193, and calculates the accuracy of the classification result 196 for each category indicated by the category information 193 for each processing unit, and creates a performance measurement result 195 showing the calculated accuracy. The model answer is the information indicated by the model answer information 194. The performance measurement unit 121 may perform a format determination process on the output of the inference model to determine whether the format is appropriate, and an accuracy determination process to determine whether the accuracy of any category indicated by the category information 193 exceeds the target accuracy corresponding to each category, and create a performance measurement result 195 based on the results of the format determination process and the accuracy determination process. The inference model may be a low-performance inference model with relatively low performance. The low-performance LLM52 is a low-performance inference model. As a specific example, the performance measurement unit 121 measures the performance of the low-performance LLM 52 for each category on a processing unit basis and registers information showing the results of the performance measurement as performance measurement results 195 in the storage unit 190. In this case, the performance measurement unit 121 may measure the performance of the low-performance LLM 52 based on the questionnaire and the analysis request and select the optimal processing unit based on the measurement results. The content of the inquiry to the low-performance LLM 52 may be an inquiry entered by the user, an inquiry created based on a template, or an inquiry included in the analysis request. A processing unit is a unit obtained by dividing the entirety of each description in the field corresponding to the free-form entry field indicated in the classification target information 192 and each category indicated in the category information 193 into one or more parts. Specific examples of processing units include a batch unit, a block unit, and an individual unit. The field corresponding to the free-form entry field may be the free-form entry field itself, or it may be a field treated as equivalent to the free-form entry field. A batch unit is a processing unit that processes the entire set of descriptions in the free-form fields indicated by the classification target information 192 and each category indicated by the category information 193 all at once. A block unit is a processing unit in which the entirety of each description in the free-form entry field indicated by the classification target information 192 and each category indicated by the category information 193 is divided into multiple blocks and processed separately for each divided block. An individual unit is a processing unit that processes each description in the field corresponding to the free-form entry field indicated in the classification target information 192, according to the category indicated in the category information 193.

[0020] The classification execution unit 122 refers to the performance measurement results 195 and selects a processing unit that satisfies the target accuracy for each category indicated by the category information 193. Using the selected processing unit, it uses an inference model to identify each classification for each description in the free-form text field indicated by the request information from the categories indicated by the category information 193. If, for each category indicated by the category information 193, there is no processing unit that satisfies the target accuracy for each category in the performance measurement results 195, the classification execution unit 122 may use an inference model with relatively higher performance than the inference model that is the performance measurement target of the performance measurement unit 121 to identify each classification for each description in the free-form text field indicated by the request information from the categories indicated by the category information 193. The classification execution unit 122 may also perform classification other than after-coding, similar to after-coding. The classification execution unit 122 is sometimes called the after-coding execution unit. As a specific example, the classification execution unit 122, based on the request information, refers to the performance measurement result 195 to select an appropriate processing unit and LLM, performs after-coding, and registers the result of the after-coding as a classification result 196 in the storage unit 190. The request information is information that indicates each description in the free-form field to be processed. The request information may be the same as the input information 191, or it may be the same as the classification target information 192.

[0021] Figure 10 shows an example of the hardware configuration of the text analysis device 100 according to this embodiment. The text analysis device 100 consists of a computer. The text analysis device 100 may consist of multiple computers.

[0022] As shown in this figure, the text analysis device 100 is a computer equipped with hardware such as a processor 11, memory 12, auxiliary storage device 13, input / output interface 14, and communication device 15. These hardware components are connected as appropriate via signal lines 19.

[0023] The processor 11 is an integrated circuit (IC) that performs arithmetic operations and controls the hardware of the computer. Specific examples of the processor 11 include a CPU (Central Processing Unit), a DSP (Digital Signal Processor), or a GPU (Graphics Processing Unit). The text analysis device 100 may have multiple processors that replace the processor 11. The multiple processors share the role of the processor 11.

[0024] Memory 12 is typically a volatile storage device, specifically RAM (Random Access Memory). Memory 12 is also called main memory. Data stored in memory 12 is saved to auxiliary storage device 13 as needed.

[0025] The auxiliary storage device 13 is typically a non-volatile storage device, specifically a ROM (Read Only Memory), an HDD (Hard Disk Drive), or flash memory. Data stored in the auxiliary storage device 13 is loaded into memory 12 as needed. The memory 12 and the auxiliary storage device 13 may be configured as a single unit.

[0026] Input / Output IF14 is a port to which input and output devices are connected. A specific example of an input / output IF14 is a USB (Universal Serial Bus) terminal. Specific examples of input devices include a keyboard and mouse. Specific examples of output devices include a display.

[0027] The communication device 15 consists of a receiver and a transmitter. A specific example of the communication device 15 is a communication chip or a NIC (Network Interface Card).

[0028] Each part of the text analysis device 100 may use the input / output IF 14 and the communication device 15 as appropriate when communicating with other devices.

[0029] The auxiliary storage device 13 stores the text analysis program. The text analysis program is a program that enables the computer to implement the functions of each part of the text analysis device 100. The text analysis program is loaded into memory 12 and executed by the processor 11.

[0030] The data used when executing the text analysis program, and the data obtained by executing the text analysis program, are appropriately stored in the memory device. The memory unit 190 is implemented by the memory device. Each part of the text analysis device 100 makes appropriate use of the memory device. The memory device consists of, as a specific example, memory 12, auxiliary memory device 13, registers in the processor 11, and at least one of the cache memory in the processor 11. Note that the terms data and information may have the same meaning. The memory device may be independent of the computer. The functions of memory 12 and auxiliary storage device 13 may be implemented by other storage devices.

[0031] The text analysis program may be recorded on a computer-readable non-volatile recording medium. Examples of non-volatile recording media include optical discs or flash memory. The text analysis program may also be provided as a program product.

[0032] ***Explanation of operation*** The operating procedure of the text analysis device 100 corresponds to the text analysis method. The program that implements the operation of the text analysis device 100 corresponds to the text analysis program. The text analysis method is sometimes called the questionnaire analysis method. The text analysis program is sometimes called the questionnaire analysis program.

[0033] Figure 11 is a flowchart illustrating an example of the operation of the preprocessing unit 110. The operation will be explained using Figure 11. Note that performance measurement questions and model answers must be prepared before executing the preprocessing.

[0034] (Step S111) The input data analysis unit 111 receives the input information 191, extracts the free-form entry fields from the received input information 191, and transmits the extracted free-form entry fields to the classification target registration unit 112.

[0035] (Step S112) The classification target registration unit 112 receives information from the input data analysis unit 111 and registers the received information in the classification target information table indicated by the classification target information 192.

[0036] (Step S113) First, the category setting unit 113 reads the registered classification target information table as the target of analysis and displays the category setting screen 181 containing the read information. Next, the user operates the category setting screen 181 to set parent category information and child category information as information indicating the analysis category. Next, the category setting unit 113 registers the parent category information and child category information set by the user in the category information 193.

[0037] (Step S114) First, the model answer registration unit 114 reads the classification target information 192 and category information 193, and displays the model answer registration screen 182 containing the read information. Next, the user registers model answers for each category registered by the category setting unit 113, based on the descriptions in the free-form text field. The expression "by category" specifically means that each parent category is the target of processing. Next, the model answer registration unit 114 registers the model answers registered by the user in the model answer information 194.

[0038] Figure 12 is a flowchart showing an example of the operation of the main processing unit 120. This operation will be explained using Figure 12.

[0039] (Step S121) The performance measurement unit 121 performs performance measurements for batch processing, block processing, and individual processing, categorized according to the processing flow described below. Here, batch processing is processing executed in batch units. Block processing is processing executed in block units. Individual processing is processing executed in individual units. The performance measurement unit 121 executes tests in order from the highest level for each category, and if the accuracy is equal to or greater than the target accuracy, it terminates the performance measurement and registers the performance measurement result in the performance measurement result 195. Here, when the test is of a higher level, the group of queries made to the LLM in the test is larger. Figure 13 shows a specific example of performance measurement results 195. In this example, the passing processing units and the accuracy of the passing processing units are shown for each category. The passing processing unit is the highest-performing processing unit among those whose corresponding accuracy meets the target accuracy.

[0040] (Step S122) The classification execution unit 122 reads the performance measurement results 195 and performs after-coding by appropriately executing "batch processing," "block processing," and "individual processing" for each category based on the read performance measurement results 195. After that, the classification execution unit 122 registers the results of the after-coding execution in the classification result 196.

[0041] Figure 14 is a flowchart illustrating an example of the operation of the performance measurement unit 121. This operation will be explained using Figure 14.

[0042] (Step S131) The performance measurement unit 121 performs a batch processing performance test to measure the performance of batch processing. Figure 15 shows a specific example of the query content to the low-performance LLM52 in a batch processing performance test. In this example, the performance measurement unit 121 queries the low-performance LLM52 for both the "communication area" category and the "user terminal" category together. Figure 16 corresponds to Figure 15 and shows a specific example of the low-performance LLM52 response in a batch processing performance test. Furthermore, the performance measurement unit 121 may edit the responses of the low-performance LLM 52 to create tabular data in order to calculate the accuracy for each category, as shown in Figure 17. Here, Figure 17 corresponds to Figure 16.

[0043] (Step S132) The performance measurement unit 121 performs a common judgment process on the output of the low-performance LLM 52. The common judgment process consists of a format judgment process and a precision judgment process. The format determination process determines whether the format of the response from the low-performance LLM52 is appropriate. The accuracy determination process determines whether the accuracy for each category exceeds the target accuracy corresponding to that category. In the accuracy determination process, the performance measurement unit 121 calculates the accuracy by comparing the model answer indicated by the model answer information 194 with the answer output by the low-performance LLM 52. If the result of the common determination process is Yes, that is, if the result of the format determination process is Yes and the result of the accuracy determination process is Yes, the performance measurement unit 121 proceeds to step S133. Otherwise, the performance measurement unit 121 proceeds to step S134.

[0044] (Step S133) The performance measurement unit 121 determines that batch processing is possible.

[0045] (Step S134) The performance measurement unit 121 performs a block processing performance test to measure the performance of block processing. Figure 18 shows a specific example of the query content to the low-performance LLM52 in a block processing performance test. In this example, the performance measurement unit 121 executes query 1, which is a query regarding the "communication area" category, and query 2, which is a query regarding the "user terminal" category, to the low-performance LLM52. Figure 19 corresponds to Figure 18 and shows a concrete example of tabular data created by compiling the responses of a low-performance LLM52 in a block processing performance test.

[0046] (Step S135) The performance measurement unit 121 performs a common judgment on the output of the low-performance LLM 52. If the result of the common judgment is Yes, the performance measurement unit 121 proceeds to step S136. Otherwise, the performance measurement unit 121 proceeds to step S137.

[0047] (Step S136) The performance measurement unit 121 determines that block processing is possible.

[0048] (Step S137) The performance measurement unit 121 performs individual processing performance tests to measure the performance of individual processing. Figure 20 shows a specific example of the queries to the low-performance LLM52 in individual processing performance testing. In this example, the performance measurement unit 121 executes queries to the low-performance LLM52 for each entry in the free-form text field, specifically for the "communication area" category and the "user terminal" category. The performance measurement unit 121 repeats the queries to the low-performance LLM52 for each category, corresponding to the number of entries in the free-form text field. Figure 21 corresponds to Figure 20 and shows a concrete example of tabular data created by compiling the responses of the low-performance LLM52 in the individual processing performance test.

[0049] (Step S138) The performance measurement unit 121 performs a common judgment on the output of the low-performance LLM 52. If the result of the common judgment is Yes, the performance measurement unit 121 proceeds to step S139. Otherwise, the performance measurement unit 121 proceeds to step S140.

[0050] (Step S139) The performance measurement unit 121 determines that individual processing is possible.

[0051] (Step S140) The performance measurement unit 121 determines that processing is impossible and returns an error. At this time, the performance measurement unit 121 returns the error details to the user so that it is clear whether the problem is related to the format or the accuracy. If accuracy is insufficient, the high-performance LLM51 is used instead of the low-performance LLM52 in after-coding related to performance measurement questions.

[0052] ***Explanation of the effects of Embodiment 1*** In Embodiment 1, model answers are created based on a questionnaire for performance measurement, and the performance of the LLM is pre-evaluated based on the created model answers. Therefore, by utilizing Embodiment 1, depending on the results of the pre-evaluation, the query content to the LLM can be automatically adjusted according to the performance of the pre-evaluated LLM so that even when using an LLM with relatively low performance, the same processing as when using an LLM with relatively high performance can be achieved. In other words, by utilizing Embodiment 1, it is possible to rationally combine an LLM with relatively high performance and an LLM with relatively low performance.

[0053] ***Other configurations*** <Example 1> Figure 22 shows an example of the hardware configuration of the text analysis device 100 according to this modified example. The text analysis device 100 includes a processor 11, a processor 11 and memory 12, a processor 11 and auxiliary storage device 13, or a processing circuit 18 instead of a processor 11, memory 12 and auxiliary storage device 13. The processing circuit 18 is hardware that implements at least some of the components of the text analysis device 100. The processing circuit 18 may be dedicated hardware, or it may be a processor that executes the program stored in memory 12.

[0054] When the processing circuit 18 is dedicated hardware, specific examples of the processing circuit 18 include a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or a combination thereof. The text analysis device 100 may include multiple processing circuits that replace the processing circuit 18. The multiple processing circuits share the role of the processing circuit 18.

[0055] In the text analysis device 100, some functions may be implemented by dedicated hardware, while the remaining functions may be implemented by software or firmware.

[0056] The processing circuit 18 can be implemented, in specific examples, by hardware, software, firmware, or a combination thereof. The processor 11, memory 12, auxiliary storage device 13, and processing circuit 18 are collectively referred to as the "processing circuitry." In other words, the functions of each functional component of the text analysis device 100 are realized by the processing circuitry. Text analysis devices 100 according to other embodiments may also have a configuration similar to this modified example.

[0057] Embodiment 2. The following will explain the differences from the previously described embodiment, primarily with reference to the drawings.

[0058] ***Explanation of the structure*** The configuration of the text analysis device 100 according to this embodiment is the same as the configuration of the text analysis device 100 according to Embodiment 1.

[0059] The performance measurement unit 121 in this embodiment creates a classification result 196 by using each inference model in the group of inference models to identify each classification of each description in the free-form text field indicated by the classification target information 192 from each category indicated by the category information 193, for each processing unit. The group of inference models consists of multiple inference models. Each inference model in the group of inference models is an inference model that has been trained to perform processing according to the content of the input language information. The performance measurement unit 121 also calculates the accuracy of the classification result 196 corresponding to each inference model for each category indicated by the category information 193 by referring to the model answer, and creates a performance measurement result 195 showing the calculated accuracy. At this time, the performance measurement unit 121 includes reference information in the performance measurement result 195 corresponding to each inference model and each processing unit. Each inference model in the inference model group may be a large-scale language model. Reference information is information based on records of creating classification results 196 corresponding to each inference model and each processing unit, and is information that users refer to when deciding whether or not to select each inference model and each processing unit. Reference information may also be information showing the processing time and usage cost when using each inference model to create classification results 196 corresponding to each inference model and each processing unit. Information showing usage costs may be information showing usage fees, or information directly related to usage fees such as consumption tokens. As a specific example, the performance measurement unit 121 reads the LLM master 291 during the performance measurement process and measures the performance of each LLM indicated by the read LLM master 291. At this time, the performance measurement unit 121 further measures the processing time and the tokens consumed during the query and response.

[0060] Figure 23 shows a specific example of the performance measurement result 195 according to this embodiment. In this example, the performance measurement result 195 further shows the processing time and the tokens consumed during the inquiry and response.

[0061] LLM Master 291 represents a group of inference models. The group of inference models consists of multiple LLMs, as a concrete example. Figure 24 shows a concrete example of LLM Master 291. In this example, in addition to the name and developer of the LLM, Web API (Application Programming Interface) usage information and the usage fee per token are shown.

[0062] The classification execution unit 122 in this embodiment further performs the process of displaying the recommendation screen 281 to the user. The classification execution unit 122 performs after-coding using the LLM and processing unit selected by the user through operation of the recommendation screen 281.

[0063] The recommendation screen 281 displays at least a portion of the performance measurement results 195 corresponding to each inference model and each processing unit, and provides reference information to help the user select an LLM and processing unit. Figure 25 shows a specific example of the recommendation screen 281. In this example, for a survey named "CCXXX", the accuracy, processing time, and cost are shown for each LLM and processing unit. "N / A" indicates that it was not possible to obtain a response in the appropriate format from the LLM. In this example, the user interface does not allow the selection of multiple types of LLMs, such as using "LLM-70b" for after-coding related to the "communication area" category and "LLM-7b" for after-coding related to the "device used" category. However, it may be possible for the user to select different LLMs for each block.

[0064] When processing is possible with either the high-performance LLM51 or the low-performance LLM52, a user requirement is expected to be the ability to select the LLM and processing unit to use, taking into consideration factors such as cost, processing time, and expected accuracy. To fulfill this requirement, in this embodiment, the performance of each LLM and each processing unit is measured using the LLM master 291. Figure 26(a) shows examples of the characteristics of the high-performance LLM51 and the low-performance LLM52, respectively. Figure 26(b) shows examples of the characteristics of each processing unit. Users may select the LLM and processing unit considering these characteristics.

[0065] A token is a unit of a word as interpreted by an LLM (Language Language Master). Even semantically identical sentences may have different numbers of tokens depending on the language and the type of LLM used. As a concrete example, the English sentence "This is a pen." has 5 tokens, including the period. On the other hand, the corresponding Japanese sentence "Kore wa pen desu." has around 7 tokens, although this varies depending on the type of LLM used. For reference, we have estimated that each Japanese character corresponds to 0.9 tokens.

[0066] ***Explanation of operation*** Figure 27 is a flowchart showing an example of the detailed operation of the performance measurement unit 121. This operation will be explained using Figure 27.

[0067] (Step S231) The performance measurement unit 121 reads the LLM master 291.

[0068] (Step S232) If there are unprocessed LLMs in the loaded LLM master 291, the performance measurement unit 121 selects one of the unprocessed LLMs as the selected LLM and proceeds to step S233. Otherwise, the performance measurement unit 121 terminates the processing of this flowchart.

[0069] (Step S233) The performance measurement unit 121 measures the performance of the selected LLM by performing a batch processing performance test and registers the performance measurement results in the performance measurement results 195.

[0070] (Step S234) The performance measurement unit 121 measures the performance of the selected LLM by performing a block processing performance test and registers the performance measurement results in the performance measurement results 195.

[0071] (Step S235) The performance measurement unit 121 measures the performance of the selected LLM by performing individual processing performance tests and registers the performance measurement results in the performance measurement results 195.

[0072] ***Explanation of the effects of Embodiment 2*** In Embodiment 2, when measuring the performance of each LLM, the average values ​​of processing time and number of tokens consumed per query for each LLM are recorded. The number of tokens consumed is, for example, the sum of the number of characters in the question and the number of characters in the answer. Then, before performing after-coding on the survey, the user is presented with the expected processing time, cost, and expected accuracy for the LLM to be used. Therefore, by utilizing Embodiment 2, the user can select the optimal LLM according to the situation.

[0073] ***Other Embodiments*** The embodiments described above can be freely combined, any component of each embodiment can be modified, or any component can be omitted in each embodiment. Furthermore, the embodiments are not limited to those shown in Embodiments 1 and 2, and various modifications can be made as needed. The procedures described using flowcharts and the like may be modified as appropriate.

[0074] The various aspects of this disclosure are summarized below as an appendix.

[0075] (Note 1) Using an inference model trained to perform processing according to the content of the input language information for each processing unit, which is a unit obtained by dividing the entirety of each description in the free-text field indicated by the information to be classified and each category indicated by the category information into one or more units, the classification of each description in the free-text field indicated by the information to be classified is identified from each category indicated by the category information to create a classification result. The performance measurement unit calculates the accuracy of the classification result for each category indicated by the category information for each processing unit, by referring to the model answer obtained by identifying each classification of each description in the free-form field indicated by the classification target information from each category indicated by the category information, and creates performance measurement results showing each calculated accuracy. A text analysis device equipped with the following features.

[0076] (Note 2) The text analysis apparatus described in Appendix 1, wherein the performance measurement unit performs a format determination process on the output of the inference model to determine whether the format is appropriate, and a precision determination process to determine whether the precision of each category indicated by the category information exceeds the target precision corresponding to each category, and creates the performance measurement result based on the results of the format determination process and the precision determination process.

[0077] (Note 3) Each category indicated by the aforementioned category information is a text analysis device as described in Appendix 1 or 2, consisting of a parent category and a child category.

[0078] (Note 4) The inference model is a low-performance inference model with relatively poor performance, as described in any one of the appendices 1 to 3 of the text analysis device.

[0079] (Note 5) The text analysis apparatus according to any one of the appendices 1 to 4, wherein the processing units handled by the performance measurement unit include a batch unit that processes the entire set of descriptions in the fields corresponding to the free-form fields indicated by the classification target information and the categories indicated by the category information all at once; a block unit that divides the entire set of descriptions in the fields corresponding to the free-form fields indicated by the classification target information and the categories indicated by the category information into multiple blocks and processes each divided block separately; and an individual unit that processes each description in the fields corresponding to the free-form fields indicated by the classification target information individually for each category indicated by the category information.

[0080] (Note 6) The inference model is a large-scale language model, and the text analysis device is one of the devices described in any one of the appendices 1 to 5.

[0081] (Note 7) The text analysis device further, A classification execution unit selects a processing unit that satisfies the target accuracy for each category indicated by the category information, based on the performance measurement results, and uses the inference model to identify each classification of each description in the free-form field indicated by the request information from the categories indicated by the category information, using the selected processing unit. A text analysis device according to any one of the appendices 1 to 6, comprising:

[0082] (Note 8) The text analysis device described in Appendix 7, wherein, for each category indicated by the category information, if there is no processing unit that satisfies the target accuracy corresponding to each category in the performance measurement results, the classification execution unit uses an inference model that is relatively more efficient than the inference model to identify each classification of each description in the free-form field indicated by the request information from each category indicated by the category information.

[0083] (Note 9) The performance measurement unit is For each processing unit, using each inference model in the group of inference models, a classification result is created by identifying each classification of each description in the free-form field indicated by the classification target information from each category indicated by the category information. Referring to the aforementioned model answer, calculate the accuracy of the classification result corresponding to each inference model for each category indicated by the category information, and create performance measurement results showing each calculated accuracy. The performance measurement results corresponding to each inference model and each processing unit include reference information based on the records used when creating the classification results corresponding to each inference model and each processing unit, which the user can use as a reference when deciding whether or not to select each inference model and each processing unit. Each of the aforementioned inference models is an inference model that has been trained to perform processing according to the content of the input language information. The classification execution unit is a text analysis device as described in Appendix 7 or 8, which displays performance measurement results corresponding to each inference model and each processing unit.

[0084] (Note 10) Each of the aforementioned inference models is a large-scale language model. The aforementioned reference information is a text analysis device described in Appendix 9, which shows the processing time and usage cost when using each inference model to create classification results corresponding to each inference model and each processing unit.

[0085] (Note 11) The aforementioned classification information refers to each of the descriptions in the free-response section of the questionnaire. The classification result is the text analysis device described in any one of the appendices 1 to 10, which are after-coding results. [Explanation of Symbols]

[0086] 11 Processor, 12 Memory, 13 Auxiliary storage device, 14 Input / Output IF, 15 Communication device, 18 Processing circuit, 19 Signal line, 51 High-performance LLM, 52 Low-performance LLM, 90 Text analysis system, 100 Text analysis device, 110 Pre-processing unit, 111 Input data analysis unit, 112 Classification target registration unit, 113 Category setting unit, 114 Model answer registration unit, 120 Main processing unit, 121 Performance measurement unit, 122 Classification execution unit, 181 Category setting screen, 182 Model answer registration screen, 190 Storage unit, 191 Input information, 192 Classification target information, 193 Category information, 194 Model answer information, 195 Performance measurement results, 196 Classification results, 281 Recommendation screen, 291 LLM master.

Claims

1. Using an inference model trained to perform processing according to the content of the input language information for each processing unit, which is a unit obtained by dividing the entirety of each description in the free-text field indicated by the information to be classified and each category indicated by the category information into one or more units, the classification of each description in the free-text field indicated by the information to be classified is identified from each category indicated by the category information to create a classification result. The performance measurement unit calculates the accuracy of the classification result for each category indicated by the category information for each processing unit, by referring to the model answer obtained by identifying each classification of each description in the free-form field indicated by the classification target information from each category indicated by the category information, and creates performance measurement results showing each calculated accuracy. A text analysis device equipped with the following features.

2. The text analysis apparatus according to claim 1, wherein the performance measurement unit performs a format determination process on the output of the inference model to determine whether the format is appropriate, and a precision determination process to determine whether the precision of any category indicated by the category information exceeds the target precision corresponding to each category, and creates the performance measurement result based on the results of performing the format determination process and the precision determination process.

3. The text analysis apparatus according to claim 1 or 2, wherein each category indicated by the category information consists of a parent category and a child category.

4. The text analysis apparatus according to claim 1 or 2, wherein the inference model is a low-performance inference model with relatively poor performance.

5. The text analysis apparatus according to claim 1 or 2, wherein the processing units handled by the performance measurement unit include a batch unit that processes the entire set of descriptions in the fields corresponding to the free-form fields indicated by the classification target information and the categories indicated by the category information all at once; a block unit that divides the entire set of descriptions in the fields corresponding to the free-form fields indicated by the classification target information and the categories indicated by the category information into multiple blocks and processes each divided block separately; and an individual unit that processes each description in the fields corresponding to the free-form fields indicated by the classification target information individually for each category indicated by the category information.

6. The text analysis apparatus according to claim 1 or 2, wherein the inference model is a large-scale language model.

7. The text analysis device further, A classification execution unit selects a processing unit that satisfies the target accuracy for each category indicated by the category information, based on the performance measurement results, and uses the inference model to identify each classification of each description in the free-form field indicated by the request information from the categories indicated by the category information, using the selected processing unit. A text analysis device according to claim 1 or 2, comprising:

8. The text analysis apparatus according to claim 7, wherein, for each category indicated by the category information, if there is no processing unit that satisfies the target accuracy corresponding to each category in the performance measurement results, the classification execution unit uses an inference model that is relatively more efficient than the inference model to identify each classification of each description in the free-form field indicated by the request information from each category indicated by the category information.

9. The performance measurement unit is For each processing unit, using each inference model in the group of inference models, a classification result is created by identifying each classification of each description in the free-form field indicated by the classification target information from each category indicated by the category information. Referring to the aforementioned model answer, calculate the accuracy of the classification result corresponding to each inference model for each category indicated by the category information, and create performance measurement results showing each calculated accuracy. The performance measurement results corresponding to each inference model and each processing unit include reference information based on the records used when creating the classification results corresponding to each inference model and each processing unit, which the user can use as a reference when deciding whether or not to select each inference model and each processing unit. Each of the aforementioned inference models is an inference model that has been trained to perform processing according to the content of the input language information. The text analysis apparatus according to claim 7, wherein the classification execution unit displays performance measurement results corresponding to each inference model and each processing unit.

10. Each of the aforementioned inference models is a large-scale language model. The text analysis apparatus according to claim 9, wherein the aforementioned reference information shows the processing time and usage cost when each inference model is used to create classification results corresponding to each inference model and each processing unit.

11. The aforementioned classification information refers to each of the descriptions in the free-response section of the questionnaire. The text analysis apparatus according to claim 1 or 2, wherein the classification result is the result of after coding.

12. Computers Using an inference model trained to perform processing according to the content of the input language information for each processing unit, which is a unit obtained by dividing the entirety of each description in the free-text field indicated by the information to be classified and each category indicated by the category information into one or more units, the classification of each description in the free-text field indicated by the information to be classified is identified from each category indicated by the category information to create a classification result. A text analysis method that calculates the accuracy of the classification result for each category indicated by the category information for each processing unit, by referring to model answers obtained by identifying each classification of each description in the free-form field indicated by the classification target information from each category indicated by the category information, and creates performance measurement results showing the calculated accuracy.

13. Using an inference model trained to perform processing according to the content of the input language information for each processing unit, which is a unit obtained by dividing the entirety of each description in the free-text field indicated by the information to be classified and each category indicated by the category information into one or more units, the classification of each description in the free-text field indicated by the information to be classified is identified from each category indicated by the category information to create a classification result. Performance measurement process: Referencing model answers obtained by identifying each category of each description in the free-response field indicated by the classification target information from each category indicated by the category information, calculate the accuracy of the classification result for each category indicated by the category information for each processing unit, and create performance measurement results showing the calculated accuracy. A text analysis program that causes a computer, which is a text analysis device, to perform the analysis.

Citation Information

Patent Citations

  • Information processing method and information processing device

    JP2021015549A