Classification processing program and method

The classification processing program leverages a large-scale language model to classify character strings composed of natural language, addressing the challenges of existing text mining methods and achieving improved accuracy and efficiency.

JP2025072591AActive Publication Date: 2025-05-09川上成年
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
JP2025020382
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-03-21
Filing Date
2025-02-11
Publication Date
2025-05-09
Estimated Expiration
2044-01-20

AI Technical Summary

Technical Problem

Existing text mining processing methods struggle with classifying character strings composed of natural language, which is a complex and challenging task.

Method used

A classification processing program that utilizes a large-scale language model to input and process multiple character strings, providing classification instruction information and obtaining classification result information to categorize the strings into various natural language categories.

Benefits of technology

Enables effective classification processing of character strings made up of natural language, improving accuracy and efficiency in categorizing complex linguistic data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025072591000001_ABST
    Figure 2025072591000001_ABST
Patent Text Reader

Abstract

To provide a program for classifying character strings.SOLUTION: A computer is caused to perform processing of inputting each of a plurality of character strings relating to a specific object and classification instruction information including instructions for classifying the plurality of character strings into one of a plurality of classifications expressed in natural language into a large-scale language model, and obtaining classification result information in which each of the plurality of character strings is classified into one of the plurality of classifications expressed in natural language from the large-scale language model.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a classification processing program. Reach and methods. [Background technology]

[0002] Patent Document 1 discloses a patent map generation program. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent No. 6586614 Summary of the Invention [Problem to be solved by the invention]

[0004] In Patent Document 1, text mining processing is used for character string classification processing. However, the processing of classifying character strings written in natural language is highly difficult.

[0005] The present invention has been made to solve such problems in the conventional technology, and has an object to perform classification processing of character strings written in natural languages. [Means for solving the problem]

[0006] The present invention is a classification processing program that causes a computer to execute a process of inputting, into a large-scale language model, each of a plurality of character strings related to a specific target and classification instruction information including an instruction to classify the plurality of character strings into one of a plurality of classifications expressed in natural language, and obtaining, from the large-scale language model, classification result information in which each of the plurality of character strings is classified into one of the plurality of classifications expressed in natural language.

[0007] The present invention is a classification processing method in which a computer inputs, into a large-scale language model, each of a plurality of character strings related to a specific target and classification instruction information including an instruction to classify the plurality of character strings into one of a plurality of classifications expressed in natural language, and obtains, from the large-scale language model, classification result information in which each of the plurality of character strings is classified into one of a plurality of classifications expressed in natural language.

[0008] The present invention is a classification processing method executed in a system in which a generation AI server that provides natural language processing using a large-scale language model and a terminal are connected via a network, in which the terminal transmits string information in a natural language and classification instruction information including classifications in multiple natural languages ​​to the generation AI server, and the generation AI server classifies the string information in natural language into classifications in the multiple natural languages ​​based on the classification instruction information and transmits it to the terminal. Effect of the Invention

[0009] Classification processing program of the present invention Reach According to this method, it is possible to perform classification processing of character strings written in natural languages. [Brief description of the drawings]

[0010] [Figure 1] Schematic diagram of a system for executing a classification processing program. [Diagram 2] Block diagram of an apparatus for executing a classification processing program [Diagram 3] Classification process flowchart [Figure 4] Data frame containing the contents of the abstract [Diagram 5] Data frame after STEP 1 processing [Figure 6] Data frame after STEP2 and 3 [Figure 7] Generated graph [Figure 8] A data frame containing the review contents [Figure 9] Data frame after STEP 2 processing [Figure 10] Generated graph DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] The classification processing program according to the embodiment will be described below. Reach The method and method will be described in detail with reference to the drawings.

[0012] 1 is a schematic diagram of a system that executes a classification processing program. An information processing device (terminal 1) for executing the classification processing program is connected to a generation AI server 2 via a network.

[0013] The generation AI server 2 is a computer that executes classification of input character strings. The generation AI server 2 of the embodiment is a generation AI server that provides a cloud-based service incorporating large-scale language models such as ChatGPT (above, conversational service), GPT-4 Turbo, GPT-4, GPT-3.5 Turbo, GPT-3.5, and GPT-3 (above, API service) from OpenAI. Note that the generation AI server 2 is not limited to this, and may be a generation AI server that provides a natural language processing service incorporating a large-scale language model that provides similar functions (for example, Bard (above, conversational service) and Gemini (above, API service) developed by Google, etc.).

[0014] The generation AI server 2 of the embodiment has a CPU, memory, input / output devices, and an external interface, and performs natural language processing (classification processing in this embodiment) incorporating a large-scale language model in response to input from an external device, and outputs the result.

[0015] FIG. 2 is a block diagram of an information processing device (terminal 1) for executing the classification processing program of the embodiment.

[0016] A CPU (Central Processing Unit) 101, a ROM (Read Only Memory) 102, a RAM (Random Access Memory) 103, a HDD (Hard Disk Drive) 105, an external I / F (Interface) 106, and an input unit 107 are connected via a system bus 108. The CPU 101, the ROM 102, and the RAM 103 constitute a control unit 104.

[0017] The ROM 102 prestores programs and thresholds executed by the CPU 101. The RAM 103 has various memory areas, such as an area for loading the programs executed by the CPU 101 and a work area that serves as a work area for data processing by the programs.

[0018] The HDD 105 stores natural language sentence data (such as patent data and questionnaire data) input from the input unit 107. The external I / F 106 is an interface for communicating with an external device such as an external server.

[0019] The external I / F 106 may be any interface that performs data communication with an external device, and may be, for example, a device (such as a USB memory) that connects locally to the external device, or may be a network interface for communication via a wired or wireless network.

[0020] The control unit 104 exchanges data with an external device (such as a generation AI server) via the external I / F 106 to obtain the Python library required to execute the program, send string data, and receive classification results.

[0021] The external I / F 106 is connected to a display device (not shown) such as a liquid crystal display, etc. The input unit 107 is an input device such as a keyboard, a mouse, a scanner (reading device), etc.

[0022] (Embodiment 1: Patent data analysis) Next, the processing procedure of the classification processing program and method of the first embodiment will be described. In the first embodiment, an example of creating a problem-solution map related to absorbent articles is shown. The following code blocks 1 to 9 are connected in succession to form the program and method of the embodiment, and should basically be written in succession. However, for ease of explanation, they will be divided into nine parts and explained.

[0023] 3 is a flowchart of the classification processing program and method according to the embodiment. In STEP 1, problem parts and solutions are extracted from the summary (character string information), in STEP 2, the extracted problems (first category classification) are classified, in STEP 3, the extracted solutions (second category classification) are classified, in STEP 4, cross-tabulation is performed on the classified problems and solutions, and in STEP 5, the tabulation results are visualized.

[0024] In the embodiment, patent data (abstract) is used as the character string to be processed, but the character string to be processed is not limited to this, and character strings within the specification body or claims may also be used, and character strings from surveys, user reviews, etc. other than patent data may also be processed.

[0025] (STEP 1: Extracting issues and solutions) First, the control unit 104 extracts character strings of the problem and character strings of the solution from the abstracts of multiple patent data. The program is as follows. The following program is a Python program stored in the control unit of the information processing device (terminal 1). However, other programming tools and programming languages ​​may also be used. The following program is merely an example, and the processing order and the libraries, functions, and variable names used may be changed.

[0026] (Code Block 1) import pandas as pd from google.colab import drive drive.mount(' / content / drive') df = pd.read_excel(' / content / drive / MyDrive / patdata.xlsx') df['problem'] = df['summary'].str.extract('[issue](.+?)[solution]', expand=False) df['solution'] = df['summary'].str.extract('[Solution](.+?)[Selection]', expand=False)

[0027] A detailed explanation of this program is as follows. First, the control unit 104 imports a module called pandas. Pandas is a tool for performing data analysis in Python. Next, a function called drive is imported from a module called google.colab. Drive is a function for accessing Google Drive.

[0028] The control unit 104 executes the drive function to mount Google Drive to the path ' / content / drive'. Mounting means that files and folders in Google Drive can be operated on Colab.

[0029] The control unit 104 uses the pd.read_excel function to read an Excel file called patdata.xlsx located in the path ' / content / drive / MyDrive / patdata.xlsx'. This Excel file stores patent data that includes abstract data for multiple items. A portion of the data frame is shown in FIG. 4. The contents of the patent abstracts are stored in the summary column of the patent data. Note that due to space limitations, FIG. 4 only shows an example of three items (three rows), but in the embodiment, a data frame of 50 items (50 rows) is used. Of course, more than this number is possible.

[0030] The control unit 104 assigns the read file to a variable called df. df is a table-format data structure called a data frame. By specifying df['summary'], only the column called summary is extracted from df. The summary column stores the contents of the patent abstract written in natural language.

[0031] The control unit 104 uses the str.extract function to extract the string from "Problem" to "Solution" from the summary column. This string represents the patent problem (problem string information). The extracted substring is added to a new column called problem. A new column problem is created in df by setting df['problem'] = ...

[0032] Similarly, the control unit 104 uses the str.extract function to extract the character string from "Solution" to "Selection diagram" from the summary column. This part of the character string represents the solution of the patent (this is called solution string information). The extracted part of the character string is added to a new column called solution. By setting df['solution'] = ..., a new column solution is created in df.

[0033] Fig. 5 shows a data frame to which the problem column and solution column have been added by the processing of STEP 1. Note that Fig. 5 shows only three cases (three rows) due to space limitations, but in the embodiment, a data frame consisting of 50 cases (50 rows) is generated. Note that the character strings in the summary column are omitted because they were shown in Fig. 4.

[0034] It should be noted that the process of extracting the character strings "Problem" and "Solution" from the summary is not essential. In other words, the entire summary may be sent as is to the generation AI server as character string information for classification. However, by extracting and classifying the character strings "Problem" and "Solution" separately, the classification accuracy by the generation AI server 2 can be improved, so it is preferable to include a process of extracting the character strings "Problem" and "Solution".

[0035] (STEP 2: Issue classification step) Next, the control unit 104 transmits task classification instruction information including multiple task string information and multiple task classifications (classifications of the first category) to the generation AI server, and receives task classification result information from the generation AI server. The program is as follows.

[0036] (Code Block 2) !pip install openai import openai import re

[0037] This command is to install the OpenAI Python library. The ! indicates that the command should be interpreted as a shell command. pip is a Python package manager, and install is the command used to install libraries.

[0038] The control unit 104 imports Python libraries called openai and re. The openai library is for executing various natural language processing tasks using the OpenAI API, and the re library is for using regular expressions. Note that the use of the openai library actually requires input of an API key, but this is omitted from the description of the program because it would be difficult to make it public.

[0039] (Code Block 3) def generate_p_class(text): response = openai.ChatCompletion.create( model="gpt-4", messages=[ {"role": "system", "content": "·Please assign one of the following five classifications to the text you enter: 'No leaks', 'No sweating', 'Easy to move in', 'Gentle on the skin', or 'Other'."}, {"role": "user", "content": text}]) p_class = response["choices"][0]["message"]["content"] return p_class

[0040] This program uses OpenAI's language model "GPT-4" to provide the ability to classify a given sentence into one of five categories. Note that the language model is not limited to "GPT-4", but may be "GPT-4 Turbo", "GPT-3.5 Turbo", or the next version of the GPT series. These are known to demonstrate high accuracy in natural language processing tasks.

[0041] The generate_p_class(text) function receives one text (text) as an argument, and the control unit 104 calls the Chat Completion API of OpenAI to classify the text. The API sends a message consisting of two roles, that of the system and that of the user, and sends an instruction from the system side saying, "Please assign one of the five classifications, 'No leaking', 'No sweating', 'Easy to move', 'Gentle on the skin', and 'Other', to the text to be entered next." This instruction information is called the task classification instruction information.

[0042] Note that 'leak-proof', 'non-sweaty', 'easy to move in', 'gentle on the skin', and 'other' are classifications of tasks (classifications of the first category) made up of predetermined natural language. Furthermore, the expression of the task is not limited to '~' and other expressions may be used. Furthermore, the instructions (prompts) in natural language may be more detailed. Furthermore, although the embodiment instructs to classify into one of the categories, instructions to classify into multiple categories may also be used. Furthermore, the number of categories is not limited to five and may be more or less than five.

[0043] The generation AI server 2 classifies the given text (task string information) into an appropriate category based on the task classification instruction information and replies. The response from the generation AI server 2 contains the classified category text (task classification result information). The task classification result information is stored in the p_class variable and returned as the output of the function.

[0044] (Code Block 4) def clean_text(text): text = re.sub(r'\s+', ' ', text) text = re.sub(r'[^\w\s]', '', text) return text

[0045] This program is a function that returns a string with special characters and extra spaces removed from the string (text) received as an argument. Specifically, it uses re (regular expression manipulation), a standard library of Python, to manipulate the string.

[0046] re.sub(r'\s+', ' ', text) replaces consecutive whitespace characters with a single whitespace character. \s represents a whitespace character, and + indicates that it can occur more than once. In other words, it replaces multiple whitespace characters with a single one.

[0047] Next, re.sub(r'[^\w\s]', '', text) removes non-alphanumeric and non-whitespace characters. [^\w\s] stands for non-alphanumeric and non-whitespace characters, and ^ stands for negation, meaning it means to remove non-alphanumeric and non-whitespace characters.

[0048] Finally, we return the cleaned string. return text indicates that we want to return the cleaned string as the output of the function. Note that this code block 4 is not required.

[0049] (Code Block 5) df['p_class'] = df['problem'].apply(lambda x: generate_p_class(clean_text(x)))

[0050] This program cleans the text of each line stored in the problem column of the Pandas data frame df using the clean_text function, and then generates a classification (p_class) using the generate_p_class function and adds the result to a new column p_class in df.

[0051] The control unit 104 uses the apply method to apply the clean_text function to the character strings in each row of the df['problem'] column to clean the text. lambda x: represents an argument applied to the clean_text function.

[0052] Next, the control unit 104 passes the cleaned text to the generate_p_class function, and assigns the result of classifying the text into one of the five categories to the df['p_class'] column. Finally, the data frame df including the new p_class column is returned, as shown in FIG.

[0053] (STEP 3: Solution classification step) Next, the control unit 104 transmits a plurality of solving means character string information and a plurality of solving means classifications (second category classifications) to the generation AI server, and receives the solving means result information from the generation AI server. The program is as follows.

[0054] (Code Block 6) def generate_s_class(text): response = openai.ChatCompletion.create( model="gpt-4", messages=[ {"role": "system", "content": "·Please assign one of the five classifications to the text you enter below: 'Absorbent structure', 'Breathable structure', 'Skin-friendly structure', 'Fixed structure', or 'Other'."}, {"role": "user", "content": text}]) s_class = response["choices"][0]["message"]["content"] return s_class

[0055] This program is a function that uses OpenAI's GPT-4 (GPT-3.5 is also possible) to classify the structure of a sentence for the string (text) received as an argument.

[0056] The generate_s_class(text) function receives one text (text) as an argument, and the control unit 104 calls the Chat Completion API of OpenAI to classify the text. The API sends a message consisting of two roles, that of the system and that of the user, and sends an instruction from the system side saying, "Please assign one of the five classifications, 'absorbent structure', 'breathable structure', 'texture structure', 'fixed structure', and 'other', to the sentence to be entered next." This instruction information is used as the solution means classification instruction information.

[0057] Incidentally, 'absorption structure', 'breathable structure', 'texture structure', 'fixed structure', and 'other' are classifications of solutions (classifications of the second category) made up of predetermined natural language. Furthermore, the expression of the solutions is not limited to 'structure', and other expressions may be used. Furthermore, the instructions in the natural language may be more detailed. Furthermore, although the instructions in the embodiment are to classify into one of the categories, instructions to classify into multiple categories may also be used. Furthermore, the number of categories is not limited to five, and may be more or less than five.

[0058] The generation AI server 2 classifies the given text (solution means string information) into an appropriate classification based on the solution means classification instruction information and returns it. The response from the generation AI server 2 contains the classified classification text (solution means classification result information). The solution means classification result information is stored in the s_class variable and returned as the output of the function.

[0059] The openai.ChatCompletion.create() method is used to make requests to the API. Messages sent to the API are given as a list with two elements. The first element is a message with the role system, which represents the classification instructions. The second element is a message with the role user, which represents the input text. When the API receives these messages, it generates and returns the appropriate classification. Finally, the classification result (s_class) is returned.

[0060] (Code Block 7) df['s_class'] = df['solution'].apply(lambda x: generate_s_class(clean_text(x)))

[0061] This program creates the s_class column by applying a function to classify the structure of each sentence in the solution column of the data frame df using OpenAI's GPT-4.

[0062] Specifically, the control unit 104 preprocesses each sentence included in df['solution'] with the clean_text function, and then stores the result of classifying the sentence structure with the generate_s_class function in the s_class column. Here, too, a function is applied to each sentence using a lambda expression.

[0063] The generate_s_class function uses the OpenAI API as explained above to classify the given text. Specifically, it prompts the user to assign one of the following classifications based on the structure contained in the text: 'absorbent structure', 'breathable structure', 'texturable structure', 'fixed structure', or 'other'. Finally, as shown in Figure 6, the classification results for each solution are stored in df['s_class'].

[0064] Fig. 6 shows a data frame to which the p_class column and the s_class column have been added by the processing in STEPs 2 and 3. Note that, due to space limitations, Fig. 6 shows only 20 items (20 rows) as an example, but in the embodiment, a data frame df consisting of 50 items (50 rows) is generated. Also, the description of the summary column, problem column, and solution column is omitted because they are described in Figs. 4 and 5.

[0065] (STEP 4: Aggregation step) Next, the control unit 104 cross-tabulates the multiple summaries (character string information) for each problem classification and solution means classification. The program is as follows.

[0066] (Code Block 8) result = df.groupby(['p_class', 's_class']).size().reset_index(name='Counts')

[0067] In this program, the control unit 104 groups the data frames by p_class and s_class, and counts (aggregates) the summaries by problem classification and solution classification, i.e., executes a so-called cross-tabulation. This operation makes it possible to know the frequency of occurrence of solutions for each problem. (STEP 5: Visualization step) Next, the control unit 104 visualizes the tallying results. The program is as follows.

[0068] (Code Block 9) ! pip install japanize-matplotlib import matplotlib.pyplot as plt import numpy as np import matplotlib.cm as cm import japanize_matplotlib y = result['p_class'].tolist() x = result['s_class'].tolist() sizes = [size * 100 for size in result['Counts'].tolist()] nums = [num for num in result['Counts'].tolist()] categories = list(set(y)) colors = cm.rainbow(np.linspace(0, 1, len(categories))) colors = {categories[i]: colors[i] for i in range(len(categories))} colors = [colors[item] for item in y] fig, ax = plt.subplots(figsize=[8,8]) ax.scatter(x, y, s=sizes, c=colors) for i, txt in enumerate(nums): ax.annotate(txt, (x[i], y[i]), textcoords="offset points", xytext=(0,0), ha='left', va='bottom') ax.set_xlabel("Solution means") ax.set_ylabel("Problem") ax.set_title("Problem - Solution means map") plt.xticks(rotation=90) plt.show()

[0069] This program creates a so-called problem-solution map. The control unit 104 creates a scatter plot using matplotlib. s_class is placed on the x-axis and p_class on the y-axis, and the size and color of each point is set based on the respective category. Specifically, the color corresponding to the y column value (problem category) of the grouped results is obtained from the color list generated by the rainbow color map. Also, the count number of the category that the point represents is displayed for each point.

[0070] Finally, the control unit 104 sets the style of the graph and displays the graph on a display device (not shown) such as a liquid crystal display connected via the external I / F 106 (FIG. 7). Note that the mapping may be done using other libraries such as seaborn instead of matplotlib, or a crosstab function may be used to perform cross-tabulation and output in the form of a two-way table or heat map.

[0071] The visualization is not limited to this, and may be a time series diagram showing the relationship between the application date of the patent data and the above classification, or a line or bar graph showing the relationship between the applicant of the patent data and the above classification. By using matplotlib, various forms of visualization are possible.

[0072] (Embodiment 2: Survey Analysis) Next, the processing procedure of the classification processing program and method of the second embodiment will be described. In the second embodiment, an example is shown in which user reviews of a vacuum cleaner are classified, tabulated, and visualized. The following code blocks 1 to 6 are connected in series to form the program and method of the embodiment, and should basically be written in series. However, for ease of explanation, the program and method will be divided into six parts and explained.

[0073] In the processing of the classification processing program and method of the second embodiment, in STEP 1, the user reviews are classified, in STEP 2, the classified user reviews are tallied, and in STEP 3, the tallied results are visualized (flowchart omitted).

[0074] (STEP 1: Classification step) First, the control unit 104 creates a data frame from a plurality of user reviews. The program is as follows. The following program is a Python program stored in the control unit of the information processing device (terminal 1). However, other programming tools and programming languages ​​may be used. The following program is merely an example, and the processing order and the libraries, functions, and variable names used may be changed.

[0075] (Code Block 1) import pandas as pd from google.colab import drive drive.mount(' / content / drive') df = pd.read_excel(' / content / drive / MyDrive / reviewdata.xlsx')

[0076] A detailed explanation of this program is as follows. First, the control unit 104 imports a module called pandas. Pandas is a tool for performing data analysis in Python. Next, a function called drive is imported from a module called google.colab. Drive is a function for accessing Google Drive.

[0077] The control unit 104 executes the drive function to mount Google Drive to the path ' / content / drive'. Mounting means that files and folders in Google Drive can be operated on Colab.

[0078] The control unit 104 uses the pd.read_excel function to read an Excel file called reviewdata.xlsx located in the path ' / content / drive / MyDrive / reviewdata.xlsx'. This Excel file stores multiple user reviews. A portion of the data frame is shown in FIG. 8. The review column of the review data stores the contents of the user reviews. Note that due to space limitations, only three items (three rows) are shown as an example in FIG. 8, but in the embodiment, a data frame with 100 items (100 rows) is used. Of course, more than this number is possible.

[0079] The control unit 104 assigns the read file to a variable called df. df is a tabular data structure called a data frame. By specifying df['review'], only the column called review is extracted from df. The review column stores the content of the user review written in natural language (review string information).

[0080] Next, the control unit 104 transmits the needs classification instruction information including the multiple review string information and the multiple needs classifications to the generation AI server, and receives the needs classification result information from the generation AI server. The program is as follows.

[0081] (Code Block 2) !pip install openai import openai

[0082] This command is to install the OpenAI Python library. The ! indicates that the command should be interpreted as a shell command. pip is a Python package manager, and install is the command used to install libraries.

[0083] The control unit 104 imports a Python library called openai. The openai library is for executing various natural language processing tasks using the OpenAI API. Note that, in order to use the openai library, it is actually necessary to input an API key, but this is omitted from the description of the program because it would be difficult to make it public.

[0084] (Code Block 3) def generate_n_class(text): response = openai.ChatCompletion.create( model="gpt-4", messages=[ {"role": "system", "content": "·Please assign one of the following 10 classifications to the text you enter: 'Comfortable cleanliness', 'Reduced effort', 'Improved efficiency of housework', 'Safe hygiene management', 'Stress-free operation', 'Pride in your home', 'Satisfaction with ease of use', 'Comfortable living environment', 'Satisfaction with time savings', 'Confidence in making a wise choice'."}, {"role": "user", "content": text}]) n_class = response["choices"][0]["message"]["content"] return n_class

[0085] This program uses OpenAI's language model "GPT-4" to provide the ability to classify a given sentence into one of five categories. Note that the language model is not limited to "GPT-4", but may be "GPT-4 Turbo", "GPT-3.5 Turbo", or the next version of the GPT series. These are known to demonstrate high accuracy in natural language processing tasks.

[0086] The generate_n_class(text) function receives one text as an argument, and the control unit 104 calls the Chat Completion API of OpenAI to classify the text. The API sends a message consisting of two roles, that of the system and the user, and sends an instruction from the system side saying, "Please assign one of the following 10 classifications to the sentence to be entered next: 'comfortable cleanliness', 'effort reduction', 'improved efficiency of housework', 'safe hygiene management', 'stress-free operation', 'pride in home', 'satisfaction with ease of use', 'comfortable living environment', 'fulfillment from time saving', and 'confidence in making a wise choice'." This instruction information is used as needs classification instruction information.

[0087] In addition, 'comfortable cleanliness', 'reduced effort', 'improved efficiency of housework', 'safe hygiene management', 'stress-free operation', 'pride in one's home', 'satisfaction with ease of use', 'comfortable living environment', 'fulfillment in time savings', and 'confidence in making wise choices' are classifications of needs made up of predetermined natural language. Other expressions may be used to express the needs. The instruction content (prompt) may be more detailed. In the embodiment, the instruction is to classify into one of the categories, but the instruction may be to classify into multiple categories. The number of categories is not limited to 10, and may be more or less than 10.

[0088] The generation AI server 2 classifies the given text (needs string information) into an appropriate classification based on the needs classification instruction information and returns it. The response from the generation AI server 2 contains the classified classification text (needs classification result information). The needs classification result information is stored in the n_class variable and returned as the output of the function.

[0089] (Code Block 4) df['n_class'] = df['review'].apply(lambda x: generate_n_class(x))

[0090] This program processes each row of text stored in the review column of the Pandas data frame df using the generate_n_class function, and adds the resulting classification (n_class) to a new column n_class in df.

[0091] The control unit 104 uses the apply method to add lambda x: to the character strings in each row of the df['review'] column, which represents an argument to be applied to x.

[0092] The control unit 104 passes the text to the generate_n_class function, classifies the text into one of the 10 categories, and assigns the result to the df['n_class'] column. Finally, the data frame df including the new n_class column is returned, as shown in FIG.

[0093] Fig. 9 shows a data frame to which the n_class column has been added by the processing of STEP 2. Note that Fig. 9 shows only 20 items (20 rows) due to space restrictions, but in the embodiment, a data frame df consisting of 100 items (100 rows) is generated. Also, the description of the review column is omitted because it was described in Fig. 8.

[0094] (STEP 2: Aggregation step) Next, the control unit 104 performs a process of tallying up a plurality of reviews (character string information) for each of a plurality of categories (needs categories) based on the category result information (needs category result).

[0095] (Code Block 5) df_n_class_counts = df['n_class'].value_counts().sort_index().reset_index() df_n_class_counts.columns = ["Need Classification", "Count"]

[0096] df['n_class'].value_counts(): This part calculates the frequency of occurrence of values ​​in a column called n_class in a dataframe df. The value_counts() function counts how frequently different values ​​occur in a dataframe and returns a new Dataframe or Series with those frequencies.

[0097] .sort_index(): This method sorts the obtained frequency data based on an index (here the n_class value). This is used to re-sort based on the original values ​​(n_class) since by default it is sorted based on the frequency of occurrence of the values.

[0098] .reset_index(): This method resets the index and converts the existing index to a regular column, which preserves the original index (the n_class value) as a column in the new DataFrame.

[0099] df_n_class_counts.columns = ["Needs Class", "Count"]: Finally, we rename the columns in the newly created dataframe df_n_class_counts. We name the first column (which contains the original n_class values) as "Needs Class", and the second column with frequency values ​​as "Count". (STEP 3: Visualization step) Next, the control unit 104 visualizes the tallying results. The program is as follows.

[0100] (Code Block 6) ! pip install japanize-matplotlib import matplotlib.pyplot as plt import japanize_matplotlib plt.figure(figsize=(10, 6)) plt.plot(df_n_class_counts['Needs classification'], df_n_class_counts['Number of cases']) plt.title('Count by Needs Category') plt.xlabel('Needs classification') plt.ylabel('number of items') plt.show()

[0101] !pip install japanize-matplotlib: This is the command to install the japanize-matplotlib package, which is used to display Japanese characters correctly in matplotlib.

[0102] import matplotlib.pyplot as plt: Here, we import the matplotlib pyplot module with the name plt.

[0103] import japanize_matplotlib: This line imports the japanize-matplotlib library that you installed earlier.

[0104] plt.figure(figsize=(10, 6)): This specifies the size of the graph. It creates a figure that is 10 inches wide and 6 inches high.

[0105] plt.plot(df_n_class_counts['Needs Classification'], df_n_class_counts['Number of Cases']): This line plots the Needs Classification column of the df_n_class_counts data frame on the x-axis and the Number of Cases column on the y-axis. This will display a line chart showing the relationship between Needs Classification and Number of Cases.

[0106] plt.title('Number of items by need category'): This sets the title of the graph. The title will be "Number of items by need category".

[0107] plt.xlabel('Needs Classification') and plt.ylabel('Count'): These lines set the labels for the x-axis and y-axis, respectively. The x-axis label is "Needs Classification" and the y-axis label is "Count".

[0108] plt.show(): Finally, this line is used to display the graph we created with matplotlib. This will display the graph on the screen based on the settings above.

[0109] The control unit 104 displays a graph on a display device (not shown) such as a liquid crystal display connected via the external I / F 106 (FIG. 10).

[0110] Also, the visualization is not limited to this, but may be a bar graph. Various forms of visualization are possible using matplotlib.

[0111] In the above embodiment, processing using Python was described, but this is not limited to this, and similar processing can also be performed by combining a spreadsheet such as Excel with ChatGPT or an API.

[0112] In addition, the classification categories are not limited to problems and solutions or needs, but may be classified into categories of quality, effect, use, and benefit in natural language according to the purpose of the analysis. Furthermore, the character strings to be classified are not limited to patent data, and may be any character string in natural language, such as a research paper.

[0113] Although the embodiment has been described above, this embodiment is presented as an example and is not intended to limit the scope of the invention. This new embodiment can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the gist of the invention. This embodiment and its modifications are included in the scope and gist of the invention, and are included in the scope of the invention and its equivalents described in the claims. [Explanation of symbols]

[0114] 1 terminal, 2 generation AI server, 101 CPU, 102 ROM, 103 RAM, 104 control unit, 105 HDD, 106 external I / F, 107 input unit, 108 system bus

Claims

1. A computer comprising:

1. A classification processing program that executes a process of inputting, into a large-scale language model, each of a plurality of character strings related to a specific object and classification instruction information including an instruction to classify the plurality of character strings into one of a plurality of classifications expressed in natural language, and obtaining, from the large-scale language model, classification result information in which each of the plurality of character strings is classified into one of the plurality of classifications expressed in natural language.

2. The classification processing program as described in claim 1, further comprising a process of visualizing the aggregated results of the classification of each of the plurality of character strings based on the classification result information.

3. A computer comprising:

1. A classification processing method comprising: inputting, into a large-scale language model, each of a plurality of character strings related to a specific object and classification instruction information including an instruction to classify the plurality of character strings into one of a plurality of classifications expressed in natural language; and obtaining, from the large-scale language model, classification result information in which each of the plurality of character strings is classified into one of the plurality of classifications expressed in natural language.

Citation Information

Patent Citations

  • Sentence classification device, sentence classification method and program

    JP2007004233A

  • Automatic information classification method, and information retrieval and analysis method

    JP2008112208A

  • Document classification apparatus, document classification method, and program

    JP2010146222A

  • Classification adding method and classification adding system

    JP2016206748A

  • Literature data analysis program and system

    JP2018049430A