Text classification system
The text classification system addresses inefficiencies in existing methods by integrating LLMs to automate document categorization with user-defined policies, ensuring high accuracy and adaptability for large datasets, including confidential content.
Patent Information
- Application Number
- JP2024064145
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-11
- Publication Date
- 2025-10-24
AI Technical Summary
Conventional rule-based and machine learning methods for document classification require significant manual effort and large training datasets, while generative AI (LLMs) face efficiency and accuracy issues with long prompts.
A text classification system that integrates with LLMs to minimize user input, allowing seamless classification of large document datasets by proposing categories and subcategories based on user-defined policies, with options for manual correction and expansion.
Enables efficient and accurate classification of large volumes of documents, including confidential data, by leveraging LLMs with minimal user interaction and adaptability.
Smart Images

Figure 2025161176000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a technology for processing natural language data, and more particularly to a technology that is effective when applied to a text classification system for classifying documents written in natural language. [Background technology]
[0002] For example, in order to analyze data sources such as documents and text data accumulated within a company, the documents and text data are classified and categorized.
[0003] As an example of a technology related to document classification and analysis, Japanese Patent Application Laid-Open No. 2023-62700 (Patent Document 1) describes the following: extracting a semantic structure from each of a plurality of documents and creating an index of the semantic structure; accepting from the user or automatically specifying semantic structure conditions based on the perspective the user wants to know based on the index; displaying documents that meet the specified semantic structure conditions to the user based on rule-based or machine learning-based processing; accepting from the user the selection of one or more documents and the specification of classification tags that tag the documents; associating the specified classification tags with each of the selected one or more documents; then searching for documents similar to at least one document associated with the specified classification tag from among the plurality of documents that do not have at least one classification tag associated; displaying similar documents to the user if any are found; and accepting the selection of the documents and the specification of classification tags for the similar documents. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Publication No. 2023-62700 Summary of the Invention [Problem to be solved by the invention]
[0005] When classifying large amounts of documents and text data stored within a company, rule-based classification, as in conventional technology, requires manual preparation of rules for classification direction, keywords, etc., which is a heavy workload. Even when using machine learning for classification, supervised learning requires the preparation of a sufficient amount of training data, which is still a high hurdle.
[0006] On the other hand, classification using generative AI (Artificial Intelligence) and large language models (LLMs) (hereinafter referred to as "LLMs"), such as ChatGPT (registered trademark), which has seen rapid growth in use in recent years, is also being considered.
[0007] However, to classify large volumes of documents or text data using LLM, the target text data must be written individually in prompts and sorted by LLM, which does not provide high processing efficiency compared to classification using conventional rule-based or machine learning technologies.Although it is possible to pass a large amount of text data to LLM at once and obtain classification results in one go, it is known that the longer and more complex the prompts, the less stable the processing accuracy of LLM becomes.
[0008] Therefore, an object of the present invention is to provide a text classification system that can efficiently and effectively classify large amounts of documents and text data using LLM.
[0009] The above and other objects and novel features of the present invention will become apparent from the description of this specification and the accompanying drawings. [Means for solving the problem]
[0010] Among the inventions disclosed in this application, the outline of representative inventions will be briefly explained as follows.
[0011] A representative embodiment of the present invention is a text classification system that classifies documents stored in a document database into categories. The system includes a classification processor that causes an LLM to propose one or more categories based on a classification policy specified by a user, a search processor that searches the document database for documents belonging to each of the proposed categories, and a UI processor that presents the user with each category and the number of documents belonging to each category. [Effects of the Invention]
[0012] The effects obtained by the representative inventions disclosed in this application can be briefly explained as follows.
[0013] That is, according to the representative embodiment of the present invention, it becomes possible to efficiently and effectively classify large amounts of documents and text data using LLM. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a diagram showing an overview of an example of the configuration of a text classification system according to an embodiment of the present invention. [Figure 2] 1 is a flowchart outlining an example of a flow of a document classification process according to an embodiment of the present invention. [Figure 3] FIG. 10 is a diagram showing an outline of an example of a screen displayed on a user terminal according to an embodiment of the present invention. [Figure 4] 10 is a flowchart outlining an example of the flow of a category creation process according to an embodiment of the present invention. [Figure 5] FIG. 10 is a diagram outlining another example of a screen displayed on a user terminal according to an embodiment of the present invention. [Figure 6] 10 is a flowchart outlining an example of the flow of a subcategory creation process according to an embodiment of the present invention. [Figure 7]FIG. 10 is a diagram outlining another example of a screen displayed on a user terminal according to an embodiment of the present invention. [Figure 8] 10 is a flowchart outlining an example of the flow of a category re-creation process according to an embodiment of the present invention. [Figure 9] FIG. 10 is a diagram outlining another example of a screen displayed on a user terminal according to an embodiment of the present invention. [Figure 10] 10 is a flowchart outlining an example of the flow of a parallel category creation process according to an embodiment of the present invention. [Figure 11] FIG. 10 is a diagram outlining another example of a screen displayed on a user terminal according to an embodiment of the present invention. [Figure 12] 1 is a flowchart outlining an example of a flow of a process for creating categories from text in one embodiment of the present invention. [Figure 13] FIG. 10 is a diagram outlining another example of a screen displayed on a user terminal according to an embodiment of the present invention. [Figure 14] FIG. 10 is a diagram outlining another example of a screen displayed on a user terminal according to an embodiment of the present invention. [Figure 15] FIG. 10 is a diagram outlining another example of a screen displayed on a user terminal according to an embodiment of the present invention. [Figure 16] FIG. 10 is a diagram outlining another example of a screen displayed on a user terminal according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0015] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In all drawings used to explain the embodiments, the same parts are generally designated by the same reference numerals, and repeated explanations will be omitted. However, parts that have been designated by reference numerals in one drawing may be referred to by the same reference numerals in other drawings, although they will not be shown again.
[0016] <Summary> The approach of dividing and classifying accumulated documents and text data into categories can be used in a variety of situations. For example, by classifying "customer feedback" accumulated through inquiries to call centers, responses to questionnaires, and word-of-mouth on websites, it is possible to analyze what customers like and dislike.
[0017] The text classification system, which is one embodiment of the present invention, uses LLM to classify large amounts of documents and text data (hereinafter referred to as "documents") stored within a company or other organization into categories. By minimizing user input and seamlessly linking the stored documents with LLM, the system enables efficient and effective document classification.
[0018] LLM cannot handle content other than that contained in the data it has been trained on. Therefore, for example, even if an attempt is made to process confidential information, such as information about a company's own products or various internal data, which requires confidentiality and is not appropriate for training on an external LLM, the information cannot be processed because the LLM has not been trained. The text classification system, which is one embodiment of the present invention, seamlessly integrates with LLM to enable classification, including for such confidential documents and text data.
[0019] <System configuration> Figure 1 is a diagram showing an overview of an example configuration of a text classification system according to one embodiment of the present invention. The text classification system 1 is configured, for example, from a server device or a virtual server built on a cloud computing service, and uses a central processing unit (CPU) (not shown) to execute middleware such as an operating system (OS), database management system (DBMS), and web server program, which are deployed to memory from a storage device such as a hard disk drive (HDD) or solid state drive (SSD), as well as software running on the OS and DBMS, thereby realizing the function of classifying documents stored within a company or other location using LLM2.
[0020] The text classification system 1 includes various components implemented as software, such as a user interface (UI) processor 11, a classification processor 12, and a search processor 13. It also includes various data stores, such as a document database (DB) 14 and a classification DB 15, implemented as databases, file tables, etc.
[0021] The UI processing unit 11 has a function of providing a user interface for classifying documents stored in the document DB 14 to a user terminal 3 such as a PC (Personal Computer) connected via a network such as the Internet or a LAN (Local Area Network), not shown. In this embodiment, the documents stored in the document DB 14 are classified into categories by a classification processing unit 12 (to be described later) based on instructions from a user via the user terminal 3, and the results are displayed in a tree format. Examples of screens displayed on the user terminal 3 by the UI processing unit 11 will be described later. Although document DB 14 may store confidential documents accumulated within a company, for example, the document DB 14 is not limited to such documents and may also store non-confidential documents. Examples of documents stored in document DB 14 include text transcripts of call center call logs, various survey data, review documents, articles about the company's products and services, patent documents, and other documents. In addition to the main text of the documents, document attribute information and vectorized data of the documents using known technology may also be stored.
[0022] The classification processing unit 12 has a function of classifying documents stored in the document DB 14 into categories by querying the LLM 2 based on instructions from the user terminal 3 (user) via the UI processing unit 11, and registering information related to the classification results in the classification DB 15. Note that the classification DB 15 stores and manages, for example, information on the categories to which each document in the document DB 14 belongs (information related to the association between the document and the category).
[0023] The search processing unit 13 has the function of searching the document DB 14 and classification DB 15 to obtain documents to be classified and information related to these documents (e.g., document attribute information, statistical information, etc.) when the classification processing unit 12 queries the LLM2.
[0024] <Processing flow> 2 is a flowchart outlining an example of the flow of document classification processing in one embodiment of the present invention. In this embodiment, as described above, documents stored in document DB 14 are classified into categories based on instructions from the user, and these are displayed in a tree format. First, UI processing unit 11 determines whether a category tree exists, i.e., whether the state is initial (S01). If classification processing has been performed at least once and a category tree already exists (No in step S01), the process proceeds to step S04, which will be described later.
[0025] On the other hand, if a category tree does not yet exist (Yes in step S01), the system is in an initial state, and first, input of information about the type of data of the documents to be classified (which may be information related to the classification policy) is accepted from the user (S02). This allows the LLM2 to be instructed about the types of documents stored in the document DB 14, or about what documents in the document DB 14 should be classified. If the user has an intention regarding the document classification policy, it is possible to reflect that intention in the instructions to the LLM2.
[0026] 3 is a diagram outlining an example of a screen displayed on the user terminal 3 in one embodiment of the present invention. The example of Fig. 3 shows a state in which the user has entered "classification related to cosmetics reviews" as information on the document to be classified (or information related to the classification policy) on the screen displayed on the user terminal 3.
[0027] Returning to FIG. 2, when the user inputs information about the documents to be classified (or information about the classification policy) in step S02, the classification processing unit 12 performs a category creation process to create categories from the documents in the document DB 14 according to this information (S03).
[0028] FIG. 4 is a flowchart outlining an example of the flow of a category creation process in one embodiment of the present invention. In the category creation process, first, the classification processing unit 12 creates a prompt to be input to the LLM2 based on instructions from the user, and queries the LLM2 (S31). The prompt is created, for example, by modifying a pre-prepared template based on instructions from the user, thereby minimizing user input. For example, if classification information such as that shown in the example of FIG. 3 is input in step S02 of FIG. 2, a prompt such as "Please output 20 classification categories related to cosmetics reviews, along with keywords and synonyms representing each classification category, in a JSON format list," is created and input to the LLM2.
[0029] In response to the input prompt, the LLM 2 sends a suggested category to the classification processing unit 12 (S32). For example, in response to the input prompt in the above example, { Id:'0', category:'price', keywords:['price', 'cost', 'cost', ...] } … The response is in JSON (JavaScript Object Notation) format, containing the category name ('price') and a list of keywords that correspond to that category ('price', 'expense', 'cost', ...).
[0030] In response to the response from the LLM 2, the classification processing unit 12 displays the proposed categories in a tree format on the screen of the user terminal 3 via the UI processing unit 11 (S33). Furthermore, for each category, the search processing unit 13 searches the document DB 14 based on the relevant keywords, and the number of documents found is also displayed via the UI processing unit 11 (S34). Then, information relating to the association between each category and the documents belonging to it (documents found by searching the document DB 14) is registered in the classification DB 15 (S35), and the category creation process is terminated.
[0031] Fig. 5 is a diagram outlining another example of a screen displayed on the user terminal 3 in an embodiment of the present invention. The example in Fig. 5 shows a state in which categories proposed by the LLM2 ('price', 'effectiveness', 'quality', ...) are displayed in a hierarchical tree structure below the classification policy input by the user. Each category also displays the number of documents belonging to it (the number of hits found by searching the document DB 14). In the example in Fig. 5, the categories are displayed in the order in which they are included in the JSON data returned from the LLM2, but they may also be sorted and displayed by the number of documents, etc.
[0032] In this embodiment, when searching for documents belonging to each category in step S34 in the example of Fig. 4, a keyword search is performed based on keywords corresponding to the category, but this is not limited to this. For example, a representative document that clearly shows the characteristics of each category may be set or selected, vectorized, and then a vector search may be performed in the document DB 14 to obtain similar documents.
[0033] 2, when a category tree is displayed on the screen of the user terminal 3 in step S03, the UI processing unit 11 accepts user operations for the category selected or specified by the user (hereinafter, sometimes referred to as "category X") (S04). Then, in response to the user operations, the classification processing unit 12 performs various processes, such as a subcategory creation process (S05), a category re-creation process (S06), a parallel category creation process (S07), and a category creation process from text (S08). These processes will be described later.
[0034] Thereafter, it is determined whether or not the user has instructed to end the process, such as by closing the screen (S09), and if the user has instructed to end the process (Yes in step S09), the process ends. On the other hand, if the user has not instructed to end the process (No in step S09), the process returns to step S04, and further user operations for categories are accepted.
[0035] FIG. 6 is a flowchart outlining an example of the flow of the subcategory creation process (step S05 in FIG. 2) in one embodiment of the present invention. In this process, subcategories are created for category X selected or specified by the user from among the categories displayed in a tree view. In this embodiment, a subcategory for category X can be created by performing a process similar to the category creation process described above in FIG. 4 (S51), with category X as information related to the classification target (or classification policy). In this case, the prompt input to LLM2 might be, for example, "Please output 10 subcategories for the classification item 'price' related to cosmetics reviews in a JSON format list."
[0036] Fig. 7 is a diagram outlining another example of a screen displayed on the user terminal 3 in an embodiment of the present invention. In the example of Fig. 7, the user presses the Create Subcategory button (a button with a + inside a circle) to the right of the 'Price' category, and the categories ('Price Range', 'Value for Money', 'Discount') proposed by LLM2 as subcategories of the 'Price' category are displayed in a tree-like format along with the number of documents. In this way, the tree can be expanded by creating subcategories.
[0037] FIG. 8 is a flowchart outlining an example of the flow of the category re-creation process (step S06 in FIG. 2) in one embodiment of the present invention. For example, even if subcategories are created using the category creation process shown in FIG. 6 above, there may be cases where the subcategories proposed by LLM2 are inappropriate or not what the user intended. In such cases, this process recreates (retries) a subcategory for category X selected or specified by the user, replacing the existing subcategory.
[0038] First, the classification processing unit 12 deletes existing subcategories of category X (S61). Then, by using category X as information related to the classification target (or classification policy) and performing processing similar to the category creation processing described above with reference to FIG. 4 (S62), subcategories of category X can be recreated.
[0039] 9 is a diagram outlining another example of a screen displayed on the user terminal 3 in one embodiment of the present invention. In the example of FIG. 9, the user has pressed the retry button (a button with a looping arrow) to the right of the 'Price' category, and the categories ('EC', 'In-Store', 'Campaign') newly proposed by LLM2 as subcategories of the 'Price' category are displayed in a tree structure along with the number of documents. In this way, by allowing the user to recreate subcategories, the user can repeatedly recreate them until they have created a subcategory that they deem practical and effective.
[0040] FIG. 10 is a flowchart outlining an example of the flow of the parallel category creation process (step S07 in FIG. 2) in one embodiment of the present invention. In this process, an additional category is created parallel to an existing subcategory for category X selected or specified by the user. As with the category creation process in FIG. 4 described above, first, the classification processing unit 12 creates a prompt for creating parallel subcategories for category X and queries the LLM2 (S71). For example, based on a pre-prepared template, a prompt such as "Please output a list of categories in JSON format that are parallel to the subcategories 'price range', 'cost performance', and 'discount' of the 'price' of cosmetics reviews" is created and input to the LLM2.
[0041] In response to the input prompt, the LLM 2 sends the proposed additional categories to the classification processing unit 12 (S72). In response to the response from the LLM 2, the classification processing unit 12 displays the proposed additional categories in a tree format on the screen of the user terminal 3 via the UI processing unit 11 (S73). Furthermore, for each category, the search processing unit 13 searches the document DB 14 based on the relevant keywords, and displays the number of documents found via the UI processing unit 11 (S74). Then, information relating to the association between each category and the documents belonging to it is registered in the classification DB 15 (S75), and the parallel category creation process ends.
[0042] Fig. 11 is a diagram outlining another example of a screen displayed on the user terminal 3 in an embodiment of the present invention. The example in Fig. 11 shows a state in which, in a state in which 'Price Range', 'Value for Money', and 'Discount' have been created as subcategories of the 'Price' category as in the example in Fig. 7 above, the user presses the subcategory creation button (a button with a + inside a circle) to the right of the 'Price' category, and as a result, categories ('Sale', 'Expensive', 'Cheap') additionally suggested by LLM2 as subcategories of the 'Price' category are displayed in a tree format along with the number of documents.
[0043] As an intermediate process between the re-creation of subcategories shown in Fig. 8 and the additional creation of parallel categories shown in Fig. 10, for example, it may be possible to leave only some of the existing subcategories designated by the user and re-create the remaining categories (additional creation of categories parallel to the remaining categories).It may also be possible to designate unnecessary categories and delete them individually.
[0044] 12 is a flowchart outlining an example of the flow of a category creation process from text (step S08 in FIG. 2) in one embodiment of the present invention. In this process, subcategories are created in a bottom-up manner based on the body text of documents belonging to category X selected or specified by the user. First, for category X specified by the user, the search processing unit 13 searches the document DB 14 to obtain the body text of documents belonging to category X, and this is displayed on the screen of the user terminal 3 via the UI processing unit 11 (S81).
[0045] Fig. 13 is a diagram outlining another example of a screen displayed on the user terminal 3 in one embodiment of the present invention. In the example of Fig. 13, the user has selected the 'Price Range' category from among the subcategories of the 'Price' category, and the body text of each document belonging to the 'Price Range' category is displayed on the right side of the screen. In this way, by selecting a category, the user can refer to and check the body text of each document associated with that category.
[0046] 12, it is determined whether the user has instructed on the screen displayed on the user terminal 3 in step S81 to create a subcategory based on the main text (S82). This instruction is given, for example, by the user pressing the "Suggest Category" button in the upper right corner of the screen in the example screen shown in Fig. 13 above. If an instruction to create a subcategory from the main text has not been given (No in step S82), the process of creating a category from text is terminated.
[0047] On the other hand, if an instruction to create a subcategory from the main text is given (No in step S82), similar to the category creation process of Fig. 4 described above, first, the classification processing unit 12 creates a prompt for creating a subcategory from the main text and queries the LLM2 (S83). For example, along with the contents of the main text, a prompt such as "Please output 10 subcategories that correspond to these texts in a JSON format list" is created and input to the LLM2.
[0048] The LLM 2 responds to the classification processing unit 12 with the proposed categories in accordance with the input prompt (S84). The classification processing unit 12 receives the response from the LLM 2 and displays the proposed categories in a tree format on the screen of the user terminal 3 via the UI processing unit 11 (S85). Furthermore, for each category, the search processing unit 13 searches the document DB 14 based on the relevant keywords, etc., and displays the number of documents found via the UI processing unit 11 (S86). Then, information relating to the association between each category and the documents belonging to it is registered in the classification DB 15 (S87), and the process of creating categories from text is completed.
[0049] 14 is a diagram outlining another example of a screen displayed on the user terminal 3 in one embodiment of the present invention. In the example of FIG. 14, the user presses the "Suggest Category" button in the upper right corner of the screen to instruct the creation of subcategories from the main text, and the categories ('Expensive,' 'Cheap,' 'Affordable,' 'Other Products') proposed by the LLM2 based on the main text (each document belonging to the 'Price Range' category) displayed on the right side of the screen are displayed in a tree format along with the number of documents. In this way, by making it possible to create categories based not only on keywords but also on the main text, it is possible to create highly accurate categories that are in line with the documents stored in the document DB 14.
[0050] 12 to 14 above show an example of creating subcategories from the main text, but this is not limited to subcategories; higher-level categories, parallel categories, etc. may also be created based on the main text of each document. For example, create a prompt such as "Please output 10 categories based on the target text in a JSON format list" and input it into LLM2.
[0051] Various variations in the method of instructing category creation are possible. For example, if you have 100 documents related to geology and provide the text information of these 100 documents to LLM2, you can create a prompt such as "Based on these texts, please output a list in JSON format of 10 categories that appropriately classify these 100 documents, based on the general academic classification of geology," or "Based on these texts, please output a list in JSON format of 10 categories that appropriately classify these 100 documents," and input this to LLM2 to instruct it to create categories.
[0052] FIG. 15 is a diagram outlining another example of a screen displayed on the user terminal 3 in an embodiment of the present invention. The example in FIG. 15 shows that when subcategories and parallel categories are created by the LLM2 in the processing of steps S03 and S05 to S08 in FIG. 2 described above, the user can specify the direction of subcategory creation in advance. The example in the figure shows a state in which a user is instructed to classify reviews into "complaints." When the user presses the "Save" button to specify the direction, and then presses, for example, the subcategory creation button (a button with a + inside a circle) to the right of the category, the classification processing unit 12 creates a prompt that instructs the suggestion of a category in a form that reflects the specified direction and inputs it to the LLM2. This makes it possible to suggest categories categorized in accordance with the user's intended direction.
[0053] Fig. 16 is a diagram outlining another example of a screen displayed on a user terminal in one embodiment of the present invention. The example in Fig. 16 shows a screen for managing keywords registered in the "price" category. Via this screen, the user can add or modify keywords used in document searches as needed, or delete unnecessary keywords, and then re-search the document DB 14 according to the results.
[0054] Furthermore, to improve the comprehensiveness of searches, the LLM2 may suggest additional keywords. For example, when the "Expand Keywords" button on the screen of FIG. 16 is pressed, the classification processing unit 12 creates a prompt (e.g., "Please suggest some synonyms and similar words for these keywords: price, expense, cost, price") that responds with a list of synonyms or similar words based on one or more keywords displayed in the input field, and inputs this into the LLM2. The LLM2 responds with a list of synonyms or similar words according to the instructions of the input prompt, and the result is reflected in the input field via the UI processing unit 11. The user can modify or delete the content reflected in the input field as needed, and can also suggest additional keywords by pressing "Expand Keywords" again as needed.
[0055] As described above, text classification system 1, which is one embodiment of the present invention, can use LLM2 to classify documents, including confidential documents and text data stored within a company, into categories with high accuracy. Furthermore, by minimizing user input and seamlessly linking documents stored in document DB 14 with LLM2, documents can be classified efficiently and effectively.
[0056] The invention made by the inventor has been specifically described above based on the embodiments, but it goes without saying that the present invention is not limited to the above embodiments and can be modified in various ways without departing from the spirit of the invention. Furthermore, the above embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to those having all of the described configurations. Furthermore, it is possible to add, delete, or replace part of the configuration of the above embodiments with other configurations.
[0057] Furthermore, the above-described configurations, functions, processing units, processing means, etc. may be partially or entirely implemented in hardware, for example, by designing them as integrated circuits. The above-described configurations, functions, etc. may also be implemented in software, with a processor interpreting and executing a program that implements each function. Information such as the programs, tables, and files that implement each function can be stored in a storage device such as a memory, hard disk, or SSD, or in a storage medium such as an IC card, SD card, or DVD.
[0058] In addition, in the above figures, the control lines and information lines shown are those that are considered necessary for explanation, and do not necessarily show all the control lines and information lines that are actually implemented. In reality, it can be assumed that almost all components are interconnected. [Industrial Applicability]
[0059] The present invention can be used in a text classification system that classifies documents written in natural language. [Explanation of symbols]
[0060] 1...Text classification system, 2...LLM, 3...User terminal, 11...UI processing unit, 12...classification processing unit, 13...search processing unit, 14...document DB, 15...classification DB
Claims
1. A text classification system that classifies documents stored in a document database into categories, comprising: a classification processor that causes a large-scale language model (hereinafter referred to as "LLM") to propose one or more categories based on a classification policy specified by a user; a search processing unit that searches the document database for documents belonging to each of the proposed categories; a UI processing unit that presents each of the categories and the number of documents belonging to each of the categories to the user.
2. 10. The text classification system of claim 1, The classification processing unit records information relating to associations between each of the categories and documents belonging to each of the categories in a classification database.
3. 10. The text classification system of claim 1, The classification processing unit causes the LLM to propose one or more subcategories for each of the categories specified by the user.
4. 10. The text classification system of claim 1, A text classification system in which the classification processing unit causes the LLM to additionally propose one or more subcategories to existing subcategories for each of the categories specified by the user.
5. 10. The text classification system of claim 1, The classification processor causes the LLM to propose one or more subcategories based on the body text of documents belonging to each category.
6. 10. The text classification system of claim 1, A text classification system in which the classification processing unit inputs instructions regarding the classification direction specified by the user when causing the LLM to propose categories.
7. 10. The text classification system of claim 1, The classification processing unit causes the LLM to suggest keywords that fall into each of the categories.
Citation Information
Patent Citations
Document analysis support system and method
JP2023062700A
Cited By
Information processing systems, information processing methods, and programs
JP7845734B1
Information processing system, information processing method, and program
JP7845735B1
Information processing systems, information processing methods, and programs
JP7853746B1