Information processing device, information processing method and computer program
The information processing device uses deep learning to predict and visually highlight candidate classifications based on scores, improving the efficiency and accuracy of classification assignment by linking them to document text, addressing the inefficiencies of conventional systems.
Patent Information
- Application Number
- JP2024039207
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-13
- Publication Date
- 2025-09-29
AI Technical Summary
Conventional classification systems struggle with efficiently assigning multiple classifications to a single document using machine learning, as they often require analyzing the entire classification table and lack models trained on relevant document-text combinations.
An information processing device with a classification estimation unit that uses deep learning to predict candidate classifications and scores, accompanied by a display processing unit that changes display modes based on scores, highlighting relevant classifications and their text associations.
Enhances the accuracy and efficiency of classification assignment by prioritizing high-scoring classifications and visually linking them to document text, facilitating quicker and more accurate classification decisions.
Smart Images

Figure 2025140053000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an information processing device, an information processing method, and a computer program for assisting in the classification of documents. [Background technology]
[0002] Traditionally, classifications have been assigned to target objects so that the target object can be easily found among other objects. For example, when the target object is a patent document, a patent classification such as the International Patent Classification (IPC), which is used internationally, or the FI and F-terms, which are classifications unique to Japan, is assigned to each patent document. Also, when the target object is a book stored in a library, a book classification is assigned to each book.
[0003] For example, when assigning a patent classification to a patent application, the classifier understands the content of the patent application, extracts a classification that corresponds to the content of the patent application from a classification table that defines patent classifications, and assigns the extracted classification to the patent application.In recent years, machine learning has been used to predict the patent classification that should be assigned to a patent application, and the classifier then determines the classification to be assigned by comparing the predicted classification with the content of the patent application.
[0004] Non-Patent Document 1 discloses that in classification work, a classification support system uses methods such as machine learning to present candidate FI / F-terms to a classifier for patent documents to be classified. The system displays a screen presenting the basis for classification, including the patent document text and an F-term information field, and when a paragraph in the patent document text is selected, the F-terms to be assigned to the selected paragraph are displayed in the F-term information field in order of score, and the F-term is assigned by checking the displayed F-term.
[0005] Patent Document 1 discloses a classification support device that can be used when classifying patent documents and the like using F-terms, etc., in which words that appear in the patent document are highlighted in an F-term definition table, and words that appear in the F-term definition table are highlighted in the patent document, making it easy to see which areas need to be checked intensively.
[0006] Patent Document 2 discloses that in a classification support system using a document processing device that displays multiple term groups or multiple documents in a list format on a display screen, by associating multiple term groups with multiple classifications, it is possible to grasp the content of the entire document and quickly find the basis for classification.
[0007] Patent Document 3 discloses that in a patent classification device that classifies patent information, in order to be able to classify patent documents with high accuracy, a learning device is trained using important information such as effect / solution means statements, effect statements, solution means statements, objective statements, effect terms, solution means terms, objective terms, characteristic words, and related words of characteristic words obtained using patent specifications, as well as classification result information that identifies the manual classification results of patents. [Prior art documents] [Patent documents]
[0008] [Patent Document 1] Japanese Patent Application Laid-Open No. 2008-171164 [Patent Document 2] Japanese Patent Publication No. 2022-103710 [Patent Document 3] Patent Publication No. 2021-036427 [Non-patent literature]
[0009] [Non-Patent Document 1] FY2017 Empirical Research Project Survey Report for the Practical Use of F-Term Granting Support Systems, March 30, 2018, Hitachi, Ltd. Summary of the Invention [Problem to be solved by the invention]
[0010] Conventional technologies display on a screen the correspondence between descriptions in a document that form the basis for classification and the candidate classification to be assigned, or display the correspondence between the document to be classified and words in a classification table or groups of terms associated with multiple classifications. However, when assigning classifications, there is a need to analyze the classification to be assigned while viewing the entire classification table as well as the estimated candidate classification. Furthermore, conventional machine learning trains a learning model using a single document and the classification assigned to that document as training data. However, when a single document is assigned multiple classifications from various perspectives, such as patent classification, it is not possible to train a learning model to learn multiple classifications for a single document. Therefore, there is a need to create a learning model suitable for patent classification.
[0011] The present disclosure has been made in consideration of the above circumstances, and aims to provide a classification support system that has a user interface that can efficiently assign classifications to documents to be classified, and that supports highly accurate classification. [Means for solving the problem]
[0012] An information processing device of a first aspect of the present disclosure is an information processing device for assisting in the assignment of classifications to documents, and is characterized by comprising: a classification estimation unit that estimates candidate classifications that are candidates for a classification to be assigned to the document and a classification score that indicates the degree of likelihood that the candidate classification will be assigned; a display processing unit that displays a classification list including the candidate classifications; and a classification display mode change unit that changes the display mode of the candidate classifications included in the classification list.
[0013] In the information processing device of a second aspect of the present disclosure, the category display mode change unit changes the display mode of the candidate categories included in the category list in accordance with the category score.
[0014] The information processing device of the third aspect of the present disclosure is characterized by further comprising a first classification display processing unit that displays only candidate classifications whose classification scores are higher than a predetermined threshold.
[0015] The information processing device of the fourth aspect of the present disclosure is characterized by further comprising a second classification display processing unit that displays candidate classifications in descending order of the classification score.
[0016] The information processing device of the fifth aspect of the present disclosure further includes a text display processing unit that displays the text of the document, and the text display processing unit is characterized in that it highlights characteristic words related to the candidate classification in the text of the displayed document.
[0017] An information processing device according to a sixth aspect of the present disclosure is characterized in that the category estimation unit estimates the candidate category and the category score using deep learning.
[0018] In a seventh aspect of the information processing device of the present disclosure, the classification estimation unit makes an estimation using a learning model that, when given text of a document, outputs a plurality of candidate classifications that are candidates for classification to be assigned to the text of the document, and a classification score that indicates the degree of likelihood that the candidate classification will be assigned, and the learning model is trained using a classification assigned to a specific document and text from the text of the specific document that is similar to a sentence that defines the classification.
[0019] The information processing device of the eighth aspect of the present disclosure is characterized in that it is trained using the text of the specific document and multiple classifications assigned to the specific document. [Effects of the Invention]
[0020] According to a first aspect of the present disclosure, by changing the display mode of estimated candidate classifications in a classification list that displays part or all of a classification table, it is possible to easily identify candidate classifications that should be noted and their classification scores while obtaining an overview of the classification list as a whole.
[0021] According to a second aspect of the present disclosure, when analyzing the classification to be assigned while visually viewing a classification list, the display mode of the candidate classification estimated by machine learning is changed according to its classification score, making it easy to identify the candidate classification that should be given priority consideration among the candidate classifications displayed in the classification list.
[0022] According to the third aspect of the present disclosure, even if a large number of candidate classifications are output by machine learning, the candidate classifications that are most likely to be assigned to a document can be analyzed preferentially, thereby enabling efficient classification assignment.
[0023] According to the fourth aspect of the present disclosure, candidate classifications output by machine learning can be analyzed in order of likelihood of being assigned to a document, enabling efficient classification.
[0024] According to the fifth aspect of the present disclosure, the candidate classification, which is the result of machine learning, can easily identify the basis for its assignment in the document text, i.e., which description in the document text it is related to, so that the validity of the candidate classification can be easily analyzed while looking at the document text.
[0025] According to the sixth aspect of the present disclosure, it is possible to predict candidate classifications with higher accuracy by using large amounts of document data, which can further improve the accuracy of classification assignment and contribute to the efficiency of classification assignment.
[0026] According to the seventh aspect of the present disclosure, when preparing teacher data to be used for training a learning model, the estimation accuracy of the learning model can be improved by using a combination of a classification and document text that is highly relevant to the classification as teacher data.
[0027] According to the eighth aspect of the present disclosure, even when multiple classifications are assigned to a single document, such as a patent document, the learning model is trained using specific text in the document and classifications that are highly relevant to the specific text as training data used for learning, thereby enabling candidate classifications to be estimated with high accuracy. [Brief explanation of the drawings]
[0028] [Figure 1] FIG. 1 is a block diagram schematically showing an example of the configuration of a classification support system 1 to which an information processing apparatus according to the first embodiment is applied. [Figure 2] FIG. 2 is a block diagram showing an example of the functional configuration of the classification assistance device 10 according to the first embodiment. [Figure 3] FIG. 3 is a diagram showing an example of machine learning by classification learning device 11 of the first embodiment. [Figure 4] FIG. 4 is a block diagram showing an example of the functional configuration of the classification terminal 12 of the first embodiment. [Figure 5] FIG. 5 is a diagram showing an example of the classification screen 5 according to the first embodiment. [Figure 6] FIG. 6 is a diagram showing an example of the feature word screen 6 according to the first embodiment. [Figure 7] FIG. 7 is a diagram showing an example of the text screen 7 of the first embodiment. [Figure 8] FIG. 8 is a flowchart illustrating the operation of the classification assistance device 10 of the first embodiment. [Figure 9] FIG. 9 is a flowchart illustrating the learning operation and classification estimation operation of classification learning device 11 of the first embodiment. [Figure 10] FIG. 10 is a diagram for explaining preprocessing of teacher data in classification learning device 11 of the third embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0029] Hereinafter, several embodiments of the present disclosure will be described in detail with reference to the drawings. In the following description, configurations and elements that are the same as or similar to configurations already described will be assigned the same reference numerals and will not be described again. Note that the following embodiments will be described using, as a non-limiting example, a case in which a patent classification is assigned to a patent document, but it goes without saying that the present disclosure can also be applied to cases in which classifications are assigned to documents other than patent documents.
[0030] First Embodiment (1) Overall configuration of classification support system 1 FIG. 1 is a bird's-eye view showing an example of the overall configuration of a classification support system 1 to which an information processing device, an information processing method, and a computer program according to a first embodiment of the present disclosure are applied. The classification support system 1 shown in FIG. 1 includes a classification support device 10, a classification learning device 11, classification terminals 12 and 13, etc. (hereinafter, only the classification terminal 12 will be described, and a description of the classification terminal 13 and other devices having a similar configuration will be omitted). The classification support device 10 is communicably connected to the classification learning device 11 and the classification terminal 12. Although FIG. 1 illustrates the classification support device 10, the classification learning device 11, and the classification terminal 12 as separate devices, these devices may be configured as a combined device, or each device may be distributed and configured as multiple devices. Furthermore, these devices may be communicably connected via a network.
[0031] The classification support device 10 is an information processing device for supporting the classification of documents. It generates information required for a classifier who assigns a classification to a document or a person involved in the classification (hereinafter also referred to as a "user") to assign a classification to a document, and causes a classification terminal 12 to display a user interface for supporting the classification. The classification support device 10 can display a classification screen 5, a feature word screen 6, or a text screen 7 on the classification terminal 12 as a user interface. The classification screen 5 is a screen for assigning a classification to a document selected by a user. The feature word screen 6 displays detailed information about an estimated candidate classification, such as feature words that are terms related to the candidate classification, and paragraphs in the document text where feature words frequently appear. The text screen 7 displays information such as candidate classifications, feature words, and assignment reasons in association with the text of the document selected by the user.
[0032] The classification support device 10 receives, via the classification screen 5, an instruction to select a document to which the user is going to assign a classification (hereinafter also referred to as a "document to be classified" or "document to be classified"), and estimates candidate classifications to be assigned to the selected document and a classification score for each candidate classification indicating the degree of likelihood that the candidate classification will be assigned. In the first embodiment, the classification support device 10 is described as estimating the candidate classifications and classification scores in cooperation with the classification learning device 11 that manages the learning model 111, but the classification support device 10 may also perform these estimations.
[0033] The classification score is an index that indicates the likelihood that a candidate classification will be assigned to a document that should be classified. The classification score can also be defined as the probability that a candidate classification will be assigned to a document that should be classified. The classification score ranges from 0 to 1, and the higher the classification score, the higher the probability of assignment.
[0034] The classification support device 10 acquires the candidate classifications and classification scores and displays a classification list including the candidate classifications on the classification screen 5. The classification list is a list of classifications to be assigned to the document to be classified, and displays classification symbols and their descriptions according to the contents of the classification table. The classification list displays part or all of the classification table. When only part of the classification table is displayed on the classification screen 5, the classification list may be displayed in a scrollable manner. The classification support device 10 changes the display mode of the candidate classifications included in the classification list based on the classification score. For example, for a classification that corresponds to a candidate classification in the classification list displayed on the classification screen 5, the background of the description of that classification is displayed so that it is different from the background of the descriptions of the other classifications, and further, the shade of the background is changed according to the classification score. This allows the existence of candidate classifications in the classification list and the degree of assignment possibility to be discerned at a glance.
[0035] The classification learning device 11 includes a learning model 111 that, when given a document text, outputs one or more candidate classifications that may be assigned to the document text, and a classification score indicating the degree of likelihood that each candidate classification will be assigned. The learning model 111 can be trained, for example, using a specific document and a classification assigned to the specific document as training data. The learning model 111 may use supervised learning, unsupervised learning, or reinforcement learning. Deep learning may also be used as machine learning.
[0036] The classification terminal 12 is an information processing device operated by a user. When activated by a user operation, the classification terminal 12 is communicably connected to the classification support device 10 and provides a user interface for classification on a display device 125. The classification terminal 12 displays a classification screen 5, a feature word screen 6, or a text screen 7 transmitted from the classification support device 10 on the display device 125. The classification terminal 12 also receives commands from the user via the classification screen 5 and transmits them to the classification support device 10.
[0037] (2) Hardware configuration The classification assistance device 10 includes a memory 102 that stores a computer program, a processor 101 that executes the computer program stored in the memory 102, a storage 103, a communication interface (communication I / F) 104, and an input / output interface (input / output I / F) 105. The processor 101 is typically a central processing unit (CPU) and / or a graphics processing unit (GPU), but may also be a microcomputer, a field programmable gate array (FPGA), a digital signal processor (DSP), or the like. The memory 102 temporarily stores programs executed by the processor 101 and data used by the processor 101. The storage 103 is a storage device also called an external storage device, and may include, for example, a non-volatile storage medium such as a hard disk drive (HDD) or a solid state drive (SSD).
[0038] Classification learning device 11 is an information processing device that includes a processor, memory, storage, a communication interface, and an input / output interface (not shown), is communicatively connected to classification assignment assistance device 10, and operates in cooperation with classification assignment assistance device 10. A learning model 111 generated by classification learning device 11 is recorded in the storage.
[0039] The classification terminal 12 includes a processor 121, a memory 122, a storage 123, an input device 124, a display device 125, and a communication interface (communication I / F) 126. The input device 124 is a device through which a user inputs various commands to the classification assistance device 10, and may include devices such as a mouse, a keyboard, a touch panel, and a microphone. The display device 125 displays a user interface for classification, and includes a liquid crystal display device, an organic EL display device, or the like.
[0040] (3) Functional configuration (3-1) Functional configuration of the classification support device 10 2 is a block diagram showing an example of the functional configuration of the classification support device 10 according to the first embodiment. The classification support device 10 includes a management unit 1010, a storage unit 1020, and a communication unit 1030. The communication unit 1030 controls communication between the classification support device 10 and the classification terminal 12 or classification learning device 11.
[0041] The storage unit 1020 includes a document data storage unit 1021 , a feature word storage unit 1022 , and a classification information storage unit 1023 . The document data storage unit 1021 stores data related to documents to be classified and documents that have already been classified. Documents broadly include written text, such as patent documents (hereinafter also referred to as "patent documents") and books. Patent documents include patent documents, i.e., published patent documents such as patent publications, published patent bulletins, and patent publications, as well as unpublished patent applications before being classified. Document data includes text data of the document and / or image data of drawings. In the case of patent documents, the document data includes text data of the claims, specification, and abstract, as well as image data of the drawings.
[0042] The feature word storage unit 1022 stores one or more feature words that characterize each classification, associated with each classification included in the classification table. Feature words are terms that serve as clues for assigning classifications, and consist of terms included in the definition sentences of each classification or any terms that characterize each classification. Feature words are associated with each classification, but feature words with similar meanings form a cross-classification group. For example, if a feature word associated with one classification is similar to a feature word associated with another classification, these feature words form a single group across classifications. Feature words are extracted in advance by analyzing patent classifications assigned to past patent documents and recorded in the feature word storage unit 1022.
[0043] The classification information storage unit 1023 stores a classification table for classifications to be assigned to documents. The classification table consists of at least a systematic classification list of multiple classification symbols and an explanation of each classification symbol (hereinafter also referred to as "definition text"). For example, patent classifications include the International Patent Classification (IPC), FI (File Index) classification, facet classification, and F-terms. The classification tables for these patent classifications specify classification symbols and their explanations (definition text). For example, the FI classification is based on the IPC, with the IPC display symbol followed by a subdivision symbol and a volume identification symbol. The FI classification table specifies FI classification symbols and their definition text. F-terms subdivide FI classifications into themes within a specific technical field from various technical perspectives so that they can be analyzed and assigned from multiple perspectives. For F-terms, an F-term list containing F-terms and their definition text is prepared for each theme code. The F-term list corresponds to a classification table.
[0044] The management unit 1010 functions as a display processing unit 1011, a classification estimation unit 1012, a classification display mode change unit 1013, a characteristic word display processing unit 1014, a text display processing unit 1015, and a classification assignment unit 1016 by the processor 101 executing a computer program stored in the memory 102.
[0045] The display processing unit 1011 generates a classification screen 5 that displays information necessary for classification, and performs processing to display the generated classification screen 5 on the display device 125 of the classification terminal 12. The display processing unit 1011 acquires the document numbers of documents to be classified from the document data storage unit 1021, and displays a list of document numbers in the document list area 50 of the classification screen 5. The display processing unit 1011 also receives candidate classifications and classification scores estimated by the classification estimation unit 1012, and acquires classification data of the classification table from the classification information storage unit 1023, and displays a classification list in the classification area 51 of the classification screen 5. When the display processing unit 1011 displays an initial screen of the classification list in the classification area 51, the category list may include candidate classifications.
[0046] The classification estimation unit 1012 works in cooperation with the classification learning device 11 to estimate candidate classifications and classification scores that may be assigned to document text. The classification estimation unit 1012 transmits document text data of a document to be classified to the classification learning device 11. The classification estimation unit 1012 receives, as classification estimation results by the classification learning device 11, one or more candidate classifications that are candidates for the classification to be assigned to the document, and the classification score for each candidate classification.
[0047] The classification display mode change unit 1013 changes the display mode of the classification list displayed in the classification area 51 of the classification screen 5 based on the classification score. That is, the classification display mode change unit 1013 changes the display mode of the candidate classifications included in the classification list according to the value of the classification score. When a classification corresponding to a candidate classification is displayed in the classification list, the classification display mode change unit 1013 changes the background density of the description (definition) of the candidate classification according to the value of the classification score. For example, the higher the classification score value, the darker the background is displayed, and the lower the classification score value, the lighter the background is displayed. The display mode may be changed by changing the background density, background color, character color, font, etc. Note that the display mode may be changed not only for the classification description (definition) but also for the classification code, or for both the classification code and description.
[0048] When the user inputs a screen switching command to the feature word screen 6, the feature word display processing unit 1014 retrieves the feature words corresponding to the estimated candidate classification from the feature word storage unit 1022 and displays the feature word screen 6 on the classification assignment terminal 12. The feature word screen 6 displays classification information such as the definition sentence and classification score of the candidate classification, feature words related to the candidate classification, and the location in the document that is the basis for assigning the candidate classification.
[0049] When the user inputs a command to switch to the text screen 7, the text display processing unit 1015 acquires document text data and drawing image data relating to the document to be classified from the document data storage unit 1021, and displays the text screen 7. At least the document text and drawings are displayed on the text screen 7, and the document text displays candidate classifications, characteristic words, and descriptions that serve as the basis for assigning the candidate classifications in an identifiable manner.
[0050] The classification unit 1016 is activated when the user selects a classification to be assigned to a document from the classification list and inputs an instruction to confirm the classification. The classification unit 1016 acquires the confirmed classification, associates it with the document number, and records it in the document data storage unit 1021. After one classification has been confirmed, if an instruction to confirm another classification is input, the classification unit 1016 performs the same operation.
[0051] (3-2) Functional configuration of classification learning device 11 Fig. 3 is a diagram showing an example of machine learning by the classification learning device 11 using deep learning according to the first embodiment. Fig. 3 shows the learning function of the learning model 111 using teacher data and the classification estimation function using the machine-learned learning model 111. The learning model 111 in Fig. 3 includes an input layer to which input document text is input, an intermediate layer (hidden layer) consisting of multiple layers, and an output layer that outputs candidate classifications and classification scores estimated from the document text.
[0052] The learning function of classification learning device 11 uses document text data of documents that have already been classified and classification data as training data to train learning model 111 so that when document text data is input, the already assigned classification data is output. Learning model 111 optimizes the weighting coefficients between neurons in each layer of the neural network so that candidate classifications corresponding to the input text data of various documents are output.
[0053] Classification learning device 11 has a classification estimation function that operates as classification estimation unit 1012. When any document text is input to learning model 111, learning model 111 outputs one or more candidate classifications and a classification score corresponding to each candidate classification. In FIG. 3, FI1 and FI2 are classifications defined in the FI classification table, and 0.5 and 0.8 are the classification scores of candidate classifications FI1 and FI2, respectively. In this example, the probability that FI1 will be assigned to the input document text is estimated to be 0.5, and the probability that FI2 will be assigned is 0.8. Classification learning device 11 transmits the estimated candidate classifications and classification scores to classification assignment assistance device 10 as classification estimation results.
[0054] (3-3) Functional configuration of the classification terminal 12 Fig. 4 is a block diagram showing an example of the functional configuration of the classification terminal 12 according to the first embodiment. The classification terminal 12 includes a control unit 1210 and a communication unit 1220. Fig. 4 illustrates a classification screen 5, a feature word screen 6, and a text screen 7 as examples of user interfaces displayed on the display device 125. The candidate classification area 59 will be described in detail in the second embodiment.
[0055] The control unit 1210 functions as an input control unit 1211 , a display control unit 1212 , a category display mode change control unit 1213 , and a classification assignment control unit 1214 by the processor 121 executing a computer program stored in the memory 122 .
[0056] The input control unit 1211 acquires operation inputs for the classification screen 5 from the user via the input device 124. Operation inputs for the classification screen 5 include a document selection command input using the document selection button 57 displayed on the classification screen 5, a screen switching command input using the characteristic word screen button 54 or text screen button 55, a classification confirmation command input using the classification confirmation button 58, or a work termination command input using the end button 56. The display control unit 1212 receives data for displaying the classification screen 5, characteristic word screen 6, or text screen 7 from the classification support device 10, and executes display control for displaying these screens on the display device 125. When a classification display mode change request is received from the classification support device 10, the classification display mode change control unit 1213 executes display control for changing the display mode of candidate classifications in the category list displayed in the classification area 51 of the classification screen 5. In response to the user selecting a category from the category list displayed in the classification area 51 and pressing the classification confirmation button 58, the classification control unit 1214 notifies the classification support device 10 that the selected category has been confirmed.
[0057] (4) Screen layout (4-1) Classification screen 5 FIG. 5 illustrates an example of a classification screen 5 displayed on the display device 125 of the classification terminal 12. The classification screen 5 includes at least a document list area 50 displaying a list of documents to be classified and a classification area 51 displaying the classifications included in the classification table in a classification list in order. The upper portion of the classification screen 5 displays a document number 52, as well as a "Classification Screen" button 53, a "Characteristic Word Screen" button 54, a "Text Screen" button 55, and an "Exit" button 56 for selecting the classification screen 5, the characteristic word screen 6, or the text screen 7. The lower portion of the classification screen 5 displays a "Select Document" button 57 for selecting a document to be classified from the document list area 50 and a "Confirm Classification" button 58 for confirming the classification selected on the classification list as the classification. The candidate classification area 59 will be described in detail in the second embodiment.
[0058] A list of document numbers of documents to be classified is displayed in the document list area 50. The document number list may be created by the user inputting document numbers, or may be stored in advance in the document data storage unit 1021 as a list of documents to be classified. Although a list of patent application numbers is displayed in Fig. 5, other numbers may be used as long as they can identify the documents.
[0059] The classification area 51 is an area for displaying a classification table relating to the classification to be assigned to a document and for determining the classification to be assigned to a document, and all or part of the classification table is displayed in the form of a classification list. The classification list displayed in the classification area 51 displays classification symbols and their descriptions in the order in which they appear in the classification table. As the initial screen of the classification area 51, a classification list may be displayed that includes candidate classifications with high classification scores. If the entire classification table cannot be displayed in the classification area 51, it may be displayed in a scrollable manner so that the entire classification table can be viewed. The classification list may also be categorized and displayed by theme code or by FI classification main group, etc.
[0060] The classification area 51 in Figure 5 displays a classification list starting with "B65G61 / 00,510" in the order of appearance in the FI classification table. Of the classification list, "B65G61 / 00,510," "B65G61 / 00,544," and "B65G61 / 00,520" are candidate classifications. As shown in Figure 5, the classification area 51 is equipped with a scroll bar, allowing the user to scroll through the entire FI classification table.
[0061] Of the classification list displayed in classification area 51, candidate classifications that are likely to be assigned to the document to be classified are displayed in a display mode different from the other classifications. That is, for a classification that falls under a candidate classification in the classification list displayed in classification area 51, the display mode of the candidate classification is changed according to the classification score so that it can be determined that the candidate classification is a candidate classification estimated by learning model 111 and what the classification score, which indicates the degree of likelihood of assignment of the candidate classification, is.
[0062] In Figure 5, for classifications that fall under the candidate classifications, the background density of the definition text in the explanation field is changed according to the classification score value. For example, the background of the explanation field for "B65G61 / 00,510," which has the highest classification score among the candidate classifications, is displayed with a dark background, while the explanation fields for "B65G61 / 00,544," which has the second highest classification score, and "B65G61 / 00,520," which has the third highest classification score, are displayed with gradually lighter backgrounds (in Figure 5, the background of the explanation text is white so that the classification explanation text can be read).
[0063] In this way, in the classification area 51, the display mode of the classification corresponding to the candidate classification is changed, and color information is used to indicate which classifications included in the document list are candidate classifications and the probability that they will be assigned, making it easy to visually identify the classification of interest while overlooking the entire classification list. Note that any method for changing the display mode can be used as long as it allows for distinguishing between the candidate classifications and the classification scores. For example, in addition to changing the shade of the background in the classification explanation field, it is also possible to change the background color or pattern.
[0064] The user selects, by clicking or otherwise selecting, the classification to be assigned to the document from the classification list displayed in classification area 51. Then, to confirm the selected classification, the user presses the "Confirm Classification" button 58. The classifications that can be confirmed are not limited to those included in the candidate classifications, but any classification defined in the classification table can be selected from the classification list and confirmed.
[0065] When a classification is confirmed in the classification area 51, the display mode of the confirmed classification may be changed. For example, the classification symbol of the confirmed classification may be highlighted among the classifications displayed in the classification area 51. In FIG. 5, when "B65G61 / 00,510" displayed in the classification area 51 is selected and the "Confirm Classification" button 58 is pressed, the background color of "B65G61 / 00,510" can be changed to highlight it differently from the other classifications. Note that the highlighting method is not limited to changing the background color, and any display mode may be adopted as long as it makes it clear that the classification has been confirmed. In this way, by displaying the classifications already assigned in the classification area 51 so that they can be easily distinguished, the efficiency of classification can be improved.
[0066] As shown in FIG. 5 , the classification screen 5 has a “Classification Screen” button 53, a “Feature Word Screen” button 54, and a “Text Screen” button 55. When any of these buttons is pressed, the classification support device 10 generates the corresponding screen and displays it on the classification terminal 12. If the user wants to obtain detailed information about the candidate classifications and classification scores displayed in the classification area 51, the user presses the “Feature Word Screen” button 54. In response to the pressing of the “Feature Word Screen” button 54, the classification support device 10 displays the feature word screen 6 on the classification terminal 12. If the user wants to confirm which content of the document the candidate classification is associated with, the user presses the “Text Screen” button 55. In response to the pressing of the “Text Screen” button 55, the classification support device 10 displays the text screen 7 on the classification terminal 12. Note that when the “Classification Screen” button 53 is pressed, the classification support device 10 displays the classification screen 5 on the classification terminal 12. These screens may be displayed in separate windows, or may be displayed simultaneously in one window.
[0067] (4-2) Characteristic Word Screen 6 Figure 6 is a diagram showing an example of the feature word screen 6. Five candidate classifications are displayed on the feature word screen 6 in Figure 6. In Figure 6, information about each candidate classification is displayed in two lines. The first line displays the classification symbol and its explanation, and the second line displays the feature words associated with each candidate classification along with a background pattern. In addition, the claim number and paragraph number that form the basis for the candidate classification are displayed along with a background color. The classification score is displayed across the first and second lines.
[0068] Feature words are terms that serve as clues for classification, but multiple feature words can form a group. Feature word groups can be formed by combinations of terms that can be ORed in a search query, combinations of synonymous terms, or combinations of co-occurring terms. For example, the term combinations "software," "program," and "application" can be grouped together because they are terms that can be ORed in a search query or are synonymous. Furthermore, terms that frequently appear together in documents, i.e., combinations of multiple terms that exhibit a co-occurrence relationship, can be grouped together. For example, if the three terms "location," "GPS," and "measurement" frequently appear together in patent documents, "location," "GPS," and "measurement" can be considered to have a co-occurrence relationship and defined as being in the same group.
[0069] Information on the feature words and groups of feature words corresponding to each classification is created in advance by analyzing documents that have already been made public, and is stored in the feature word storage unit 1022. The feature word display processing unit 1014 obtains information on the feature words and groups associated with each candidate classification from the feature word storage unit 1022, and generates the feature word screen 6.
[0070] The candidate classifications, characteristic words, and the basis points of the candidate classifications displayed on the characteristic word screen 6 are displayed according to the following rules. These rules also apply when displaying the candidate classifications, characteristic words, and the basis points of the candidate classifications on the document text on the text screen 7. Each candidate classification is assigned a different color. The color assigned to the candidate classification is used for the background patterns of the characteristic words related to the candidate classification, and for the background colors of the claim numbers and paragraph numbers that form the basis for assigning the candidate classification. A different background pattern is applied to each feature word. However, the same background pattern is applied to multiple feature words that make up a group. The claim numbers and paragraph numbers that form the basis for assigning the candidate classification are displayed in the background color assigned to the candidate classification, and the display mode of the background color changes depending on the degree to which the characteristic words are included.
[0071] In the feature word screen 6 in Figure 6, a different color is assigned to each of the five candidate classifications (in Figure 6, the patterns in the first column and explanations in speech bubbles are used to represent the colors). The colors assigned to the candidate classifications are also applied to the background patterns of the feature words associated with each candidate classification and the background of the claim numbers and paragraph numbers that serve as the basis for assignment. For example, the first candidate classification, "B65G61 / 00,510," is assigned blue, and the background patterns of the feature words of that candidate classification, such as cargo, surveillance, tracking, vehicle, and location, are displayed in blue. In addition, the background colors of the [Claim 1] and
[0008] , which serve as the basis for assignment, are displayed in shades of blue. The second candidate classification, "B65G61 / 00,544," is assigned orange, and the background patterns of the feature words of that candidate classification and the background of the basis for assignment are highlighted in orange.
[0072] A different background pattern, such as vertical stripes, horizontal stripes, or polka dots, is applied to each feature word (in Figure 6, each pattern is numbered (a) to (g) so that it is easy to distinguish between different patterns). However, the same pattern is applied to feature words that form a single group. For example, "cargo (a)," "load (a)," and "luggage (a)" are groups that can be ORed, and "monitoring (b)," "tracking (b)," and "location (b)" are groups that have a co-occurrence relationship, so the same pattern is applied to the feature words that form each group.
[0073] In the characteristic word screen 6 in Figure 6, after the characteristic words, the claim numbers and paragraph numbers that contain many of the characteristic words of the candidate classification are displayed as the basis for the candidate classification. The color assigned to the candidate classification is applied to the background of these claim numbers and paragraph numbers, and they are displayed in shades according to the degree to which the characteristic words are included. The characteristic words of the first candidate classification "B65G61 / 00,510" appear frequently in claim 1, so the background of [Claim 1] is displayed in dark blue, and because they appear less frequently in paragraph
[0008] , the background of
[0008] is displayed in light blue.
[0074] (4-3) Text Screen 7 Fig. 7 is a diagram showing an example of the text screen 7. The text screen 7 in Fig. 7 is an example of displaying the document text and drawings of a patent document, and includes a heading area 71 for selecting a heading of the patent document, a document text area 72 for displaying the text of the claims and specification, and a drawing area 73 for displaying the drawings.
[0075] The heading area 71 is an area for selecting the items to be displayed in the document text area 72, and displays a list of the document's headings. For example, when [Claims] is selected in the heading area 71, the document text of claims 1 to 5 of the claims is displayed in a scrollable manner in the document text area 72. The drawings included in the document are displayed in a scrollable manner in the drawing area 73. Note that the heading area 71 may also be used to select a drawing number.
[0076] The display mode of the document text displayed in the document text area 72 is changed so that it is possible to determine which terms in the document text are related to the candidate classifications or characteristic words, or which claims or paragraphs are related to the candidate classifications. Specifically, the same display mode as that adopted in the characteristic word screen 6 of Figure 6 is applied.
[0077] In the text screen 7 in Figure 7, among the terms listed in [Claim 1], "baggage" is applied with a blue vertical stripe, "vehicle" with a blue horizontal stripe, and "location" and "tracking" with a blue polka dot background pattern. The blue background pattern indicates that they are characteristic words of the candidate classification "B65G61 / 00,510." It also indicates that "vehicle" (horizontal stripe) in claim 1 and "vehicle" (horizontal stripe) in claim 2 belong to the same group.
[0078] In the text screen 7 of Figure 7, [Claim 1] is displayed against a dark blue background, indicating that claim 1 corresponds to the candidate classification "B65G61 / 00,510" and has a high degree of inclusion of characteristic words of that candidate classification. Also, [Claim 2] is displayed against a dark orange background, indicating that claim 2 corresponds to the candidate classification "B65G61 / 00,544" and has a high degree of inclusion of characteristic words of that candidate classification. Also, [Claim 3] is displayed against a light green background, indicating that claim 3 corresponds to the candidate classification "B65G61 / 00,520" and has a low degree of inclusion of characteristic words of that candidate classification.
[0079] In this way, after displaying candidate classifications on the classification assignment screen 5, the user can switch between displaying the characteristic word screen 6 and the text screen 7, and confirm the classification to be assigned to the document while referring to the characteristic words and assignment reasons related to the candidate classification, or the document text and drawings.This not only improves the accuracy of classification assignment, but also enables the process from classification estimation to classification assignment to be carried out seamlessly and efficiently.
[0080] (5) Operation (5-1) Operation of the classification support device 10 8 is a flowchart illustrating an example of the operation of the classification support device 10 in the first embodiment. A computer program stored in the memory 102 causes the processor 101 of the classification support device 10 to execute each step illustrated in FIG.
[0081] In response to the user starting up the categorization terminal 12, the categorization support device 10 executes processing for displaying the categorization screen 5 on the display device 125 of the categorization terminal 12. The categorization support device 10 generates data for displaying the categorization screen 5 including the document list area 50, the categorization area 51, and predetermined buttons, and displays the categorization screen 5 on the display device 125 of the categorization terminal 12 (step S81).
[0082] The classification support device 10 acquires a document number list of documents awaiting classification from the document data storage unit 1021, and executes a process of displaying the document number list in the document list area 50 of the classification screen 5 (step S82). In the case of patent documents, a patent application number list is displayed in the document list area 50.
[0083] When the classification assistance device 10 detects that the user has selected a document number from the document list area 50 and pressed the "Select Document" button 57, it retrieves document data, including document text data and drawing image data, of the specified document from the document data storage unit 1021. The classification assistance device 10 estimates candidate classifications and classification scores for the specified document (step S83). For example, the classification assistance device 10 transmits the document text of the selected document to the classification learning device 11, and receives from the classification learning device 11 candidate classifications that may be assigned to the document and classification scores that indicate the degree of likelihood.
[0084] The classification support device 10 generates a classification list from the classification table and the estimated candidate classifications and classification scores, and displays it in the classification area 51 of the classification screen 5 (step S84). At this time, the classification support device 10 can display the classification list from the beginning of the classification table, or it can display a classification list including the candidate classification with the highest classification score among the candidate classifications received from the classification learning device 11.
[0085] The classification support device 10 changes the display mode of a classification that matches a candidate classification in the classification list displayed in the classification area 51 based on the classification score value (step S85). The classification support device 10 changes the shading of the background of the explanation field for a classification that corresponds to a candidate classification based on the classification score value.
[0086] When the classification support device 10 detects that the user has pressed the "characteristic word screen" button 54, it reads out the characteristic words and group information related to the candidate classification from the characteristic word storage unit 1022, generates the characteristic word screen 6, and displays it on the classification terminal 12 (step S86). Also, when the classification support device 10 detects that the user has pressed the "text screen" button 55, it obtains the document text data and drawing data of the relevant document from the document data storage unit 1021, generates the text screen 7, changes the display mode of the candidate classification, characteristic words, and classification basis in the document text, and displays it on the classification terminal 12 (step S86).
[0087] When the categorization support device 10 detects that the user has selected one of the categories displayed in the categorization area 51 and pressed the "Confirm Category" button 58, it executes a categorization process (step S87). The categorization support device 10 stores the confirmed category in the document data storage unit 1021 in association with the document number. (5-2) Classification Learning Device 11 9 is a flowchart illustrating an example of the operation of classification learning device 11 of the first embodiment. Classification learning device 11 includes a learning flow using a computer program for training learning model 111 shown in FIG. 9(a), and a classification estimation flow using a computer program for performing classification estimation using trained learning model 111 shown in FIG. 9(b).
[0088] In the learning flow of FIG. 9(a), the classification learning device 11 acquires training data (step S101). The classification learning device 11 reads documents to which classifications have already been assigned from the document data storage unit 1021, and acquires the document text and the patent classification assigned to the document text as training data. The classification learning device 11 uses the acquired training data to train the learning model 111 (step S102). When document text that serves as training data is input, the classification learning device 11 trains the learning model 111 so that the classification assigned to the document text is output. Note that the learning model 111 may be trained using deep learning.
[0089] In the classification estimation flow of Figure 9(b), when the classification learning device 11 receives a document text to be estimated together with a classification estimation request from the classification assignment assistance device 10, it inputs the received document text to the trained learning model 111 (step S103). The learning model 111 outputs candidate classifications that may be assigned to the input document text and classification scores for each candidate classification (step S104). The classification learning device 11 transmits the candidate classifications and classification scores, which are outputs of the learning model 111, to the classification assignment assistance device 10 as classification estimation results in response to the classification estimation request.
[0090] According to the first embodiment, by changing the display mode of the candidate classifications displayed on the classification assignment screen 5, it becomes possible to easily distinguish the classification to be noted and the possibility of assigning that classification while getting an overview of the classification table as a whole, thereby improving the accuracy and efficiency of classification assignment.
[0091] Second Embodiment A classification support system 1 according to a second embodiment will now be described. In the second embodiment, as shown in FIG. 5, a candidate classification area 59 is provided for displaying candidate classifications estimated by a learning model 111. The candidate classification area 59 displays a candidate classification symbol, a description, and estimated information (classification score) for each candidate classification. To generate display data for the candidate classification area 59, the classification support device 10 includes a first classification display processing unit that displays only candidate classifications whose classification scores are higher than a predetermined threshold, or a second classification display processing unit that displays candidate classifications in descending order of classification score.
[0092] The first classification display processor extracts only candidate classifications whose classification scores are higher than a predetermined threshold from among the pairs of candidate classifications and classification scores received from classification learning device 11, and displays them in candidate classification area 59. In candidate classification area 59 of FIG. 5, five candidate classifications output by learning model 111 are displayed in descending order of classification score. The user can activate the first classification display processor when, for example, the number of candidate classifications output by classification learning device 11 is large, and display a predetermined number of candidate classifications with high classification scores. The number of candidate classifications to be displayed in candidate classification area 59 may be set in advance, or may be specified by the user when activating the first classification display processor.
[0093] The second classification display processor displays the pairs of candidate classifications and classification scores received from classification learning device 11 in order of classification score in candidate classification area 59. When analyzing the candidate classifications output by classification learning device 11 in order of classification score, the user can activate the second classification display processor and display the output candidate classifications in order of likelihood of assignment.
[0094] In this way, the user can select the first classification display processing unit or the second classification display processing unit according to the candidate classification and classification score output by the learning model 111. This allows the candidate classification to be displayed in a display format suitable for analysis, enabling efficient classification. Note that display by the first classification display processing unit and display by the second classification display processing unit may be performed simultaneously in the candidate classification area 59.
[0095] <Third embodiment> In the third embodiment, in the preparation stage of training data used to generate a learning model, preprocessing is performed to associate the text of a specific document to be input to the training model with a classification assigned to the specific document. As preprocessing of the training data, a process is performed to associate the classification assigned to the specific document with text in the specific document that is similar to a sentence defining the classification. If a specific document is assigned multiple classifications, similar preprocessing is performed on the text of the specific document and the multiple classifications assigned to the specific document.
[0096] When training a learning model using the text of a patent document and the classification assigned to the patent document as training data, the classification assigned to the patent document may relate to only a portion of the text of the patent document, and therefore, if the text of the patent document that is not related to the classification is selected as training data, the estimation accuracy of the learning model 111 may decrease. For example, in a case where the number of words in a document that can be input to the learning model 111 as training data is limited to 512 words, if the input 512 words are not related to the classification, the learning model 111 may not be trained correctly.
[0097] In addition, in cases where multiple patent classifications are assigned to a single patent document, if a first description in the patent document corresponds to the first classification and a second description different from the first description corresponds to the second classification, using the first description as training data for the second classification may result in a decrease in the estimation accuracy of the learning model 111.
[0098] FIG. 10(a) is a diagram illustrating such a case. FIG. 10(a) shows text data of the claims in a patent document, where claim 1 is a description related to the first classification (a) and claim 2 is a description related to the second classification (b). Furthermore, assume that the number of words in the training data that can be input to the learning model 111 is limited to 512 words. FIG. 10(a) illustrates a case in which only the middle of claims 1 and 2 (the portions indicated by black circles) can be input to the learning model 111. In such a case, learning is performed so that when the contents of claim 1 are input, the first classification (a) and the second classification (b) are output, resulting in a decrease in the estimation accuracy of the second classification (b).
[0099] Figure 10(b) shows a general flow of preprocessing of training data. Figure 10(b) shows a document (original) as a specific document used in training data. The document (original) includes sentence 1 (claim 1), sentence 2 (claim 2), and sentence 3 (claim 3), and is assigned a first classification (a) and a second classification (b).
[0100] In the preprocessing of the training data, the sentences contained in the document (original) are rearranged so that the combination of sentence 2 and the second classification (b) becomes the training data. In the example of Figure 10(b), sentences 1 to 3 contained in the document (original) are rearranged to place sentence 2 (shown in bold) at the beginning of the document, since sentence 2 is highly related to the definition sentence of the second classification (b). As a result, a document (after rearrangement) is generated in which sentences 2, 1, and 3 are rearranged in this order for the second classification (b). Here, sentences highly related to the second classification (b) can be selected by determining the similarity between the contents of each paragraph in the document and the definition sentences of the second classification (b). To determine the similarity, the distance between the original patent document and the definition sentences of the second classification (b) may be calculated, or a learning model that outputs the similarity between the two may be used.
[0101] After such preprocessing, the learning model 111 is trained using the document (after rearrangement) and the second classification (b) as training data. In this case, the learning model 111 is trained so that it outputs the second classification (b) when the first 512 words of the document (after rearrangement) are input. Since the first 512 words of the document include sentence 2, the learning model 111 can be trained with high accuracy. Note that although the semantic content of the sentence changes when sentence 1 and sentence 2 are swapped, this does not affect the accuracy of the learning model because the learning model 111 does not determine the semantic content of the sentence. In Figure 10(b), the first classification (a) has also been assigned to the document (original), so the same preprocessing can be performed on the first classification (a).
[0102] In Fig. 10(b), an example in which two classifications are assigned to the document (original) has been described, but similar preprocessing may also be applied when multiple classifications are assigned. Also, even when only one classification is assigned to the document (original), preprocessing similar to that for the first classification (b) can be performed.
[0103] According to the third embodiment, when one document is assigned one or more classifications in training the learning model 111 used for classification estimation, the accuracy of classification estimation can be improved by preprocessing the training data.
[0104] The present disclosure is not limited to the above-described embodiments, and in the implementation stage, the components can be modified and embodied without departing from the spirit of the present disclosure. Furthermore, various inventions can be formed by appropriately combining multiple components disclosed in the above-described embodiments. For example, some components may be omitted from all the components shown in the embodiments. Furthermore, components from different embodiments may be appropriately combined. [Explanation of symbols]
[0105] 1...Classification support system, 5...Classification screen, 6...Feature word screen, 7...Text screen, 10...Classification support device, 11...Classification learning device, 12...Classification terminal, 50...Document list area, 51...Classification area, 52...Document number display, 53...Classification screen button, 54...Feature word screen button, 55...Text screen button, 56...Exit button, 57...Document selection button, 58...Classification confirmation button, 59...Candidate classification area, 71...Heading item area, 72...Document text area, 73...Drawing area, 101...Processor, 102...Memory, 103...Storage, 104...Communication I / F, 105...Input / output I / F, 111...Learning model , 121...processor, 122...memory, 123...storage, 124...input device, 125...display device, 126...communication I / F, 1010...management unit, 1011...display processing unit, 1012...category estimation unit, 1013...category display mode change unit, 1014...topic word display processing unit, 1015...text display processing unit, 1016...categorization assignment unit, 1020...storage unit, 1021...document data storage unit, 1022...topic word storage unit, 1023...categorization information storage unit, 1030...communication unit, 1210...control unit, 1211...input control unit, 1212...display control unit, 1213...categorization display mode change control unit, 1214...categorization assignment control unit, 1220...communication unit
Claims
1. An information processing device for supporting classification of documents, a classification estimation unit that estimates a candidate classification that is a candidate for a classification to be assigned to the document and a classification score that indicates a degree of likelihood that the candidate classification will be assigned; a display processing unit that displays a classification list including the candidate classifications; a classification display mode change unit that changes the display mode of candidate classifications included in the classification list, Information processing device.
2. The information processing device according to claim 1 , wherein the category display mode change unit changes the display mode of the candidate categories included in the category list in accordance with the category score.
3. The information processing apparatus according to claim 1 , further comprising a first classification display processing unit that displays only candidate classifications whose classification scores are higher than a predetermined threshold.
4. The information processing apparatus according to claim 1 , further comprising a second classification display processing unit that displays candidate classifications in descending order of the classification score.
5. The information processing apparatus according to claim 1 , further comprising a text display processing unit that displays the text of the document, wherein the text display processing unit highlights characteristic words related to the candidate classification in the text of the displayed document.
6. The information processing device according to claim 1 , wherein the classification estimation unit estimates the candidate classifications and the classification scores using deep learning.
7. the classification estimation unit, when given a text of a document, performs estimation using a learning model that outputs a plurality of candidate classifications that are candidates for classification to be assigned to the text of the document, and a classification score that indicates a degree of likelihood that the candidate classification will be assigned; The information processing apparatus according to claim 1 , wherein the learning model is trained using a classification assigned to a particular document and text from the particular document that is similar to a sentence defining the classification.
8. The information processing apparatus according to claim 7 , wherein the learning model is trained using the text of the particular document and a plurality of classifications assigned to the particular document.
9. An information processing method for assisting in classification of documents, comprising: estimating a candidate classification that is a candidate for a classification to be assigned to the document and a classification score that indicates a degree of likelihood that the candidate classification will be assigned; displaying a classification list including the candidate classifications; changing the display mode of candidate categories included in the category list; Information processing methods.
10. A computer program that causes a computer to function as each unit of the information processing device according to any one of claims 1 to 8.
Citation Information
Patent Citations
Classification support apparatus and method, and program
JP2008171164A
Algorithm determining device, learning equipment production device, patent classifier, classification information determining device, algorithm determination method, learning equipment production method, classification information determination method, and program
JP2021036427A
Document processing device and classification assignment support system
JP2022103710A