A system for estimating important words, a method for constructing the system, and a method for classifying words.

The document display support system uses a machine-learning based word selection model to classify and visually highlight important words, addressing the inefficiencies of manual highlighting and keyword-based systems by enhancing relevance accuracy.

JP7854631B2Active Publication Date: 2026-05-07XCOO INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
XCOO INC
Filing Date
2021-04-30
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing document display technologies require manual highlighting of important words, which is time-consuming and subjective, and do not accurately indicate relevance to a specific field, while keyword-based systems need manual selection of words for importance, also lacking accuracy.

Method used

A document display support system using a trained word selection model through machine learning to classify words as target, non-target, or indeterminate based on context, visually distinguishing important words without manual input.

Benefits of technology

Accurately estimates and visually distinguishes important words in documents relevant to a specific field, reducing manual effort and improving relevance determination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007854631000008
    Figure 0007854631000008
  • Figure 0007854631000009
    Figure 0007854631000009
  • Figure 0007854631000010
    Figure 0007854631000010
Patent Text Reader

Abstract

To provide a document display assistance system, a method, and a program, which estimates important words in a specific field and visually distinguishes and displays the estimated words.SOLUTION: A document display assistance system comprises: a DB in which words to be selected and words not to be selected are registered; a trained word selection model that has undergone machine learning for estimating whether a word is a word to be selected; a text pre-processing section that extracts each word from a received document to be displayed; a word classifying section that classifies each word as a word to be selected, a word not to be selected, or an uncertain word on the basis of the DB; a text post-processing section that generates output data by assigning a predetermined attribute to a predetermined word in the document to be displayed; and an output section that outputs the output data. If a label indicating that the uncertain word classified by the word classifying section is a word to be selected is estimated, the word selection model classifies the uncertain word as a word to be selected. The text post-processing section assigns the predetermined attribute to the classified word to be selected.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to document display support technology, and more particularly, to a document display support system, a document display support method, and a program for executing the method for visually distinguishing and displaying specific words in a document in a specific field.

Background Art

[0002] Currently, a huge number of documents are accumulated in databases around the world. Usually, in order for a user to determine whether a certain document is relevant to a specific field and whether it is important, the user needs to read the document and understand its content. For example, in the field of cancer genomics medicine, medical professionals such as doctors and researchers annotate information such as gene mutations, drugs, and clinical trial information while referring to a knowledge database or an integrated database that stores information on gene mutations, drugs, and clinical trials in order to create, for example, a report. Among these, information related to gene mutations and clinical significance is mostly published in the form of papers, and medical professionals are required to read and correctly understand the difficult text of the papers.

[0003] However, the number of papers that need to be referred to is huge, the content of each paper is difficult to understand, and the amount is also very large. Therefore, in reality, it is difficult to quickly determine whether the paper of interest is useful in cancer genomics medicine and which part of the paper contains important information. For this reason, several techniques for devising the display of the document that the user intends to refer to have been proposed.

[0004] For example, Patent Document 1 below discloses a technology for checking only the sentences to be checked in long patent information data. Specifically, Patent Document 1 discloses an information processing device in which the processing means comprises a content discrimination character search unit, a highlighting display processing unit, a sentence delimiter discrimination unit, a display determination unit, and a hide setting unit, and the display means displays sentences separated by paragraph marks that the display determination unit has determined to contain the highlighted characters, for each case data, and also displays the highlighted characters that have been colored by the highlighting display processing unit.

[0005] Furthermore, Patent Document 2, described below, discloses a technology for automatically searching for relevant case reports from past medical reports, including the sections for examination purpose, findings, and diagnosis, using keywords that are weighted according to the frequency and importance of the cases. [Prior art documents] [Patent Documents]

[0006] [Patent Document 1] Japanese Patent Publication No. 2016-207071 [Patent Document 2] Japanese Patent Publication No. 2012-141797 [Overview of the Initiative] [Problems that the invention aims to solve]

[0007] However, the technology disclosed in Patent Document 1 has the problem that the user has to think about each character they want to highlight and input it through the input means, which is time-consuming and cumbersome. Furthermore, the highlighted characters are subjective to the user and do not necessarily accurately indicate that such highlighted characters are important or highly relevant in a particular field.

[0008] Furthermore, in the technology disclosed in Patent Document 2 mentioned above, in order to use words or word sequences weighted by frequency of occurrence and importance as search keywords, it was necessary for doctors or others to select words or word sequences they considered important from the displayed report and add them to the importance keyword list in advance. Therefore, similar to the above, it did not necessarily accurately indicate whether the word or word sequence was important or highly relevant in a particular field.

[0009] Therefore, the present invention aims to provide a document display support technology that can estimate words that are considered important or highly relevant in a document of a specific field that a user intends to refer to, and visually distinguish and display the estimated words. [Means for solving the problem]

[0010] The present invention, which solves the above problems, is comprised of the following inventive features or technical characteristics.

[0011] The present invention, in accordance with a certain perspective, is a document display support system that may comprise: a database in which a set of selectable words consisting of words to be selected and a set of non-selectable words consisting of words not to be selected in documents of a particular field are registered; a trained word selection model that has been subjected to machine learning to estimate a label indicating at least one of a selectable word and a non-selectable word for a given word, according to the context related to the word; a document reception unit that receives a document to be displayed; a text pre-processing unit that extracts multiple words from the received document to be displayed; a word classification unit that classifies each of the multiple words in the received document to be displayed into at least one of a selectable word, a non-selectable word, and an indeterminate word based on the database; a text post-processing unit that generates output data by assigning predetermined character attribute information to predetermined words in the document to be displayed; and an output unit that outputs the output data. The trained word selection model can classify the classified indeterminate word as a selectable word if a label indicating a selectable word is estimated for that indeterminate word. The text post-processing unit can also assign the predetermined character attribute information to the classified selectable word.

[0012] The word classification unit may classify each of the plurality of words, based on the database, as an undetermined word if it is not classified as either a selected word or a non-selected word.

[0013] Furthermore, the present invention, from a different perspective, is a document display support system that may comprise: a database in which multiple documents are registered; a trained word selection model subjected to machine learning to estimate a label indicating at least one of a target word and a non-target word for a given word, according to the context related to the word; a document reception unit that receives input of a document to be displayed; a text pre-processing unit that extracts multiple words from the received document to be displayed; a word classification unit that calculates a predetermined score for each of the extracted multiple words based on the database and classifies the extracted words into at least one of a target word, a non-target word, and an indeterminate word according to the calculated predetermined score; a text post-processing unit that generates output data by assigning predetermined character attribute information to predetermined words in the document to be displayed; and an output unit that outputs the output data. The trained word selection model classifies the classified indeterminate word as a target word if a label indicating a target word is estimated for that indeterminate word. The text post-processing unit also assigns the predetermined character attribute information to the classified target word.

[0014] The database may include a corpus of correct example documents defined as correct examples. The word classification unit may classify a word as a target word if a first score calculated using the corpus of correct example documents exceeds a first threshold.

[0015] The database may further include a corpus of negative example documents defined as incorrect cases. The word classification unit may calculate the first score based on the total number of documents in which the word appears in all positive example documents in the positive example document corpus and all negative example documents in the negative example document corpus, relative to the number of documents in which the word appears in all positive example documents in the positive example document corpus.

[0016] The word classification unit may classify a word as a non-selection target word based on a second score corresponding to the frequency of documents in which the word appears, based on the entire positive example document corpus and the entire negative example document corpus. For example, the word classification unit may classify a word as a non-selection target word if the second score corresponding to a predetermined inverse document frequency does not exceed a second threshold, based on the entire positive example document corpus and the entire negative example document corpus. Alternatively, the word classification unit may classify a word as a non-selection target word if the second score corresponding to a predetermined document frequency exceeds a second threshold.

[0017] Furthermore, the document display support system may further include a trained document classification model that has undergone machine learning to calculate the accuracy of a document according to the frequency of word occurrences in the specified field, using a first group of documents in a specific field document corpus related to the specified field and a second group of documents in a general field document corpus covering a broader field than the specified field as training data.

[0018] Furthermore, the document display support system may further include a document extraction unit that randomly extracts a second set of documents from the general field document corpus such that the number of documents in the second set of documents is substantially the same as the number of documents in the first set of documents.

[0019] The trained document classification model can receive a document to be classified and calculate the accuracy of the received document, thereby estimating whether the document is a correct or incorrect document based on a predetermined threshold.

[0020] Furthermore, the text post-processing unit described above can generate output data configured such that the predetermined words to which the predetermined character attribute information is attached are visually distinguishable and displayed.

[0021] Furthermore, the present invention, in a different view, may be a document display support method that displays predetermined words in a document to be displayed in such a way that they are visually distinguishable. The method may include registering a set of words to be selected and a set of words not to be selected in a database in a document of a particular field; constructing a trained word selection model that has been subjected to machine learning to estimate a label indicating at least one of the selected words and non-selected words for a given word, according to the context related to the word; extracting multiple words from a received document to be displayed; classifying each of the multiple words in the received document to be displayed into at least one of the selected words, non-selected words and indeterminate words based on the database; generating output data by assigning predetermined character attribute information to predetermined words in the document to be displayed; and outputting the output data. The trained word selection model may classify the classified indeterminate words as selected words if a label indicating the selected words is estimated for those indeterminate words. Furthermore, generating output data may include assigning the predetermined character attribute information to the classified selected words.

[0022] Also, according to another aspect, the present invention can be a document display support method for performing display so that a predetermined word in a display target document is visually distinguished. The method includes constructing a learned word selection model that has been subjected to machine learning to estimate a label indicating at least one of a selection target word and a non-selection target word for a given word according to the context related to the word, extracting a plurality of words from the received display target document respectively, calculating a predetermined score for each of the plurality of extracted words based on a database in which a plurality of documents are registered, classifying the extracted word into at least one of a selection target word, a non-selection target word, and an uncertain word according to the calculated predetermined score, generating output data by assigning predetermined character attribute information to a predetermined word in the display target document, and outputting the output data. The learned word selection model can classify an uncertain word as a selection target word when a label indicating the selection target word is estimated for the classified uncertain word. Further, generating the output data may include assigning the predetermined character attribute information to the classified selection target word.

[0023] The method may further include constructing the learned document classification model that has been subjected to machine learning to calculate the accuracy of a document according to the frequency of occurrence of words in the specific field, using a first document group in a specific field document corpus related to the specific field and a second document group in a general field document corpus in a field wider than the specific field as learning data.

[0024] [[ID=!7]] Also, the present invention can be understood as a product including a computer program for realizing the above method by a computing device and a recording medium that non-temporarily records the program.

[0025] It should be noted that there seems to be a typo in your original text where "!7" should probably be "7" in . I translated it as it is but just for your reference.In the present disclosure, the term "means" does not simply refer to a physical means, but also includes cases where the function of the means is realized by software. Further, even if the function of one means is realized by two or more physical means, or the functions of two or more means are realized by one physical means, it is acceptable.

[0026] In the present disclosure, the term "system" includes a logical aggregation of a plurality of devices (or functional modules that realize specific functions), and it does not matter whether each device or functional module is physically configured as a single entity or as separate entities.

Advantages of the Invention

[0027] According to the present invention, for a document in a specific field that a user intends to refer to, using a word selection model, words that are considered important or highly relevant in the specific field are estimated as target selection words, and the estimated target selection words can be visually distinguished and displayed.

[0028] Other technical features, objectives, operational effects, and advantages of the present invention will be clarified by the following embodiments described with reference to the attached drawings.

Brief Description of the Drawings

[0029] [Figure 1] It is a block diagram showing an example of the schematic configuration of a document display support system according to an embodiment of the present invention. [Figure 2] It is a block diagram for explaining the functional configuration of a document display support system according to an embodiment of the present invention. [Figure 3A] It is a diagram for explaining the machine learning of a word selection model in a document display support system according to an embodiment of the present invention. [Figure 3B] It is a diagram for explaining the estimation by a learned word selection model in a document display support system according to an embodiment of the present invention. [Figure 4]This figure shows an example of a document displayed on a browser screen in a document display support system according to one embodiment of the present invention. [Figure 5] This is a block diagram illustrating the functional configuration of a document display support system according to one embodiment of the present invention. [Figure 6] This diagram illustrates an example of the configuration of a document database in a document display support system according to one embodiment of the present invention. [Figure 7] This is a block diagram showing an example of the schematic configuration of a document display support system according to one embodiment of the present invention. [Figure 8] This is a flowchart illustrating the process of building a document database in a document display support system according to one embodiment of the present invention. [Figure 9] This is a flowchart illustrating the learning process of a word selection model in a document display support system according to one embodiment of the present invention. [Figure 10] This is a flowchart illustrating the learning process of a word selection model in a document display support system according to one embodiment of the present invention. [Figure 11] This figure shows an example of the hardware configuration of a document display support system according to one embodiment of the present invention. [Modes for carrying out the invention]

[0030] Embodiments of the present invention will be described below with reference to the drawings. However, the embodiments described below are merely illustrative, and there is no intention to exclude various modifications or applications of techniques not explicitly stated below. The present invention can be implemented by various modifications (for example, by combining each embodiment) without departing from its spirit. In the following drawings, identical or similar parts are denoted by the same or similar reference numerals. The drawings are schematic and do not necessarily correspond to actual dimensions or proportions. There may be parts in the drawings where the dimensional relationships and proportions differ from those of other parts.

[0031] This disclosure describes a document display support system and a document display support method for displaying a document to be displayed, which uses highly relevant documents selected through filtering from a large number of documents in a specific field, and then uses a selection process to select words that characteristically represent the content of the document, so that the selected words are visually distinguishable. As described later, the filtering process uses a document classification model that has been subjected to machine learning using known machine learning algorithms such as neural networks. In addition, the word selection process uses a word selection model that has been subjected to machine learning. In this disclosure, "word" is not limited to terms or technical terms, but broadly refers to something that represents a unit in meaning and has a grammatical function, and is not limited to a single word, but includes combinations of multiple words or phrases.

[0032] [First Embodiment] This embodiment relates to a document display support system and method that uses a word database containing words related to a specific field to classify each word in the text body of a document into one of three categories: a target word, a non-target word, or an undetermined word. Furthermore, it uses a pre-trained word selection model, which has been subjected to machine learning using a predetermined machine learning algorithm, to estimate whether the classified undetermined word is a target word or a non-target word. The system then displays a document in which the classified or estimated target words are visually distinguished.

[0033] Figure 1 is a block diagram showing an example of a schematic configuration of a document display support system according to one embodiment of the present invention. As shown in the figure, the document display support system 100 is configured to include, for example, a cloud computer 110 and a client computer 120. Here, the document display support system 100 is configured as a cloud-type system, but it is not limited to this, and may be configured as a client / server system or a standalone system.

[0034] The cloud computer 110 is a computing device that has the function of analyzing a document specified by, for example, the client computer 120. The hardware configuration of the cloud computer 110 is shown in Figure 11, for example, but since such a configuration is known, the details are omitted. The cloud computer 110 sends its analysis results back to the client computer 120. The cloud computer 110 may be configured to include a word database 210 and a pre-trained word selection model 236 that has undergone machine learning, as described with reference to Figure 2. The cloud computer 110 may be configured by one or more computing devices.

[0035] The client computer 120 is a computing device operated by a user of the document display support system 100. The client computer 120 may, for example, implement a web browser and access the cloud computer 110 via the web browser. Alternatively, instead of a web browser, a dedicated browser or application adapted for the document display support system 100 may be implemented. As an example, the client computer 120 sends a document (for example, the document data itself or the URL of the document data) to the cloud computer 110 for the user to determine whether or not it is relevant to a particular field. In response, the cloud computer 110 sends the results of the document analysis back to the client computer 120, and the client computer 120 displays the analysis results on the user interface.

[0036] The document display support system 100 will be described in more detail below. Figure 2 is a block diagram illustrating the functional configuration of a document display support system according to one embodiment of the present invention. In this figure, the document display support system 100 includes, for example, a word database 210, a document reception unit 220, a document analysis unit 230 equipped with a learned word selection model 236, a text post-processing unit 240, and a document display unit 250. In the cloud-type document display support system 100 described above, for example, the word database 210, the document reception unit 220, the document analysis unit 230, and the text post-processing unit 240 are implemented as functions of a cloud computer 110, and the document display unit 250 is implemented as a function of a client computer 120, but it is not limited to this. For example, all or part of the text post-processing unit 240 may be implemented as a function of the client computer 120. The document reception unit 220, the document analysis unit 230, and the text post-processing unit 240 form a document processing unit 260.

[0037] The word database 210 is a database system in which, for documents in a specific field, sets of words that should be selected (hereinafter referred to as "selected word sets") and sets of words that should not be selected (hereinafter referred to as "non-selected word sets") are registered, classified according to word class classification. The word database 210 is a form of dictionary database. Word class classification refers to, for example, using a corpus of documents defined in advance as correct examples (hereinafter referred to as "correct example documents") and a corpus of documents defined in advance as incorrect examples (hereinafter referred to as "incorrect example documents"), scoring each word using a predetermined statistical model (e.g., on a scale of -100 to 100), and classifying words according to the score assigned based on a predetermined threshold. Correct examples refer to cases where a document is relevant to a specific field, and incorrect examples refer to cases where a document is not relevant to a specific field. In other words, the selected word sets are sets of words that the document display support system 100 definitely wants to select in a given display target document, and the non-selected word sets are sets of words that it definitely does not want to select. Furthermore, in the word classification results, sets of words with scores that are neither low nor high (for example, scores of 0 or around that) are considered indeterminate word sets.

[0038] The document reception unit 220 functions as an interface for receiving documents to be displayed, for which it is desired to visually determine whether or not they are relevant to a specific field. Documents received by the document reception unit 220 are displayed by the document display unit 250, described later, with their relevance to the specific field visually represented. For example, the document reception unit 220 can receive documents obtained from an internet-based literature database. Alternatively, the document reception unit 220 may receive documents obtained from a local computer. As another example, the document reception unit 220 can receive documents specified by the user via a web browser. The document reception unit 220 passes the received documents to the document analysis unit 230 and also to the text post-processing unit 240.

[0039] The document analysis unit 230 analyzes the received document and identifies words that are considered important in a particular field within the document. The identified words are displayed visually (e.g., highlighted). As described above, in this example, the document analysis unit 230 is composed of a text preprocessing unit 232, a word classification unit 234, and a word selection model 236.

[0040] The text preprocessing unit 232 extracts multiple words from the text body of the received document. A morphological analysis unit that performs morphological analysis of sentences obvious to those skilled in the art is one form of the text preprocessing unit 232. The text preprocessing unit 232 passes each of the extracted words to the word classification unit 234.

[0041] The word classification unit 234 identifies and classifies each word in the received document by referring to the word database 210, determining whether it is a target word, a non-target word, or an indeterminate word. Specifically, for words extracted by the text preprocessing unit 232, the word classification unit 234 refers to the word database 210 and, if the word exists in the target word set, assigns or associates a learning label (e.g., "1") indicating it is a target word. If it exists in the non-target word set, assigns or associates a learning label (e.g., "-1") indicating it is a non-target word. Furthermore, if the extracted word does not exist in either the target word set or the non-target word set, the word classification unit 234 classifies it as an indeterminate word. In this example, no learning label is assigned to indeterminate words. The word classification unit 234 passes each word (word sequence) classified as a target word, a non-target word, or an indeterminate word to the word selection model 236 and also to the text postprocessing unit 240. As a variation, the word classification unit 234 may pass words classified as indeterminate words to the word selection model 236, and pass words classified as target words and non-target words to the text post-processing unit 240.

[0042] The word selection model 236 of this embodiment is a trained model that has undergone machine learning to estimate the label of each word extracted from the text of a document (for example, whether it is a word to be selected or not), while considering the context and meaning of the word. Context refers to the logical relationships and continuity between individual words in a sentence and / or between individual sentences or phrases. For example, if a word is used with the same meaning, it is in the same context, and if it is used with a different meaning, it is in a different context. Context depends on the sequence of words (which may be sentences or phrases) before and after the word. In this disclosure, the word selection model 236 before or during training may also be simply referred to as "word selection model 236," but it may also be referred to as "trained word selection model 236" to indicate that it is trained. The word selection model 236 may, but is not limited to, sequential labeling models such as hidden Markov models (HMMs) or long short-term memory (LSTMs). In machine learning, for example, each word extracted by morphological analysis is converted into a word embedding vector. A word embedding vector represents a vocabulary by embedding vectors representing the features of individual words into a vector space that represents the features of the vocabulary. In this disclosure, word embedding vectors are used, but the disclosure is not limited to this, and other feature vectors may be used. As described later, the word selection model 236 is input with indeterminate words classified into an indeterminate word set, and labels for these indeterminate words are estimated.

[0043] For example, consider the case of machine learning a word selection model 236 using the text sentence "...the c-Myc-Skp2-Miz1 complex that activates RhoA and the c-MycpTEFB complex..." from a document in the field of cancer genome medicine as training data. As shown in Figure 3A, the text sentence is extracted into multiple words, and each extracted word is converted into a word embedding vector. Each word converted into a word embedding vector is input to the word selection model 236, and the word selection model 236 performs machine learning so that it obtains the same output as the pre-given training labels for the words corresponding to the words to be selected and the words not to be selected. In the example shown in Figure A, the words "activates" and "RhoA" are indeterminate words, so no training labels are given to them.

[0044] Returning to Figure 2, the word selection model 236, after machine learning as described above, estimates a label for the input word while considering the context (e.g., the surrounding word sequence). That is, the word selection model 236 assigns either a target word or a non-target word label to the input word and outputs the word-label pair as the estimation result. As an example, the words input to the word selection model 236 include not only indeterminate words but also words other than indeterminate words (e.g., target words and / or non-target words) as contextual information. This allows estimation to be performed using the generalization ability of the trained model. As another example, the words to be input may be, for example, words that have been classified as indeterminate words by the word classification unit 234.

[0045] In other words, suppose that the word classification unit 234 classifies the words "activates" and "RhoA" as indeterminate words in a document containing the above-mentioned text sentence "...the c-Myc-Skp2-Miz1 complex that activates RhoA and the c-MycpTEFB complex ...". In this case, as shown in Figure 3B, the trained word selection model 236 estimates labels for these indeterminate words "activates" and "RhoA" according to the context of the words and their surroundings. Figure 3B shows that, based on the context, the word "activates" is estimated to have the label "-1" as a word not to be selected, and the word "RhoA" is estimated to have the label "+1" as a word to be selected, based on the context.

[0046] The text post-processing unit 240 modifies the text of the received document so that the selected words are displayed visually distinctly in the document, based on the estimation results output by the word selection model 236. Specifically, for the document handed over from the document receiving unit 220, the text post-processing unit 240 generates output data configured so that the selected words are displayed in a visually distinguishable manner in the text of the document, based on the selected and non-selected words handed over from the word classification unit 234 and the selected words estimated and handed over by the word selection model 236. A visually distinguishable manner typically includes, but is not limited to, highlighting, such as underlining, coloring, italics, bolding, enlargement, bordering, and / or markers. The output data, for example, stores character attribute information indicating such manners, associated with the selected words in the text of the document. The text post-processing unit 240 hands over the generated output data to the document display unit 250.

[0047] The document display unit 250 displays the modified document based on the output data received from the text post-processing unit 240. As described above, the document display unit 250 is implemented, for example, as part of the functions of the client computer 120. For example, the document display unit 250 interprets the output data received from the text post-processing unit 240 and, according to the character attribute information therein, displays the selected words in the text of the document on the user interface in a visually distinguishable manner, as shown in Figure 4. That is, the figure shows an example of a browser screen 400, and more specifically, it shows how words identified as selected words are highlighted in the document displayed on the browser screen (represented as enclosed characters in the figure). In this example, the document score and related keywords are also shown on the right side of the browser screen 400.

[0048] As described above, according to this embodiment, the document display support system 100 displays words related to a specific field in a visually distinguishable manner for documents whose relevance to a specific field is to be determined. As a result, users can easily and quickly follow the words displayed in a visually distinguishable manner, and immediately determine whether or not the document is related to a specific field.

[0049] In particular, according to this embodiment, the document display support system 100 can more accurately estimate whether a word that was previously undeterminable belongs to a set of selected words related to a specific field by inputting words classified as uncertain words by the word classification unit 234 into a word selection model 236 that has been subjected to machine learning to calculate the contextual similarity of a given word in a document of a specific field using a registered set of selected words and a set of non-selected words.

[0050] Furthermore, according to this embodiment, there is no need to manually set or register target words related to a specific field in advance, freeing the user from that hassle.

[0051] [Second Embodiment] This embodiment is a modification of the first embodiment described above and is directed toward a document display support system and a document display support method that use a document database comprising a positive example document corpus and a negative example document corpus.

[0052] Figure 5 is a block diagram illustrating the functional configuration of a document display support system according to one embodiment of the present invention. In the figure, components identical to those in the above embodiment are denoted by the same reference numerals, and therefore, their descriptions will be omitted as appropriate below. That is, as shown in the figure, the document display support system 100 of this embodiment differs from the first embodiment in that it employs a document database 510 instead of the word database 210 in the first embodiment. In other words, the document database 510 is one form of the database 210 shown in Figure 1. Accordingly, the word classification unit 234 is configured to operate in accordance with the document database 510.

[0053] The document database 510 is composed of, for example, a positive example document corpus 512 and a negative example document corpus 514. As described above, the positive example document corpus 512 is a database in which documents defined in advance as correct examples (positive example documents) are registered. The negative example document corpus 514 is a database in which documents defined in advance as incorrect examples are registered.

[0054] Such a document database 510 is constructed by inputting a large number of documents into a pre-trained document classification model that has been machine-learned using a document corpus of a specific field with a high frequency of selected words, although it is a small domain, and a document corpus of a general field with a broader domain than the document corpus of the specific field, as described in other embodiments, and classifying each of the input documents as either a positive example document or a negative example document. The document database 510 can store documents collected and classified from, for example, document corpora related to journals and academic papers.

[0055] A vast number of documents, such as journals and academic papers, exist worldwide. While some can be obtained free of charge from open-access databases, many must be purchased from paid databases. Such documents generally consist of bibliographic information, abstracts, and full text. Some databases (e.g., PubMed Central) make the full text available free of charge or at low cost. Therefore, this disclosure envisions, but is not limited to, the accumulation of the full text of such free or low-cost documents in document database 510 as positive and negative example documents.

[0056] The word classification unit 234 identifies and classifies each word in the received document by referring to the document database 510, determining whether it is a target word, a non-target word, or an indeterminate word. Specifically, the word classification unit 234 uses the positive example document corpus 512 and the negative example document corpus 514 to score each word (e.g., on a scale of 0 to 100) according to a predetermined statistical model, and classifies the word as a target word, a non-target word, or an indeterminate word based on the assigned score relative to a predetermined threshold. For example, the word classification unit 234 classifies words that may or may not be target words depending on the context, or words whose calculated score is a neutral score (e.g., near 0), as indeterminate words and reserves the classification of whether they are target words or not. As described above, the trained word selection model 236 estimates whether an indeterminate word is a target word or a non-target word.

[0057] In this example, the classification of target words according to the scoring by the word classification unit 234 is performed as follows. That is, the word classification unit 234 calculates the score (accuracy) of the extracted words using a scoring function, for example, shown in the following formula.

number

number

[0058] In other words, Equation 1 above represents the ratio of the number of documents in the entire document corpus D (i.e., the positive example document corpus 512 and the negative example document corpus 514) in which a certain word appears (for example, if the word appears (exists) even once in a document, it is counted as +1) to the number of documents N+ in which the word appears in positive example document D+ of the positive example document corpus 512.

[0059] For example, consider the scores of the words t1 "EGFR" (Epidermal Growth Factor Receptor) and t2 "the" contained in a document in a specific field. For simplicity, α = 0.

[0060] The document database 510, for example as shown in Figure 6, registers a document set D consisting of a positive example document set D+ and a negative example document set D-. Document d in the positive example document set D+ and the negative example document set D- contains "EGFR" and "the".

[0061] From Equation 1, the number of positive example documents containing word t, N+, is: N+=|d∈D+:t∈d| Therefore, the number of documents N that contain word t in the entire document D is N = |d∈D:t∈d| Therefore, the score for word t1 "EGFR" is,

number

[0062] Similarly, the score for word t2 "the" is,

number

[0063] The word classification unit 234 calculates a score for each of the extracted words as described above, identifies words whose calculated scores exceed the first threshold (in the example above, word t1 "EGFR") as target words, and classifies (labels) them.

[0064] For example, in a document in the field of cancer genomics, the word classification unit 234 can identify and classify the following group of words as target words: L858R / E545K / H1047R / T790M / EGFR-TKIs / EGFR-TKI / G12C / G13D / MuTect / GISTIC / E542K / ASXL1 / FLT3-ITD / V600E / NRAS / G12V / …

[0065] On the other hand, the classification of non-selected words by the word classification unit 234 is performed as follows. Specifically, the word classification unit 234 calculates the score (incorrectness) of the extracted words using the inverse document frequency function (idf), which is shown by the following formula, for example.

number

[0066] Similarly, we consider the scores using the inverse document frequency function for the entire set of positive example documents D+ and the entire set of negative example documents D- shown in Figure 6.

[0067] In other words, from Equation 5, the score of word t1 "EGFR" using the inverse document frequency function is:

number

number

[0068] In this example, we used the inverse document frequency function (idf) for explanation, but the document frequency function (df) may also be used. In this case, words whose calculated scores exceed a predetermined threshold will be identified and classified as non-selection words.

[0069] In this way, the word classification unit 234 scores each word according to a predetermined statistical model and, based on a predetermined threshold, classifies the word into one of the following categories: selected words, non-selected words, or indeterminate words, according to the score assigned. Then, as described above, the word classification unit 234 passes the words classified as selected words and non-selected words to the text post-processing unit 240, and also passes each word classified as selected words, non-selected words, or indeterminate words to the word selection model 236. As a result, as described above, the word selection model 236 estimates the label of the input indeterminate word (i.e., whether the word is a selected word or not) while taking its context into consideration.

[0070] As described above, this embodiment offers the same advantages as the first embodiment, and by utilizing the text of documents that can be obtained free of charge or at low cost as positive and negative example documents, it becomes possible to construct a document display support system 100 efficiently and at low cost.

[0071] [Third Embodiment] This embodiment is a variation of the second embodiment described above and is directed toward a method for constructing a document database comprising a positive example document corpus and a negative example document corpus.

[0072] Figure 7 is a block diagram showing an example of a schematic configuration of a document display support system according to one embodiment of the present invention. In the figure, components that are the same as those in the above embodiment are denoted by the same reference numerals, and therefore, their descriptions will be omitted as appropriate below. That is, as shown in the figure, the document display support system 100 of this embodiment differs from the configuration shown in the second embodiment in that a document sorting unit 710 is added. Also, in the figure, the document receiving unit 220, document analysis unit 230, and text post-processing unit 240 shown in Figure 2, etc., are simply shown as a document processing unit 260. The document sorting unit 710 may be configured to include a document extraction unit 712 and a document classification model 714, as shown in the figure.

[0073] The Specific Field Document Corpus 720 is a corpus of documents with a high frequency of occurrence of words that should be selected from documents in a specific field or domain. For example, when constructing the Standard Example Document Corpus 512 to select words related to the cancer genomics field, papers referenced from knowledge databases related to the cancer genomics field, such as "OncoKB" and "COSMIC," may be selected. In this disclosure, the documents registered in the Specific Field Document Corpus 720 are, for example, abstracts of documents such as journals and academic papers, and do not necessarily include the full text. The Specific Field Document Corpus 720 has a smaller number of registered documents (i.e., abstracts) compared to the General Field Document Corpus 730 described later (for example, about 15,000).

[0074] The General Subject Document Corpus 730 is a general subject document corpus with a broader domain than the specific subject document corpus 720. For example, the General Subject Document Corpus 730 targets abstracts of papers provided by PubMed, a free search engine for life sciences and biomedical fields, including cancer genomics. Because the General Subject Document Corpus 730 has a massive number of documents (for example, about 40 million), a predetermined number of documents can be randomly extracted by the document extraction unit 712.

[0075] The document extraction unit 712 randomly extracts a predetermined number of documents (i.e., abstracts) from the general domain document corpus 730 (for example, substantially the same number as the number of documents registered in the specific domain document corpus 720). By making the number of documents from the general domain document corpus 730 substantially the same as the number of documents registered in the specific domain document corpus 720, bias in the trained model can be prevented. The document extraction unit 712 inputs the extracted documents as training data into the document classification model 714. Note that "substantially the same number" does not mean that the two numbers are exactly the same, but rather that the difference is within a range that is obvious to a person skilled in the art and does not affect the training results.

[0076] The document classification model 714 is a pre-trained model that has undergone machine learning to calculate the accuracy of document recognition according to the frequency of word occurrences in a specific field, using a set of documents registered in the specific field document corpus 720 and a set of documents extracted from the general field document corpus 730 via the document extraction unit 712. In this disclosure, the document classification model 714 uses "Bag of Words" to feature vectorize the words contained in the text of the documents and performs machine learning using a known machine learning algorithm. Bag of Words is known in natural language processing as a method for representing data written in natural language, i.e., text, as vectors. In other words, the document classification model 714 calculates the frequency of occurrence of each word appearing in each document registered in the specific field document corpus 720, while also calculating the frequency of occurrence of each word appearing in each general document extracted by the document extraction unit 712, and then forms a pre-trained model by feature vectorizing the words according to the calculated frequency of occurrence. Note that other feature vectors, such as the word embedding vectors mentioned above, may be used instead of Bag of Words.

[0077] After machine learning as described above, the document classification model 714 calculates the accuracy of a document based on the input document set, thereby determining whether the document is correct. example Is it a document? Negative examples It estimates whether something is a document. That is, the document classification model 714 compares the calculated accuracy with a predetermined threshold, and if the calculated accuracy exceeds the predetermined threshold, it classifies the document as a document. example The document is classified and registered in the positive example document corpus 512. On the other hand, if the calculated accuracy of the document classification model 714 does not exceed a predetermined threshold, the document is removed. Negative examplesThe documents are classified and registered in the negative example document corpus 514. In this disclosure, the documents input to the trained document classification model 714 are, for example, the text body of a document and do not need to include bibliographic information or an abstract. More specifically, the documents may be, for example, the text body of a document referenced from PubMed Central (PMC). Furthermore, the number of documents (text bodies) input to the trained document classification model 714 is very large compared to the number of documents (abstracts) input for machine learning, and may be, for example, 1 million or more.

[0078] In this way, the document sorting unit 710 uses the trained document classification model 714 to classify a very large number of documents into either positive example documents or negative example documents, thereby enabling the construction of a document database 510 that includes the positive example document corpus 512 and the negative example document corpus 514.

[0079] As described above, this embodiment offers the same advantages as the above embodiment. In particular, since this embodiment uses a limited number of documents, and only a portion thereof (such as abstracts), as training data, the document classification model 714 can be efficiently trained. In particular, since the documents used as training data are extracted from a document corpus with a high frequency of words that should be selected in relation to a specific domain, and from a general field document corpus, which is a broader domain than the specific domain, the document classification model 714 can efficiently learn whether a document is a positive example or a negative example.

[0080] Furthermore, according to this embodiment, the machine learning-based document classification model 714 can calculate the accuracy of the input document according to the frequency of word occurrences in a specific field, and based on a predetermined threshold, the input document can be classified as correct. example Is it a document? Negative examples It will become possible to determine whether something is a document.

[0081] Furthermore, according to this embodiment, even words that are not registered in the conventional dictionary database but which should be identified as target words in documents of a specific field can be correctly estimated as target words. For example, in the field of cancer genome medicine, the word representing "mutation" (e.g., written as "L747S") is described in a great many patterns in documents of that field. Therefore, such words are practically impossible to handle with the conventional dictionary database, but according to this embodiment, they can be estimated as target words.

[0082] Next, the flow of a series of processes in the document display support system 100 shown in each of the above embodiments will be explained using Figures 8 to 10.

[0083] Figure 8 is a flowchart illustrating the process of constructing a document database in a document display support system 100 according to one embodiment of the present invention. This process is realized, for example, by a cloud computer 110 executing a predetermined program under the control of a processor.

[0084] As shown in the figure, in the document display support system 100, the document selection unit 710 acquires a predetermined number of document groups from the specific domain document corpus 720 (S801) and substantially the same number of document groups from the general domain document corpus 730 (S802). As described above, the document groups acquired here may be parts of documents, such as abstracts. By making the number of documents acquired from the specific domain document corpus 720 substantially the same as the number of documents acquired from the general domain document corpus 730, bias in the trained model can be prevented.

[0085] Next, the pre-trained document classification model 714 performs machine learning to calculate the accuracy of each document based on the frequency of word occurrences in a specific field, using each acquired document set (S803). In machine learning, for example, the frequency of word occurrences in the text is calculated using a Bag of Words and then vectorized.

[0086] After building the trained model, the document classification model 714 accepts input of a document to be classified and calculates the accuracy of the document (S804). The input document is a part of the document, for example, the main text. The document classification model 714 compares the calculated accuracy of the document with a predetermined threshold and determines whether the accuracy of the document exceeds the predetermined threshold (S805).

[0087] If the document classification model 714 determines that the accuracy of the document exceeds a predetermined threshold (Yes in S805), it classifies the document. Good example The document is registered as a positive example document corpus 512 (S806). On the other hand, if the document classification model 714 determines that the accuracy of the document does not exceed a predetermined threshold (No. in S805), the document is... Negative examples Register the document as a negative example document in Corpus 514 (S807).

[0088] The document sorting unit 710 then determines whether all documents to be classified have been classified (S808). If it determines that there are unclassified documents (No in S808), it returns to the process in S804 and repeats until all documents to be classified have been classified. On the other hand, if the document sorting unit 710 determines that all documents to be classified have been classified (Yes in S808), it terminates the document database construction process.

[0089] Figure 9 is a flowchart illustrating the learning process of a word selection model in a document display support system 100 according to one embodiment of the present invention. This process is realized, for example, by a cloud computer 110 executing a predetermined program under the control of a processor.

[0090] As shown in the figure, the word selection model 236 receives positive example documents from the positive example document corpus 512 and negative example documents from the negative example document corpus 514 of the document database 510, which were constructed according to the above-described process, as training data (S901). Subsequently, the word selection model 236 performs machine learning on each word in the positive example documents and negative example documents, taking into account its context, in order to enable the estimation of the label of that word (S902). The word selection model 236, having undergone machine learning in this manner, becomes capable of estimating and outputting a label for a given word as input.

[0091] Figure 10 is a flowchart illustrating the document display support process in a document display support system 100 according to one embodiment of the present invention. This process is realized, for example, by a cloud computer 110 executing a predetermined program under the control of a processor.

[0092] In other words, as shown in the figure, the document receiving unit 220 receives a document as a document to be displayed, in which it is desired to visually determine whether or not it is related to a specific field (S1001). Subsequently, the text preprocessing unit 232 extracts multiple words from the text body of the received document, for example, using morphological analysis (S1002).

[0093] Next, the word classification unit 234 refers to the document database 510 for the extracted words and calculates a score based on a predetermined scoring function (S1003). For example, the calculated score is the correctness (e.g., a value from 0 to 100) and the incorrectness (a value from -100 to 0). The word classification unit 234 compares the calculated score with a predetermined threshold (S1004) and classifies the word. That is, if the calculated score (but a positive value) of the word exceeds the first threshold, the word classification unit 234 classifies it as a word to be selected (S1005). On the other hand, if the calculated score (but a negative value) of the word does not exceed the second threshold, the word classification unit 234 classifies it as a word not to be selected (S1006). Furthermore, if the word classification unit 234 finds that the word is otherwise, that is, if the calculated score does not exceed the first threshold but exceeds the second threshold, it classifies it as an uncertain word (S1007).

[0094] Next, the word selection model 236 estimates whether an indeterminate word is a target word or not, based on the input words (e.g., target words, non-target words, and indeterminate words), according to the surrounding context (S1008). As a result of this estimation, the word selection model 236 assigns, for example, a label to the input word indicating whether it is a target word or a non-target word.

[0095] After the word selection model 236 performs estimation, the text post-processing unit 240 modifies the text of the received document based on the estimation results output by the word selection model 236 so that the selected words are visually distinguishable and displayed (S1009). For example, the text post-processing unit 240 generates output data in which character attribute information is added to the document so that the selected words are highlighted. The document display unit 250 then displays the modified document on the user interface based on the output data generated by the text post-processing unit 240.

[0096] As described above, according to the embodiment described above, the document display support system 100 can use the word selection model 236 to estimate words that are important or highly relevant in a particular field of a document that the user intends to refer to, and can display the estimated words in a visually distinguishable manner.

[0097] The embodiments described above are illustrative examples for illustrating the present invention and are not intended to limit the invention to these embodiments only. The present invention can be implemented in various forms without departing from its spirit.

[0098] For example, this disclosure uses the field of cancer genomics as an example of a specific area, but it is not limited to this. Furthermore, documents may include not only academic papers, but also reports, tables, electronic medical records, and so on.

[0099] Furthermore, while the document display support system 1 described above is configured to visually distinguish and display selected words in a document, it is not limited to this configuration. For example, it may be configured to further utilize the selected words. That is, words that the document display support system 1 classifies as selected words and highlights in the displayed document indicate that they are semantically important words in the document. Therefore, the selected words may be used as metadata (tags) or keywords in the document.

[0100] For example, the document display support system 1 registers words classified as target words by the word selection model 236 as tags and / or keywords associated with a document in a database. When a user refers to the document, the document display support system 1 provides the user with a preview of these associated words, allowing the user to quickly grasp the content of the document.

[0101] As another example, the document display support system 1 may have a function to suggest search queries based on selected words registered in association with documents, and a function to weight the search results.

[0102] As yet another example, the document display support system 1 may have a function to identify parts of a document, such as sentences, paragraphs, or chapters, in which the selected word appears frequently, and to generate a summary of the document using natural language processing, focusing on the identified parts.

[0103] As another example, the document display support system 1 may, in addition to or instead of visually distinguishing and displaying the selected words in the document, also have a configuration that reads the selected words aloud using voice or other means.

[0104] Furthermore, in the methods disclosed herein, steps, operations, or functions may be performed in parallel or in different orders, provided that this does not result in inconsistent outcomes. The steps, operations, and functions described are provided merely as examples, and some of these steps, operations, and functions may be omitted, combined into a single step, operation, or function, without departing from the spirit of the invention.

[0105] Furthermore, although various embodiments are disclosed herein, specific features (technical matters) in one embodiment can be added to or replaced in other embodiments, with appropriate modifications, and such forms are also included in the gist of the present invention. [Explanation of Symbols]

[0106] 100…Document display support system 110…Cloud Computing 120…Client Computers 210...Word Database 220... Document Reception Department 230...Document Analysis Department 232...Text preprocessing 234...Word Classification Section 236…Word selection model 240...Text post-processing 250...Document display section 260…Document Processing Department 510…Document Database 512…Example Document Corpus 524... Negative example document corpus 710...Document Sorting Department 712...Document Extraction Unit 714…Document classification model 720...Specific field document corpus 730... General Subject Document Corpus

Claims

1. A system for estimating and outputting important words in a specific field from a document received in that specific field, A database including a positive example document corpus consisting of positive example documents defined as cases related to a specific field, and a negative example document corpus consisting of negative example documents defined as cases not related to the said specific field, A word classification unit that classifies each of the multiple words in the received document into at least one of the following: selected words, not selected words, and uncertain words, according to a score calculated based on the number of times the word appears in the positive example document and the negative example document, respectively. A word selection model is subjected to machine learning, using pre-assigned learning labels for words corresponding to at least one of the target words and one of the non-target words, and taking multiple words extracted from the positive example documents in the positive example document corpus and the negative example documents in the negative example document corpus as input, to output as an estimation result a label indicating either a target word or a non-target word according to the context in which the word relates. It comprises an output unit that outputs output data, The word classification unit inputs at least the words classified as indeterminate words into the word selection model, The word selection model, when estimating a label indicating a target word for a word that has been classified as an indeterminate word input from the word classification unit, classifies the word as a target word. The output unit outputs the words classified as the target words as output data indicating words important in the specific field. system.

2. The word classification unit classifies the input word into at least one of the selected word, the non-selected word, and the undetermined word based on the number of times the word appears in the positive example document corpus and the number of times the word appears in the negative example document corpus. The system described in claim 1.

3. The word classification unit classifies each of the multiple words in the received document into at least one of the following: a word to be selected, a word not to be selected, or an undetermined word, according to a score calculated based on the number of positive example documents in which the word appears in the positive example document corpus and the number of negative example documents in which the word appears in the negative example document corpus. The system according to claim 1.

4. The word classification unit classifies the input word into at least one of the selected word, the non-selected word, and the undetermined word, based on the number of positive example documents in which the word appears in the positive example document corpus and the number of negative example documents in which the word appears in the negative example document corpus. The system described in claim 3.

5. The system further comprises means for distinguishing and outputting, from other words in the received document, a word that has been classified as an indeterminate word input to the word selection model subjected to machine learning, a word for which a label indicating the target word has been estimated, from other words in the received document. The system according to any one of claims 1 to 4.

6. A method performed in a system for estimating and outputting important words in a document of a specific field, To construct a database that includes a positive example document corpus consisting of positive example documents defined as cases related to the aforementioned specific field, and a negative example document corpus consisting of negative example documents defined as cases not related to the aforementioned specific field, Using pre-assigned learning labels for words corresponding to at least one of the target words and one of the non-target words, a word selection model is constructed by applying machine learning to multiple words extracted from the positive example documents in the positive example document corpus and the negative example documents in the negative example document corpus of the constructed database, and outputting a label indicating either a target word or a non-target word as an estimation result, according to the context related to the word. Each of the multiple words in the received document is classified into at least one of the following: selected words, non-selected words, and undetermined words, according to a score calculated based on the number of times that word appears in the positive example document and the negative example document, respectively. At least the words classified as indeterminate words are input into the word selection model, When estimating a label indicating the target word for a word classified as an indeterminate word by the word selection model, the word is classified as a target word. This includes outputting the words classified as target words as important words in the specific field, method.

7. This further includes constructing a document classification model that uses machine learning to calculate the accuracy of document classification according to the frequency of word occurrences in the specified field, using a first set of documents in a specific field document corpus related to the specified field and a second set of documents in a general field document corpus covering a broader field than the specified field as training data. Building the aforementioned database means Accepting documents to be classified, Using the document classification model that has undergone machine learning, the accuracy of the received documents to be classified is calculated. This includes classifying the documents to be classified as either good example documents or bad example documents according to the calculated accuracy score. The method described in claim 6.

Citation Information

Patent Citations

  • Medical report creation support device

    JP2012141797A

  • Information processing apparatus

    JP2016207071A

  • Display control apparatus, program, display system, display control method and display data

    JP2019105751A

  • Medical data system and method

    US20150046190A1