Text generation system, text generation method, and program

The text generation system addresses the instability of deep learning-based medical record generation by integrating multiple inferences and user-editable candidates, enhancing accuracy and stability.

JP7799735B2Active Publication Date: 2026-01-15CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024050946
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-03-27
Publication Date
2026-01-15
Estimated Expiration
2044-03-27

Smart Images

  • Figure 0007799735000001
    Figure 0007799735000001
  • Figure 0007799735000002
    Figure 0007799735000002
  • Figure 0007799735000003
    Figure 0007799735000003
Patent Text Reader

Abstract

To provide a sentence generation system, a sentence generation method, and a program that prevent instability of an inference result when deep learning is applied to sentence generation, and stably generate sentence candidates.SOLUTION: A sentence generation system according to the present invention comprises: an input candidate information acquisition unit that acquires input candidate information to be input to a language model that outputs sentence information from input information; a sentence information acquisition unit that acquires a plurality of pieces of sentence information different from each other by inputting the input information based on the input candidate information to the language model; and a determination unit that determines report sentence candidates by using the plurality of pieces of sentence information.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a text generation system, a text generation method, and a program. [Background technology]

[0002] In the medical field, electronic medical records are created by doctors and technicians, and while efforts have been made to reduce the labor required through voice input and other methods, the writing of the records has still had to be done manually.

[0003] To further reduce the burden on medical professionals, the application of deep learning technology is being considered. It is expected that it will be applied not only to interpreting medical images, but also to writing electronic medical records in the future. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] OpenAI, “GPT-4 Technical Report,”arXiv:2303.08774v3,2023. Summary of the Invention [Problem to be solved by the invention]

[0005] Since the creation of electronic medical records requires text generation technology, it is expected that deep learning technology, such as the generative AI disclosed in Non-Patent Document 1, will be used. However, text generation based on deep learning has the problem of instability, where inference results fluctuate stochastically (randomly) and even slight differences in input information can cause the inference results to fluctuate significantly.

[0006] The present invention has been made in consideration of the above-mentioned problems, and aims to provide a text generation system, a text generation method, and a program that suppress the instability of inference results when deep learning is applied to text generation, and that stably generate text candidates to be input into reports, etc.

[0007] In addition to the above-mentioned objectives, the achievement of effects derived from the various configurations shown in the description of the invention below, which cannot be obtained by conventional techniques, can also be positioned as another objective of the disclosure of this specification. [Means for solving the problem]

[0008] In order to solve the above problems, the text generation system according to the present invention includes an input candidate information acquisition unit that acquires input candidate information to be input from input information to a language model that outputs text information, a text information acquisition unit that acquires a plurality of mutually different pieces of text information by inputting input information based on the input candidate information to the language model, and a determination unit that determines report text candidates using the plurality of pieces of text information. a display control unit that causes the report sentence candidates determined by the determination unit to be displayed on a display unit, and the display control unit displays adopted information that has been adopted from the text information by the determination unit as information to be used for the report sentence candidates and non-adopted information that has not been adopted in a distinguishable manner. It is characterized by the following. [Effects of the Invention]

[0009] According to the present invention, it is possible to suppress instability of inference results and generate sentence candidates stably. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 1 is a diagram showing a functional configuration of a sentence generation system according to a first embodiment. [Figure 2] 4 is a flowchart illustrating an example of a processing procedure of the sentence generation system according to the first embodiment. [Figure 3] FIG. 4 is a diagram showing an example of input candidate information according to the first embodiment. [Figure 4] FIG. 3 is a diagram showing an example of text information according to the first embodiment. [Figure 5] FIG. 10 is a diagram showing an example of report sentence candidates according to the first embodiment input into a medical record. [Figure 6] FIG. 10 is a diagram showing an example of report sentence candidates input into a medical record according to the embodiment of Modification 1-1. [Figure 7] FIG. 10 is a diagram showing an example of report sentence candidates input into a medical record according to the embodiment of Modification 1-1. [Figure 8] FIG. 10 is a diagram showing a functional configuration of a sentence generation system according to a second embodiment. [Figure 9] 10 is a flowchart illustrating an example of a processing procedure of a sentence generation system according to a second embodiment. [Figure 10] FIG. 11 is a diagram illustrating an example of an input dialog for input candidate information according to the second embodiment. [Figure 11] FIG. 10 is a diagram showing an example of input information according to the second embodiment. [Figure 12] FIG. 10 is a diagram showing an example of text information according to the second embodiment. [Figure 13] FIG. 10 is a diagram showing an example of report sentence candidates according to the second embodiment input into a medical record. [Figure 14] FIG. 10 is a diagram showing an example of report sentence candidates input into a medical record according to the embodiment of modified example 2-1. DETAILED DESCRIPTION OF THE INVENTION

[0011] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. However, the components described in the embodiments are merely examples, and the technical scope of the present invention is determined by the claims, and is not limited to the individual embodiments described below.

[0012] First Embodiment The sentence generation system according to the first embodiment is a system that generates candidate sentences for reports to be recorded in a medical record (hereinafter referred to as report sentence candidates) from sentences input by a user, which are examples of sentence candidates. This system has the function of performing inference using a language model based on the sentences input by the user and acquiring sentence information that will become report sentence candidates. The language model used here assumes that the output results obtained when the same sentence is input are not consistent. This system acquires multiple different sentence information by having the language model perform inference multiple times on the same input sentence. This system is characterized by its ability to create more stable report sentence candidates by integrating multiple different sentence information acquired from the language model.

[0013] According to the present invention, it is possible to provide the user with more accurate report sentence candidates than when report sentence candidates are generated by simply inputting the user's input into a language model.

[0014] In this embodiment, an example will be described in which report sentence candidates to be written in the diagnosis column are generated from a report of medical history and findings input by a user as an image interpretation report.

[0015] However, even in cases where the distinction between medical history and findings is unclear, or where the information is other than text, such as a radiological report, or medical images referenced during radiological interpretation, the effects of the present invention can be obtained as long as the information can be used for inference using a language model.

[0016] The functional configuration and processing of a text generation system 10 including the text generation system of this embodiment will be described below with reference to Fig. 1. Fig. 1 is a block diagram showing an example of the configuration of the text generation system 10 of this embodiment. The text generation system 10 is communicably connected to a language model 22 via a network 21.

[0017] The network 21 includes, for example, a LAN (Local Area Network) and a WAN (Wide Area Network).

[0018] The language model 22 has a function of generating sentences from structured sentences, keywords, unstructured sentences, images, etc. The language model 22 also has a function of predicting sentences that are suitable as answers based on randomness and probability from input information. The language model 22 is realized, for example, by GPT (Generative Pre-trained Transformer) or BERT (Bidirectional Encoder Representations from Transformers). The sentence generation system 10 can acquire sentences predicted by the language model 22 via the network 21.

[0019] The sentence generation system 10 includes a communication IF (Interface) 31 (communication unit), a ROM (Read Only Memory) 32, a RAM (Random Access Memory) 33, a storage unit 34, an operation unit 35, a display unit 36, and a control unit 37.

[0020] The communication IF 31 (communication unit) is configured with a LAN card or the like, and realizes communication between an external device (e.g., the language model 22) and the text generation system 10. The ROM 32 is configured with a non-volatile memory or the like, and stores various programs. The RAM 33 is configured with a volatile memory or the like, and temporarily stores various pieces of information as data. The storage unit 34 is configured with an HDD (Hard Disk Drive) or the like, and stores various pieces of information as data. The operation unit 35 is configured with a keyboard, mouse, touch panel, etc., and inputs instructions from a user (e.g., a doctor or radiologist) to various devices.

[0021] The display unit 36 ​​is configured with a display or the like, and displays various information to the user. The control unit 37 is configured with a CPU (Central Processing Unit) or the like, and controls the overall processing in the sentence generation system 10. The control unit 37 is configured to include, as its functional configuration, an input candidate information acquisition unit 51, a generation unit 52, a sentence information acquisition unit 53, a determination unit 54, a display control unit 55, and an edit acceptance unit 56.

[0022] The input candidate information acquisition unit 51 acquires information to be processed (input candidate information) from the operation unit 35. This information is, for example, a report created by a doctor or a radiologist. That is, the input candidate information acquisition unit 51 corresponds to an example of an input candidate acquisition unit that acquires input candidate information. In this embodiment, an example will be described in which an interpretation report into which a medical history and findings have been input is acquired as input candidate information, but an interpretation report into which information other than a medical history and findings has been input may also be acquired as input candidate information. Furthermore, text other than an interpretation report may also be acquired as input candidate information. Furthermore, the input candidate information is not limited to text information, and may be medical information other than text, such as a medical image referred to during interpretation.

[0023] The generation unit 52 generates input information necessary for inference using the language model 22, from the sentences etc. acquired by the input candidate information acquisition unit 51. That is, the generation unit 52 corresponds to an example of a generation means that generates input information for the language model from the input candidate information.

[0024] The text information acquisition unit 53 inputs the input information generated by the generation unit 52 into the language model 22 multiple times, and acquires the results of inference made by the language model 22 based on the input information each time. The language model 22 in this embodiment is assumed to have randomness, and the output text changes even with the same input. In other words, the multiple pieces of text information acquired by the text information acquisition unit 53 are not necessarily identical and may differ from one another. In other words, the text information acquisition unit 53 corresponds to an example of a text information acquisition means that acquires multiple pieces of text information that are different from one another from the language model 22 based on the input information.

[0025] The determination unit 54 generates report sentence candidates to be recorded in the medical record from the multiple pieces of text information acquired by the text information acquisition unit 53, using information that is frequently included in each piece of text information. The determination unit 54 generates the report sentence candidates, for example, by integrating the pieces of text information. This integration may be performed internally by the determination unit 54, or an external function such as the language model 22 may perform the integration to generate the report sentence candidates. In other words, the determination unit 54 corresponds to an example of an integration processing means that integrates the text information acquired by the text information acquisition unit 53.

[0026] The display control unit 55 causes the display unit 36 ​​to display the report sentence candidates determined by the determination unit 54. The display control unit 55 also switches the display information in response to a user's operation on the edit receiving unit 56, which will be described later.

[0027] The editing accepting unit 56 accepts editing information by the user for the report sentence candidate displayed on the display unit 36 ​​by the display control unit 55. The report sentence candidate may be displayed directly in the diagnosis field of the radiology report, or may be displayed elsewhere and reflected in the diagnosis field of the radiology report based on the user's editing work. In other words, the editing accepting unit 56 corresponds to an example of an editing accepting means that reflects the results of user editing of the report sentence generated by the determining unit 54 in the radiology report. Note that editing refers to, for example, corrections such as adding or deleting sentences to the report sentence candidate, or an instruction to confirm the reflection of the report sentence candidate in the report.

[0028] Each of the components of the above-described sentence generation system 10 functions according to a computer program. For example, the control unit 37 (CPU) uses the RAM 33 as a work area to read and execute a computer program stored in the ROM 32 or the storage unit 34, thereby realizing the function of each component. Note that some or all of the functions of the components of the sentence generation system 10 may be realized using dedicated circuits. Also, some of the functions of the components of the control unit 37 may be realized using a cloud computer.

[0029] For example, a computing device located at a different location from the text generation system 10 may be communicably connected to the text generation system 10 via the network 21. Then, the text generation system 10 and the computing device may transmit and receive data to each other, thereby realizing the functions of the components of the text generation system 10 or the control unit 37.

[0030] Next, an example of a flow of generating report sentence candidates in the sentence generation system 10 according to this embodiment will be described with reference to FIG.

[0031] 2 is a flowchart showing an example of the processing procedure of the sentence generation system 10. In this embodiment, an example will be described in which report sentence candidates to be written in the diagnosis section of a radiology report are generated from medical history and findings information written in the radiology report. However, this embodiment can also be applied to sentences in which the distinction between medical history and findings is unclear, sentences other than radiology reports, or cases in which processing is performed based on information other than sentences, such as medical images referenced during radiology interpretation.

[0032] (Step S101: Obtaining input candidate information) In step S101, the input candidate information acquisition unit 51 acquires the medical history and findings text of the radiology report input by the user via the operation unit 35, and stores it in the RAM 33. An example of the radiology report input by the user is shown in FIG. 3. FIG. 3 shows an example of a radiology report in which the medical history and findings columns have been filled in by the radiology doctor who is the user, but the diagnosis column is left blank. In this embodiment, such a radiology report is acquired as input candidate information.

[0033] (Step S102: Generation of input information) In step S102, the generation unit 52 processes the input candidate information acquired by the input candidate information acquisition unit 51 into input information that can be inferred by the language model 22. The processing in this step corresponds to preprocessing of the subsequent text information acquisition processing, and aims to set inference instruction content in the input candidate information.

[0034] In this embodiment, a sentence is generated by adding the inference instruction content "Generate a sentence of {diagnosis} from the following {medical history} and {findings}. {diagnosis} shall be in bullet points" to the beginning of the input candidate information shown in Fig. 3, and this is used as input information to the language model (hereinafter simply referred to as input information). Note that the information added to the input candidate information is not limited to the above character string, and may be any information that can convey the inference instruction content to the language model 22. Alternatively, it may be in a form other than a character string, such as parameter information.

[0035] (Step S103: Input the input information to the language model) In step S103, the text information acquisition unit 53 inputs the input information generated by the generation unit 52 to the language model 22 via the communication IF 31 and the network 21. As a result, inference by the language model 22 is executed.

[0036] (Step S104: Obtain sentence information from the language model) In step S104, the text information acquisition unit 53 acquires the inference result performed by the language model 22 in step S103.

[0037] The sentence generation system 10 repeatedly executes steps S103 to S104 a plurality of times, thereby receiving a plurality of inference results from the language model 22 for one piece of input information.

[0038] The configuration of the language model 22 used in this embodiment is described below. The language model 22 in this embodiment is a probabilistic language model (a model that probabilistically predicts the word that follows the immediately preceding sentence and constructs sentences by repeating this process), and when selecting candidate words, it randomly selects from among words with high probabilities. Therefore, even if the text information 53 repeatedly performs inference on the language model 22 using the same input information, the text information received from the language model 22 will not be constant.

[0039] FIG. 4 shows an example of text information received by the text information acquisition unit 53 from the language model 22. FIG. 4 shows inference results as text to be entered in the diagnosis section of a radiology report, with each cell in the table listing the results of multiple inferences. The inference results differ from one another, and differences may occur not only in cases where similar information is expressed using different character strings, but also in the information itself. For example, "SCC" and "lung squamous cell carcinoma" are simply expressed differently, but the meaning (information) they indicate remains the same. On the other hand, the terms "sarcoidosis" and "metastasis to the spleen" are not included in other text information. In this way, the text information 53 acquires multiple pieces of text information that differ due to the randomness of the language model 22, etc.

[0040] (Step S105: Integrate multiple pieces of text information) In step S105, the determination unit 54 integrates the multiple pieces of text information acquired by the text information acquisition unit 53 to generate a single report sentence candidate. In this embodiment, the language model 22 generates sentences probabilistically, and each piece of text information differs from the others. The determination unit 54 determines that information frequently included in the multiple pieces of text information is important or highly accurate, and conversely, determines that information that is only frequently included is unimportant or low-accuracy information (i.e., noise) that appears randomly. A specific example of this determination is a procedure in which, after standardizing terminology variations using a medical dictionary, information indicating possibility (such as "it is possible" or "it is possible") and information indicating definite content (such as "it is" are separately aggregated). Figure 4 shows that "lung cancer (SCC)" appears four times as information indicating possibility, and "metastasis (lymph nodes, bones, liver)" appears three or more times. Furthermore, "old granuloma" and "nonspecific post-inflammatory changes" also appear, but their frequencies are not high, being only once each. However, no definite content information is included. Therefore, the determining unit 54 generates a report sentence candidate by combining two pieces of information that have a high frequency among the pieces of information indicating the possibility. Fig. 5 shows an example in which the report sentence candidate generated by the determining unit 54 is entered in the diagnosis column 512 of the radiology report shown in Fig. 3.

[0041] The determination process is not limited to this example, and for example, the language model 22 may be instructed to integrate multiple pieces of text information and generate report sentence candidates from the language model 22. In this case, it is desirable to instruct the language model 22 to prioritize information indicating definite content over information indicating possibility, or to prioritize information with a high frequency of occurrence when integrating the text information.

[0042] (Step S106: User confirmation and correction) In step S106, the display control unit 55 causes the display unit 36 ​​to display the report sentence candidates generated by the determination unit 54 as input candidates for the radiology report. An example of a display screen displayed by the display control unit 55 on the display unit 36 ​​is shown in FIG. 5. In FIG. 5, the report sentence candidates generated by the determination unit 54 are displayed in a diagnosis column 512 on the radiology report 511 together with the medical history and findings shown in FIG. 3. After the user views this result, the edit receiving unit 56 receives information on corrections made to the diagnosis column via the operation unit 35 as necessary, and finalizes the contents of the radiology report.

[0043] In the above description, the input candidate information is a sentence, but similar processing can be performed even if image information is combined.

[0044] According to this embodiment, by performing multiple inferences on a language model and integrating the inference results, it is possible to provide the user with more stable report sentence candidates from unstable inference results.

[0045] Although the determination unit 54 generates report sentence candidates by placing importance on information that appears multiple times in step S105, other rules may also be applied. For example, medically important information may be actively adopted as report sentence candidates, such as by adopting keywords important for determining cancer metastasis even if they appear infrequently. Target keywords may be flagged in the medical dictionary, or individual keyword tables may be prepared for each purpose of radiological interpretation and switched depending on the purpose. Information not included in the question may be actively adopted as report sentence candidates as knowledge that is difficult to infer but should be emphasized. By actively adopting information that appears infrequently as report sentence candidates in this way, it is possible to provide the user with report sentences that suppress noise in inference and reduce the possibility of oversight.

[0046] Although the present embodiment has been described above, the present invention is not limited to this embodiment, and can be modified and changed within the scope of the claims.

[0047] (Variation 1-1) In the first embodiment, an example was shown in which the display control unit 55 displays the report sentence candidates generated by the determination unit 54 on the display unit 36 ​​in step S106. This method is capable of ultimately presenting highly accurate report sentence candidates, but has the disadvantage that unaccepted information, which is information rejected by the determination unit 54, is hidden and unavailable to the user. In this modified example, by presenting the report sentence candidates generated by the determination unit 54 along with the text information before integration to the user, highly accurate report sentence candidates are presented and a process for assisting the user in making corrections is exemplified. That is, the display control unit 55 displays the report sentence candidates in association with the presence or absence of unaccepted information that was not adopted by the determination unit 54 as an input candidate for the report from the text information.

[0048] FIG. 6 is a display example of report sentence candidates displayed on the display unit 36 ​​by the sentence generation system 10 in this modified example. Unlike FIG. 5, in FIG. 6, the display control unit 55 also displays a list of sentence information 612 before integration on the display unit 36. Furthermore, the characters in the range adopted in the integrated report sentence candidates are grayed out, differentiating the display method from information not adopted as targets for integration (non-adopted information). That is, the display control unit 56 displays information adopted by the determination unit 54 as an input candidate for the report from the sentence information and non-adopted information in a distinguishable manner. By displaying the adopted information and non-adopted information in different display formats in this way, the user can easily identify the keywords that have become non-adopted information, and can also easily understand the context of the keywords by reading the entire sentence information.

[0049] The user can edit the report sentence candidate via the operation unit 35, referring to the additionally displayed text information 612 in FIG. 6, and complete the radiology report. The determination unit 54 may also be configured to monitor the user's selection operation status on the operation unit 35, and when the editing acceptance unit 56 detects the user's click on unacceptable information in the text information, the unacceptable information may be added to the accepted information and the report sentence updated. This allows the user to modify the report candidate simply by clicking on the rejected keyword in the text information. In other words, the editing acceptance unit 56 is characterized in that it accepts the user's editing information for the unacceptable information, and updates the report sentence candidate by replacing the selected keyword with the accepted information for the unacceptable information.

[0050] FIG. 7 illustrates an example in which the pre-integration information is used by selection, rather than displaying the entire information before integration. The editing reception unit 56 does not display the report sentence candidates as they are, but instead displays a portion of information that is strongly related but not adopted as a report sentence candidate due to its low frequency of occurrence in a user-selectable format. Alternatively, the editing reception unit 56 displays a portion of information that was adopted as information but not adopted during integration in a user-selectable format. In FIG. 7(a), to indicate the presence of a portion of the metastasis site (pancreas) that was not adopted as a report sentence candidate, an underline 721 is displayed in the sentence of the site adopted as a report sentence candidate, indicating that it can be selected. Furthermore, to indicate that there was inconsistency in the expression in the information before integration, and that unified information (SCC, lung squamous cell carcinoma) exists, an underline 722 is displayed in the unified portion (SCC), indicating that it can be selected. In FIG. 7(b), information 712 in the form of checkboxes is displayed, including information that was not adopted. When the user moves the mouse pointer 730 over the list of metastasis destinations 721, the editing reception unit 56 displays a list of adopted / unadopted information in the form of checkboxes 712. Lymph nodes, bones, and liver have already been adopted, so their checkboxes are checked, while pancreas has not been adopted and is therefore unchecked. The user can change the metastasis site by checking or unchecking the checkboxes. Also, in FIG. 7( c), a list of terms before integration is displayed as radio button information 713. When the user moves the mouse 730 over the list of metastasis destinations 722, the editing reception unit 56 displays a list of adopted terms before integration as radio button information 713. Of SCC and lung squamous cell carcinoma, SCC has been adopted, so SCC is selected. The user can change the notation by selecting a term using the radio button.

[0051] The display control unit 55 may display on the display unit 36 ​​any portion of the report sentence candidates containing information that has not been adopted in a visually easy-to-identify format, such as by underlining, coloring the text or background, or making the text bold. Alternatively, the display control unit 55 may not make any changes to the initial display, and may only display that there are options when a selection is made with the mouse or the like or when the target range is in an editing state.

[0052] According to this modification, by presenting the information before integration to the user, it is possible to provide the user with additional auxiliary information for correcting the report text.

[0053] (Variation 1-2) The above embodiment and modified examples illustrate an example in which a report sentence candidate is edited by a user. While this method has the advantage of enabling the final report sentence to be created as intended by the user, in some cases, similar corrections may be required frequently. For example, the user may need to frequently correct the wording of the sentences obtained as report sentence candidates (the different uses of multiple words with the same meaning), the tone and style of the sentences, etc. In this modified example, the edits made by the user to the report sentence candidate by the edit receiving unit 56 are reflected when the determination unit 54 generates the next report sentence, thereby providing a function for correcting the report sentence candidate and illustrating a process that prevents the user from having to repeatedly make similar corrections.

[0054] 5 to 7, when the user modifies the report sentence candidate, the edit receiving unit 56 passes the edited content to the determination unit 54. The determination unit 54 incorporates the received edited content as a condition for the next and subsequent integrations, and uses it to combine the text information in the next step S105.

[0055] For example, if the user changes the description of "SCC" to "lung squamous cell carcinoma" in FIG. 7(c), the edit receiving unit 56 notifies the decision unit 54 of the change. Based on the received information, the decision unit 54 lowers the selection priority of the term "SCC" in the dictionary used during integration and increases the selection priority of "lung squamous cell carcinoma." By doing so, the next time a term meaning "SCC" or "lung squamous cell carcinoma" is output when generating report sentence candidates, "lung squamous cell carcinoma" can be preferentially adopted as a report sentence candidate rather than "SCC."

[0056] Furthermore, for example, if the user changes the tone of a report sentence candidate from "possibly" to "there is a possibility," the edit receiving unit 56 similarly notifies the determination unit 54 that it has been switched to "polite." Based on the received information, the determination unit 54 switches the tone parameter from "plain" to "polite," thereby changing the tone used when connecting the selected information in the next step S105 to "desumasu."

[0057] According to this modification, by reflecting the edits made by the user in the decision unit 54, it is possible to provide a mechanism that automatically brings the report sentence candidates for the next time and thereafter closer to the style of sentences desired by the user based on the results of the corrections made to the previous report sentence, thereby reducing the editing work load on the user.

[0058] While the above examples have been described with respect to examples of changing the selection priority of words used in report sentence candidates and changing parameters related to the tone of report sentence candidates, the present invention is not limited to these examples. For example, a report sentence edited by a user may be recorded, and the determination unit 54 may set conditions for the next processing. For example, when requesting the language model 22 to integrate text information, the determination unit 54 may set a condition such as "Generate a report sentence similar to the following report sentence: ~ (hereinafter, the recorded report sentence edited by the user) ~." Furthermore, multiple pieces of text information and the report sentence edited by the user may be passed to the determination unit 54. When requesting the language model 22 to integrate text information, the determination unit 54 may set a condition such as "Below are samples of the report sentence before integration and the report sentence after integration. This time, integrate so that it resembles this sample: ~ (hereinafter, the text information and report sentence) ~." This allows for the generation of text in a form similar to the text desired by the user, thereby reducing the editing load on the user.

[0059] (Variation 1-3) In the above embodiment and modified example, the same input information is input multiple times to the same language model 22, and multiple sentence information is acquired through multiple inferences. While this method is capable of presenting report sentence candidates to the user that are less affected by the randomness of the language model, there is a possibility that report sentence candidates that are biased toward a specific language model may be generated. In this modified example, the sentence information acquisition unit 53 performs inference using multiple different language models, thereby generating report sentence candidates that are not biased toward a specific language model.

[0060] In step S103, the text information acquisition unit 53 inputs the input information generated by the generation unit 52 into each of the multiple language models to perform inference, and acquires the results in step S104. At this time, inference for each language model may be performed once for each language model, or inference may be performed multiple times for each language model. The determination unit 54 integrates the text information acquired by the text information acquisition unit 53 from the multiple language models 22 to generate report sentence candidates.

[0061] According to this modification, by integrating the inference results of a plurality of language models, it is possible to obtain report sentence candidates that are not biased towards a specific language model.

[0062] (Variation 1-4) In the above-described first embodiment, an example has been described in which the acquired input candidate information is converted into input information by the generation unit 52, but the present invention is not limited to this. For example, the input candidate information acquisition unit 51 may omit the process of step S102 in the first embodiment and directly input the input candidate information acquired by input from the user to the language model 22 as input information.

[0063] For example, a model specialized for generating report text from free text may be used as the language model 22, and the input candidate information exemplified in the first embodiment may be directly inferred. In this way, it becomes unnecessary to add inference conditions to the input candidate information, and the language model 22 can generate sentence information directly from the input candidate information.

[0064] Furthermore, the text information acquisition unit 53 may omit the processing of step S102 by inputting the input candidate information acquired through user input and a pre-stored inference request template sentence into the language model 22. The content of the template sentence may be, for example, "generate a report sentence similar to the attached report sentence." In this way, it becomes unnecessary to add an inference condition to the input candidate information, and the language model 22 can directly generate text information without processing the input candidate information.

[0065] <Second embodiment> In the first embodiment described above, an example was shown in which inference using a language model is performed multiple times based on the same input information, and the results are integrated to suppress randomness in the language model and bias between language models, thereby stably generating report sentence candidates that include important information.

[0066] In the second embodiment, the generation unit 52 provides a language model with multiple pieces of input information that are different from each other, performs multiple inferences based on the input information, and integrates the results. This reduces the instability of report sentence candidate generation caused by the instability of input candidate information provided by the user. According to the present invention, it is possible to provide the user with report sentence candidates that are robust against the instability of input candidate information entered by the user.

[0067] Regarding the configuration of this embodiment, only the parts of an information processing system 80 according to this embodiment that differ from the first embodiment will be described using FIG. 8. An input candidate information acquisition unit 81 acquires information to be processed (input candidate information) from the operation unit 35. This information is parameters entered by a doctor or radiologist. In other words, the input candidate information acquisition unit 81 corresponds to an example of an input candidate acquisition means that acquires input candidate information. In this embodiment, an example will be described in which parameters used for report generation, such as the primary site and size of the primary tumor, are acquired as input candidate information, but other parameters may also be acquired as input candidate information.

[0068] As in the first embodiment, the generation unit 82 generates input information from the input candidate information acquired from the input candidate information acquisition unit 51. Furthermore, the generation unit 82 generates other input candidate information by partially changing the parameters of the input candidate information, and generates new input information from the other input candidate information.

[0069] The text information acquisition unit 83 acquires a plurality of pieces of text information by having the language model 22 perform inference on each of the plurality of pieces of input information generated by the generation unit 82.

[0070] The determination unit 54, the display control unit 55, and the edit receiving unit 56 perform the same processes as in the first embodiment.

[0071] Next, a processing procedure of the text generation system 80 in this embodiment will be described. The processing procedure of the text generation system 80 in this embodiment will be described with reference to FIG.

[0072] In this embodiment, an example will be described in which a report sentence relating to an image interpretation report of primary lung cancer is generated from parameters specified by a user. However, this embodiment can also be applied to generating report sentences other than those for lung cancer and sentences other than image interpretation reports.

[0073] (Step S901: Obtaining input candidate information) In step S901, the input candidate information acquisition unit 81 acquires parameters entered by the user via the operation unit 35 and stores them in the RAM 33. Examples of parameters entered by the user are shown in FIG. 10. FIG. 10(a) shows a dialog (an example of an input template) for the user to enter parameters, presenting options for determining the status of metastasis to the primary lung tumor. This dialog is displayed by the display unit 36 ​​under the control of the display control unit 55. The sentence generation system 80 also acquires user input to the dialog via the operation unit 35. In this embodiment, for simplicity's sake, parameters for the primary tumor site, primary tumor size, and metastasis site can be entered. In practice, other parameters necessary for diagnosis may also be entered, and the parameters that can be entered may be narrowed down to avoid complicating the dialog. FIG. 10(b) shows an example of parameters actually entered by the user. In this embodiment, the user determines the primary tumor site to be the "left hilum," and the user can specify the location on the dialog where "left hilum" is written by operating the operation unit 35. The display control unit 55 displays a mark on the dialog to visually identify the specified location. The user also determines the size of the primary tumor to be "3.1 cm," and the display control unit 55 similarly acquires this through user operation via the operation unit 35 and displays the result as a numerical value on the dialog. Furthermore, the user indicates the "right hilar lymph node" as a site with a high probability of metastasis, and indicates multiple sites with a low probability of metastasis. These can also be designated by the user operating the operation unit 35. In this embodiment, the input candidate information acquisition unit 81 acquires the radiology report acquired as described above as input candidate information.

[0074] (Step S902: Generate input information) In step S102, the generation unit 82 generates a plurality of pieces of input information that can be inferred by the language model 22 from the parameters acquired by the input candidate acquisition unit 51. An example of the input information generated by the generation unit 82 is shown in FIG. 11. FIG. 11(a) shows input information in which the input candidate information acquired in step S901 is directly converted into text. However, the sentence "Please create a sentence to be included in the radiology report" is added as an inference condition when the language model 22 makes an inference in step S904 (to be described later). Here, the primary source information "a 3.1 cm tumor in the left hilum" and the metastasis destination information "left and right bronchial lymph nodes, left hilar lymph nodes, left and right mediastinal lymph nodes, right hilum" can be generated by a known technique, such as by inserting the input candidate information acquired in step S101 into a template sentence.

[0075] The generation unit 82 further modifies some of the input candidate information to generate input information different from that shown in FIG. 11(a). That is, the generation unit 82 generates first input information including input candidate information selected by the user to be used as an input template, and second input information including other input candidate information selected by the user. FIG. 11(b) shows multiple pieces of input information generated by the generation unit 82 using sizes obtained by increasing or decreasing the tumor size of 3.1 mm obtained as input candidate information in 0.1 mm increments. In this manner, the generation unit 52 generates multiple different pieces of input information from the acquired parameters. The reason for slightly varying the input candidate information provided by the user in this manner is that it implicitly assumes that the input candidate information provided by the user is unstable or ambiguous. In other words, the tumor size "3.1 mm" entered by the user in this embodiment may vary slightly if entered by a different user due to differences in measurement methods and rounding methods to two decimal places between users. Such slight variations may affect the results of inference by the language model 22. More specifically, if the user had entered 3.2 mm as the tumor size, an important result would have been obtained as an inference result from the language model, but because 3.1 mm was entered, it may not have been obtained. Therefore, in this embodiment, inference is performed using the language model even when a small variation is intentionally given to the input candidate information entered by the user, and the results are integrated using the method described below.

[0076] In the above description, a case where multiple pieces of input information are generated with slight changes to tumor size has been described as an example, but the present invention is not limited to this. For example, slight changes may be made to the location of the primary site or the location of the metastatic site (e.g., changing to another adjacent site), or the information on the likelihood of metastasis may be changed (e.g., changing "suspected" to "strongly suspected" or "possible").

[0077] (Step S903: Input information to the language model) In step S103, the text information acquisition unit 83 inputs each of the multiple pieces of input information generated by the generation unit 82 to the language model 22. As a result, the language model 22 executes inference for each piece of input information.

[0078] (Step S904: Obtaining text information) In step S104, the text information acquisition unit 83 acquires the inference results performed by the language model 22 in step S903. Because the language model performs inference based on multiple different pieces of input information, the text information received by the text information acquisition unit 83 may not be identical. Here, FIG. 12 shows an example of the content of the text information received by the text information acquisition unit 83 from the language model 22. To simplify the explanation of this embodiment, unlike the example of text information in FIG. 4, FIG. 12 shows the content of the multiple pieces of text information acquired from the language model 22 in a unified form for comparison. FIG. 12(a) shows the content obtained by inference based on the input information in FIG. 11(a), and includes information on the size determination of the primary tumor (T2 / 5 cm or less) and the extent of metastasis (N3 / contralateral hilar lymph node enlargement). On the other hand, FIG. 12(b) shows the information obtained by inference based on the input information in FIG. 11(b). Unlike FIG. 12(a), there is a difference in the size determination of the primary tumor. In step S903, the number of times that each piece of input information is inferred may be one for each piece of input information, or multiple times as in the first embodiment. Even if only one piece of input information is inferred, multiple inferences are performed using similar input information, which makes it possible to reduce the instability of the inference to a certain extent while keeping down the calculation cost. Furthermore, by inferring each piece of input information multiple times, more stable report sentence candidates can be obtained, as in the first embodiment.

[0079] (Steps S905 to S906: same as in the first embodiment) In steps S905 to S906, as in the first embodiment, the determination unit 54 integrates the text information acquired by the text information acquisition unit 53, and the edit reception unit 56 displays the integrated report text candidate by the determination unit 54 on the display unit 36. In FIG. 12, different pieces of information, "T2 / 5 cm or less" and "T1c / 3 cm or less," are obtained, but the former appears three times while the latter appears twice, so "T2 / 5 cm or less" is adopted by majority vote or the like to generate the final report text. FIG. 13 shows an example in which a final report text candidate 1112 is displayed on an interpretation report 1111.

[0080] According to this embodiment, by having the language model 22 perform inference based on multiple types of input information and integrating the results, it is possible to provide the user with stable report sentence candidates that are less susceptible to the instability and ambiguity of the user input.

[0081] In addition to the method of generating input information to the language model 22 by applying slight variations to the input candidate information as exemplified above, the description of the inference condition assigned by the generation unit may also be changed. By generating multiple input sentences by replacing the inference condition "Please write a sentence to be written in the radiology report" shown in FIG. 11 with "Please write a report sentence to be written in the findings section," etc., it is possible to obtain multiple sentence information that has changed due to the instability of the input. By integrating the sentence information obtained in this way, it becomes possible to present the user with report sentence candidates that are less affected by the instability of the user's input.

[0082] (Variation 2-1) In the second embodiment, an example was shown in which multiple pieces of input information were generated in step S902. As with the first embodiment, this method has the disadvantage that information that the determination unit 54 has rejected is hidden and cannot be used by the user. In this modified example, as with modified example 1-1, by presenting the report sentence candidates generated by the determination unit 54 to the user along with the sentence information before integration, highly accurate report sentence candidates are presented, and an example process is shown for assisting the user in making corrections.

[0083] FIG. 14 shows an example of a display of report sentence candidates displayed on the display unit 36 ​​of the display control unit 55 constituting the sentence generation system 80 in this modified example. Unlike FIG. 13 , in FIG. 14(a), the display control unit 55 also displays pre-integration sentence information 1213 in the findings field 1112. The display control unit 55 also grays out the characters included in the integrated report sentence candidates, allowing the user to easily identify the keywords excluded from integration. The display control unit 55 may also display characters not included in the report sentence candidates in radio button format, as shown by 1214 in FIG. 14(b). The display control unit 55 may also display, in a different display format from FIGS. 14(a) and 14(b), differences in parameters used when each piece of sentence information that forms the report sentence candidates was acquired, as shown by 1215 in FIG. 14(c). This allows the user to easily infer the reason why a character in each piece of sentence information was not included in the report sentence candidate.

[0084] According to this modified example, by presenting the user with sentence information before it is integrated into a report sentence candidate, it is possible to provide the user with additional auxiliary information for correcting the report sentence candidate.

[0085] (Variation 2-2) The second embodiment and the modified example described above provide a method for providing a mechanism for allowing a user to modify a report sentence candidate. As mentioned in the explanation of Modifications 1-2, this method may require the user to make similar modifications frequently. This modified example provides a function for modifying a report sentence candidate by having the edit acceptance unit 56 reflect the edits made by the user to the report sentence candidate when the next report sentence is generated, and also illustrates a process for preventing the user from having to repeatedly make the same modifications. While Modifications 1-2 input the user's modifications to the determination unit 54 and use them in the next process, this modified example inputs them to the generation unit 82 and uses them in the next process.

[0086] 13 and 14, the edit receiving unit 56 transmits the edited content to the generation unit 82. The generation unit 82 incorporates the received edited content as a condition for generating input information from the next time onward, and uses it to generate input information in the next step S902.

[0087] For example, if a user changes the description of T2~ to T1c~ based on the display shown in FIG. 14(b), the edit receiving unit 56 transmits information regarding the difference in parameter conditions when a difference between T2~ and T1c~ occurs to the generation unit 82. Specifically, the edit receiving unit 56 transmits information regarding the parameter "nuclear power plant size 3.1 cm" input by the user and the parameter "nuclear power plant size 2.9 to 3.0 cm" when the inference result of T1c~ is obtained to the generation unit 82. The generation unit 52 uses this information to calculate a correction value for the nuclear power plant size required for the next input information generation unit, "-0.2 to -0.1 cm," and adds to the next input information an inference condition stating, "However, the nuclear power plant size shall be corrected by -0.2 to -0.1 cm from the description below." By automatically correcting the nuclear power plant size in the next and subsequent input information in this way, it becomes possible to obtain sentence information with the corrected nuclear power plant size from the language model 22.

[0088] According to this modification, subsequent report sentence candidates can be automatically made closer to the results of the previous report sentence correction by reflecting the edits made by the user in the generation unit 82. This brings the effect that subsequent report sentence candidates will be closer to the user's expectations, and reduces the editing workload on the user.

[0089] The user's corrections may be input to the input candidate information acquisition unit 81 rather than the generation unit 82. As shown in FIG. 10(a), there is a limit to the amount of information that can be displayed on the screen, and it is not possible to display selection items that cover all possibilities. Therefore, the input candidate information acquisition unit 81 can dynamically change the dialog displayed to the user in step S901 by adding selection items necessary to reproduce the user's corrections or deleting selection items unnecessary for reproducing the corrections. Specifically, if the user adds a description of abdominal organs to the report text as metastasis destination information, the input candidate information acquisition unit 81 adds items for various abdominal organs to the list of tumor confirmation sites from the next time. If the user adds a description of various blood test indicators to the report text, the input candidate information acquisition unit 81 adds a list of options for blood test results from the next time. If the user deletes information about the foci size from the report text, the input candidate information acquisition unit 81 deletes the input field for the foci size from the next time. The above methods are conceivable. This allows report text candidates from the next time onward to meet the user's expectations, thereby reducing the user's editing workload.

[0090] In the above example, information about parameters used to generate input information in the input candidate information acquisition unit 81 has been described as an example, but the present invention is not limited to this. For example, the report text edited by the user may be recorded, and in the next process, the generation unit 82 may add a description such as "Generate a report text similar to the report text below: ~ (hereinafter, the report text) ~" when generating input information. This makes it possible to easily build a system that prevents the user from having to make similar corrections repeatedly.

[0091] (Variation 2-3) In the second embodiment described above, the generation unit 82 acquires parameters selected by the user from options displayed on a dialog and generates multiple pieces of input information by varying the parameters. However, the present invention is not limited to this. For example, the generation unit 52 may analyze free-text input candidate information as shown in the first embodiment, extract keywords or numerical values ​​that serve as parameters, and vary the parameters in the corresponding locations. By automatically detecting parameters and varying them in this way, it is no longer necessary to create a dialog for the user to select parameters, and the user can freely input information without having to worry about how to use the dialog.

[0092] Furthermore, for items that remain ambiguous during automatic analysis, the parameter variation range may be increased. By changing the variation range in this way, it is possible to generate input information that reflects the instability and ambiguity of the user's input.

[0093] The above is one example of an embodiment, but the present invention is not limited to the embodiment described above and shown in the drawings, and can be practiced by making appropriate modifications within the scope that does not change the gist of the present invention.

[0094] <Other embodiments> Furthermore, the disclosed technology can be embodied as, for example, a system, a device, a method, a program, or a recording medium (storage medium), etc. Specifically, it may be applied to a system consisting of multiple devices (for example, a host computer, an interface device, an imaging device, a web application, etc.), or it may be applied to an apparatus consisting of a single device.

[0095] Needless to say, the object of the present invention can be achieved by the following: A recording medium (or storage medium) on which software program code (computer program) that realizes the functions of the above-described embodiments is recorded is supplied to a system or device. The storage medium is, of course, a computer-readable storage medium. The computer (or CPU or MPU) of the system or device then reads and executes the program code stored on the recording medium. In this case, the program code itself read from the recording medium realizes the functions of the above-described embodiments, and the recording medium on which the program code is recorded constitutes the present invention. [Explanation of symbols]

[0096] 10 Text Generation System 21 Network 22 Language Models 31 Communication Interface 32 ROM 33 RAM 34 Storage section 35 Control section 36 Display section 37 Control Unit 51 Input candidate information acquisition section 52 Generation part 53 Text information acquisition section 54 Decision Section 55 Display control unit 56 Editorial Reception Department 80 Text Generation System 81 Input candidate information acquisition section 82 Generation part 83 Text information acquisition section

Claims

1. an input candidate information acquisition unit that acquires input candidate information to be input from input information to a language model that outputs sentence information; a text information acquisition unit that acquires a plurality of pieces of text information that are different from each other by inputting input information based on the input candidate information into the language model; a determination unit that determines report sentence candidates using the plurality of pieces of text information; a display control unit that displays the report sentence candidates determined by the determination unit on a display unit; and The display control unit is characterized in that it displays in an identifiable manner the adopted information that has been adopted by the determination unit from the text information as information to be used for the report sentence candidates and the unadopted information that has not been adopted.

2. a generation unit that generates the input information to be input to the language model based on the input candidate information; 2. The text generation system according to claim 1.

3. the generation unit generates a plurality of pieces of input information that are different from each other based on the input candidate information; The text information acquisition unit acquires the plurality of pieces of text information by inputting the plurality of pieces of input information into the language model.

3. The sentence generation system according to claim 2, wherein:

4. the input candidate information acquisition unit acquires a user's selection of an input template as input candidate information; The generation unit generates first input information including input candidate information selected by the user, and second input information including other input candidate information related to the input candidate information selected by the user.

4. The sentence generation system according to claim 3,

5. The determination unit determines the report sentence candidates by further inputting the plurality of pieces of sentence information into the language model.

4. The text generation system according to claim 1, wherein:

6. The report sentence candidate further includes an edit receiving unit that receives edits from a user to the report sentence candidate.

4. The text generation system according to claim 1, wherein:

7. An input candidate information acquisition unit that acquires input candidate information to be input from input information to a language model that outputs sentence information; a text information acquisition unit that acquires a plurality of pieces of text information that are different from each other by inputting input information based on the input candidate information into the language model; a determination unit that determines report sentence candidates using the plurality of pieces of text information; a display control unit that displays the report sentence candidates determined by the determination unit on a display unit; an edit receiving unit that receives edits by a user to the report sentence candidates; and The display control unit displays the report sentence candidate in association with the presence or absence of non-adopted information that was not adopted by the determination unit from the text information as information to be used for the report sentence candidate.

8. the editing reception unit receives editing information by the user for the non-recruitment information; 7. The writing generation system according to claim 6, wherein the report sentence candidates are updated by using the selected keyword as employment information for the non-employment information.

9. An input candidate information acquisition unit that acquires input candidate information to be input from input information to a language model that outputs sentence information; a generation unit that generates the input information to be input to the language model based on the input candidate information; a text information acquisition unit that acquires a plurality of pieces of text information that are different from each other by inputting the input information generated by the generation unit into the language model; a determination unit that determines report sentence candidates using the plurality of pieces of text information; an edit receiving unit that receives edits by a user to the report sentence candidates; and the generation unit updates rules for generating input information based on the editing information of the user accepted by the editing acceptance unit; Generate the input information using the updated rules. A text generation system characterized by:

10. An input candidate information acquisition unit that acquires input candidate information to be input from input information to a language model that outputs sentence information; a text information acquisition unit that acquires a plurality of pieces of text information that are different from each other by inputting input information based on the input candidate information into the language model; a determination unit that determines report sentence candidates using the plurality of pieces of text information; a display control unit that displays the report sentence candidates determined by the determination unit on a display unit; an edit receiving unit that receives edits by a user to the report sentence candidates; and the determination unit updates a rule for determining the report sentence candidates based on the editing information of the user received by the editing receiving unit; The report sentence candidates are determined based on the updated rules. A text generation system characterized by:

11. 11. The writing generation system according to claim 10, wherein the determination unit changes the priority of keywords to be adopted as report sentence candidates based on the editing information of the user.

12. the language model is composed of a plurality of language models that are different from each other, The text information acquisition unit acquires the plurality of pieces of text information by inputting input information into the plurality of language models.

4. The text generation system according to claim 1, wherein:

13. an input candidate information acquisition unit that acquires input candidate information to be input from input information to a language model that outputs sentence information; a text information acquisition unit that acquires a plurality of pieces of text information that are different from each other by inputting input information based on the input candidate information into the language model; a determination unit that determines sentence candidates using the plurality of sentence information; a display control unit that displays the sentence candidates determined by the determination unit on a display unit; and The display control unit is characterized in that it displays in an identifiable manner the adopted information that has been adopted by the determination unit from the sentence information as information to be used for the sentence candidate and the non-adopted information that has not been adopted.

14. A computer-implemented method for generating text, comprising: an input candidate information acquisition step for acquiring input candidate information to be input from the input information to a language model that outputs sentence information; a sentence information acquisition step of acquiring a plurality of sentence information different from each other by inputting input information based on the input candidate information into the language model; a determining step of determining report sentence candidates using the plurality of pieces of text information; a display step of displaying the report sentence candidates determined in the determination step on a display unit, The display step is a text generation method characterized in that adopted information that was adopted from the text information in the determination step as information to be used for the report sentence candidate and non-adopted information that was not adopted are displayed in an identifiable manner.

15. A program for executing the sentence generation method according to claim 14 on a computer.

Citation Information

Patent Citations

  • Question and answer processing method, device and equipment based on artificial intelligence and storage medium

    CN117149982A

  • Information processing program, information processing method and information processing device

    JP2022185799A