Learning device, learning method, trained model, and program

By using three-dimensional X-ray CT images to create pseudo-three-dimensional plain X-ray images and correcting errors, the learning device enhances the accuracy of radiological interpretation reports, addressing the limitations of two-dimensional models.

JP7776455B2Active Publication Date: 2025-11-26FUJIFILM CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2022578243
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-01-26
Filing Date
2022-01-17
Publication Date
2025-11-26
Estimated Expiration
2042-01-17

AI Technical Summary

Technical Problem

Existing machine learning models trained on two-dimensional plain X-ray images struggle with low accuracy due to the challenge of overlapping organs and loss of three-dimensional shape information, leading to low-quality interpretation reports.

Method used

A learning device and method that utilizes three-dimensional X-ray CT images to generate pseudo-three-dimensional plain X-ray images and corresponding interpretation reports, incorporating error correction and conversion processes to enhance the accuracy of interpretation models.

Benefits of technology

The approach enables the generation of highly accurate interpretation reports by leveraging three-dimensional information, improving the precision of radiological assessments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007776455000001
    Figure 0007776455000001
  • Figure 0007776455000002
    Figure 0007776455000002
  • Figure 0007776455000003
    Figure 0007776455000003
Patent Text Reader

Abstract

Provided are: a learning device that uses high-accuracy, high-quality learning data to generate a learned model which outputs a high-accuracy diagnostic interpretation report; a learning method; a program; and a learned model which has undergone learning using said learning method. A learning device comprises: a processor (129); a memory (114); and a learning model (130). The processor (129) executes: a process for projecting an X-ray CT image (202) to generate a pseudo pure X-ray image (204), and inputting the pseudo pure X-ray image (204) into a learning model (126); a process for converting a first diagnostic interpretation report (206) to generate a second diagnostic interpretation report (208) for the pseudo pure X-ray image (204); a process for acquiring an error between the second diagnostic interpretation report (208) and an estimated report (210) for the pseudo pure X-ray image (204) that was output by the learning model (126) on the basis of the input pseudo pure X-ray image (204); and a process for making the learning model (126) learn using the error.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a learning device, a learning method, a trained model, and a program, and more particularly to a learning device, a learning method, a trained model, and a program that perform learning regarding the output of radiology reports. [Background technology]

[0002] Traditionally, doctors have interpreted plain X-ray images to identify diseases and compiled the results of the interpretation into an interpretation report. However, interpreting plain X-ray images is not easy even for doctors, and the accuracy of the interpretation report can be low. Here, a plain X-ray image is a two-dimensional image obtained by irradiating X-rays and projecting the resulting shadow onto a plane.

[0003] In recent years, machine learning techniques have been used to propose trained models that are trained to output interpretation reports for input plain X-ray images.

[0004] For example, Non-Patent Documents 1 and 2 describe machine learning techniques that input chest X-ray images (plain X-ray images) and output radiology reports. [Prior art documents] [Patent documents]

[0005] [Non-Patent Document 1] Yuan, Jianbo, et al., "Automatic radiology report generation based on multi-view image fusion and medical concept enrichment.", MICCAI, 2019. [Non-patent document 2] Li, Christy Y., et al. "Knowledge-driven encode, retrieve, paraphrase for medical image report generation.", AAAI, 2019. Summary of the Invention [Problem to be solved by the invention]

[0006] Here, the techniques described in Non-Patent Document 1 and Non-Patent Document 2 use plain X-ray images having two-dimensional information and their interpretation reports as training data. As described above, it is not easy even for doctors to create interpretation reports for plain X-ray images, and the accuracy of the interpretation reports may be low. One reason for this is that plain X-ray images depict organs, which originally have three-dimensional shapes, as two-dimensional images, which may result in organs being displayed overlapping with each other or making it difficult to grasp the original shapes of the organs. Therefore, a trained model trained using such low-accuracy interpretation reports may not be able to output highly accurate interpretation reports.

[0007] The present invention has been made in consideration of these circumstances, and its purpose is to provide a learning device, a learning method, a program, and a trained model trained using the learning method, which uses high-precision, high-quality training data to generate a trained model that outputs highly accurate radiological reports. [Means for solving the problem]

[0008] To achieve the above-mentioned object, one aspect of the present invention is a learning device comprising a processor, a memory for storing a learning dataset of an X-ray CT image having three-dimensional information and a first interpretation report for the X-ray CT image, and a learning model for generating an interpretation report from a plain X-ray image having two-dimensional information, wherein the processor performs the following processes: projecting the X-ray CT image to generate a pseudo-simple X-ray image and inputting the pseudo-simple X-ray image into the learning model; converting the first interpretation report to generate a second interpretation report for the pseudo-simple X-ray image; obtaining the error between the estimated report for the pseudo-simple X-ray image output by the learning model based on the input pseudo-simple X-ray image and the second interpretation report; and training the learning model using the error.

[0009] According to this aspect, a pseudo plain X-ray image and a second interpretation report for the pseudo plain X-ray are generated from a learning dataset of an X-ray CT image having three-dimensional information and a first interpretation report for the X-ray CT image, and learning is performed using the pseudo plain X-ray image and the second interpretation report. As a result, this aspect performs learning using the pseudo X-ray image and the second interpretation report based on the X-ray CT image and the first interpretation report, which have a large amount of information, and therefore can perform learning to output a highly accurate interpretation report.

[0010] Preferably, the process of generating the second interpretation report generates the second interpretation report from the first interpretation report by converting organ labels included in the first interpretation report into organ labels of the second interpretation report.

[0011] Preferably, the process of generating the second interpretation report generates the second interpretation report from the first interpretation report by converting a disease label included in the first interpretation report into a disease label of the second interpretation report.

[0012] Preferably, the process of generating the second interpretation report converts a first knowledge graph corresponding to the first interpretation report into a second knowledge graph corresponding to the second interpretation report, and generates the second interpretation report based on the conversion.

[0013] Preferably, the memory stores an X-ray CT image of a subject in a first position, and when the learning model generates an interpretation report from a plain X-ray image of a subject in a second position, the process of inputting the pseudo plain X-ray image generates a pseudo plain X-ray image of the second position from the X-ray CT image of the first position and inputs the pseudo plain X-ray image of the second position to the learning model.

[0014] Preferably, the process of inputting the pseudo plain X-ray image involves generating a pseudo plain X-ray image projected in a first direction from the X-ray CT image and a pseudo plain X-ray image projected in a second direction, and inputting the pseudo plain X-ray image projected in the first direction and the pseudo plain X-ray image projected in the second direction into the learning model.

[0015] Preferably, the memory stores an additional learning data set of plain X-ray images and disease labels of the plain X-ray images, and the process of obtaining an error obtains an error between an estimated report for the pseudo plain X-ray image output by the learning model with reference to the disease label and the second interpretation report.

[0016] Preferably, the memory stores an additional learning data set of a plain X-ray image and a third interpretation report for the plain X-ray image, and the process of obtaining the error obtains the error between the estimated report for the pseudo plain X-ray image output by the learning model based on the input pseudo plain X-ray image and the second interpretation report, and the error between the estimated report for the plain X-ray image output by the learning model based on the input plain X-ray image and the third interpretation report.

[0017] Another aspect of the present invention is a learning method in which a processor uses a learning dataset of X-ray CT images having three-dimensional information and a first interpretation report for the X-ray CT image stored in a memory to train a learning model that generates an interpretation report from a plain X-ray image having two-dimensional information, and includes the steps of projecting the X-ray CT image to generate a pseudo-plain X-ray image and inputting the pseudo-plain X-ray image into the learning model, converting the first interpretation report to generate a second interpretation report for the pseudo-plain X-ray image, obtaining the error between the estimated report for the pseudo-plain X-ray image output by the learning model based on the input pseudo-plain X-ray image and the second interpretation report, and training the learning model using the error.

[0018] Preferably, the step of generating the second interpretation report generates the second interpretation report from the first interpretation report by converting organ labels included in the first interpretation report into organ labels of the second interpretation report.

[0019] Preferably, the step of generating the second interpretation report generates the second interpretation report from the first interpretation report by converting a disease label included in the first interpretation report into a disease label of the second interpretation report.

[0020] Preferably, the step of generating the second interpretation report includes converting a first knowledge graph corresponding to the first interpretation report into a second knowledge graph corresponding to the second interpretation report, and generating the second interpretation report based on the conversion.

[0021] A learning program according to another aspect of the present invention causes a processor to execute the processing of each step in the above-described learning method.

[0022] The trained model, which is another aspect of the present invention, is trained using the above-mentioned training method. [Effects of the Invention]

[0023] According to the present invention, a pseudo-simple X-ray image and a second interpretation report for the pseudo-simple X-ray are generated from a learning dataset of an X-ray CT image having three-dimensional information and a first interpretation report for the X-ray CT image, and learning is performed using this pseudo-simple X-ray image and the second interpretation report.Therefore, learning is performed using a pseudo-X-ray image and a second interpretation report based on an X-ray CT image with a large amount of information and the first interpretation report, and learning can be performed to output an interpretation report with high accuracy. [Brief explanation of the drawings]

[0024] [Figure 1] FIG. 1 is a block diagram showing an embodiment of the hardware configuration of a learning device. [Figure 2] FIG. 2 is a block diagram illustrating the main functions of the learning device. [Figure 3] FIG. 3 is a diagram illustrating an X-ray CT image and a first radiological report, which are examples of the learning data set. [Figure 4] FIG. 4 is a diagram illustrating the pseudo image generating unit. [Figure 5] FIG. 5 is a diagram illustrating the report generating unit. [Figure 6] FIG. 6 is a diagram showing an example of an organ label conversion list provided in the report generating unit. [Figure 7] FIG. 7 is a diagram illustrating the correspondence between the three-dimensional organ labels and the two-dimensional organ labels. [Figure 8] FIG. 8 is a diagram illustrating the disease label conversion list. [Figure 9] FIG. 9 is a diagram illustrating the conversion from the first report to the second report by the report generating unit. [Figure 10] FIG. 10 is a functional block diagram illustrating the learning model, the error acquisition unit, and the learning control unit. [Figure 11] FIG. 11 is a diagram illustrating a learning method using a learning device and each step executed by a processor according to a program. [Figure 12]FIG. 12 is a diagram illustrating a position conversion unit that converts an X-ray CT image in a supine position into an X-ray CT image in an upright position. [Figure 13] FIG. 13 is a diagram illustrating that the pseudo image generating unit generates pseudo X-ray images in two directions. [Figure 14] FIG. 14 is a diagram for explaining an example of conversion of an anatomical knowledge graph provided in the report generating unit. [Figure 15] FIG. 15 is a diagram conceptually showing an anatomical knowledge graph for an X-ray CT image. [Figure 16] FIG. 16 is a diagram conceptually showing an anatomical knowledge graph for an X-ray CT image. [Figure 17] FIG. 17 is a diagram conceptually showing an anatomical knowledge graph for a plain X-ray image. [Figure 18] FIG. 18 is a diagram showing an example of conversion of a disease knowledge graph provided in the report generating unit. [Figure 19] FIG. 19 is a diagram illustrating conversion from a first report to a second report by the report generating unit equipped with the anatomical knowledge graph and the disease knowledge graph. [Figure 20] FIG. 20 is a diagram illustrating an additional training data set. [Figure 21] FIG. 21 is a diagram illustrating the learning of the learning model. [Figure 22] FIG. 22 is a diagram illustrating an additional training data set. [Figure 23] FIG. 23 is a diagram illustrating the learning of the learning model. DETAILED DESCRIPTION OF THE INVENTION

[0025] Hereinafter, preferred embodiments of a learning device, a learning method, a trained model, and a program according to the present invention will be described with reference to the accompanying drawings.

[0026] FIG. 1 is a block diagram showing an embodiment of the hardware configuration of a learning device.

[0027] The learning device 100 shown in FIG. 1 is configured as a computer. The computer may be a personal computer, a workstation, or a server computer. The learning device 100 includes a communication unit 112, a memory (storage unit) 114, a learning model 126, an operation unit 116, a central processing unit (CPU) 118, a graphics processing unit (GPU) 119, a random access memory (RAM) 120, a read-only memory (ROM) 122, and a display unit 124. The CPU 118 and the GPU 119 constitute a processor 129. The GPU 119 may be omitted from the processor 129.

[0028] The communication unit 112 is an interface that performs communication processing with an external device via wire or wirelessly, and exchanges information with the external device.

[0029] The memory 114 includes a storage device configured using, for example, a hard disk drive, an optical disk, a magneto-optical disk, or a semiconductor memory, or an appropriate combination of these. The memory 114 stores various programs and data required for image processing such as learning processing and / or image generation processing. The programs stored in the memory 114 are loaded into the RAM 120 and executed by the processor 129, causing the computer to function as a means for performing various processes defined by the programs. The memory also stores a learning dataset, which will be described below.

[0030] Operation unit 116 is an input interface that accepts various operational inputs to study device 100. Operation unit 116 may be, for example, a keyboard, a mouse, a touch panel, operation buttons, or a voice input device, or an appropriate combination of these.

[0031] The processor 129 reads out various programs stored in the ROM 122 or the memory 114, etc., and executes various processes. The RAM 120 is used as a working area for the processor 129. The RAM 120 is also used as a storage unit that temporarily stores the read programs and various data.

[0032] The display unit 124 is an output interface that displays various types of information. The display unit 124 may be, for example, a liquid crystal display, an organic electro-luminescence (OEL) display, a projector, or an appropriate combination of these.

[0033] The learning model 126 is configured by a CNN (Convolutional Neural Network). As will be described later, a pseudo plain X-ray image generated from an X-ray CT image is input to the learning model 126, and an interpretation report is generated based on the input pseudo plain X-ray image. The learning model 126 in the learning device 100 is an untrained model, and the learning device 100 according to the present invention trains the learning model 126 by machine learning.

[0034] First Embodiment The first embodiment will be described below. The following description will be directed to learning of a learning model that generates pseudo plain X-ray images from X-ray CT images having three-dimensional information of a chest image and outputs an interpretation report of the pseudo plain X-ray images.

[0035] FIG. 2 is a block diagram illustrating the main functions of the learning device 100 of this embodiment.

[0036] The learning device 100 mainly comprises a memory 114, a processor 129, and a learning model 126 (see FIG. 1). The processor 129 realizes the functions of a learning data acquisition unit 130, a pseudo image generation unit 132, a report generation unit 134, an error acquisition unit 136, and a learning control unit 138.

[0037] The learning data acquisition unit 130 acquires a learning data set to be used for learning stored in the memory 114. For example, the learning data set is composed of an X-ray CT image of a patient's chest and a first interpretation report for the X-ray image. The first interpretation report is a report created by a doctor or the like by interpreting the X-ray CT image.

[0038] FIG. 3 is a diagram illustrating an X-ray CT image, which is an example of the learning data set, and a first radiology report 206. In FIG.

[0039] The training data set 200 is composed of a set of X-ray CT images 202 and a first radiological report 206. The memory 114 stores a plurality of training data sets 200, and the training model 126 is trained using these plurality of training data sets 200.

[0040] The X-ray CT image 202 is obtained by actually imaging a patient, who is the subject. The X-ray CT image 202 has three-dimensional information (three-dimensional spatial information). Therefore, when generating an interpretation report (first interpretation report 206) based on the X-ray CT image 202, a doctor can observe organs and the like using the three-dimensional information. Therefore, compared to creating an interpretation report based on a plain X-ray image having two-dimensional information, a doctor can create a more detailed and accurate interpretation report based on the X-ray CT image 202 having three-dimensional information. In the X-ray CT image 202, cross sections 600S, 600C, and 600A are cross sections in the sagittal direction, coronal direction, and axial direction, respectively. The illustrated X-ray CT image 202 of the chest is an example of an X-ray CT image, and X-ray CT images of other parts of the body can also be used in this embodiment.

[0041] The first interpretation report 206 includes information interpreted from the X-ray CT image 202. The first interpretation report 206 includes anatomical structure information that can be interpreted from the X-ray CT image 202. Because the X-ray CT image 202 includes three-dimensional information, a doctor can observe, for example, the lungs by dividing them into smaller sections. Therefore, the first interpretation report 206 includes a statement that "an irregular solid mass is observed in the right segments S4 and S5." The first interpretation report 206 also includes a disease label that can be interpreted from the X-ray CT image 202. Because the X-ray CT image 202 includes three-dimensional information, a doctor can observe, for example, the shape of the margin in more detail. Therefore, the first interpretation report 206 includes a statement that "the margin is saw-toothed and accompanied by spicules, and pleural indentation is also observed."

[0042] The training data acquisition unit 130 acquires the training data set 200 from the memory 114, sends the X-ray CT image 202 to the pseudo image generation unit 132, and sends the first radiology report 206 to the report generation unit 134.

[0043] FIG. 4 is a diagram illustrating the pseudo image generating unit 132. As shown in FIG.

[0044] The pseudo image generator 132 generates a pseudo plain X-ray image 204 having two-dimensional information from an input X-ray CT image 202 having three-dimensional information. The pseudo image generator 132 can generate the pseudo plain X-ray image 204 from the X-ray CT image 202 by various methods. For example, the pseudo image generator 132 generates the pseudo plain X-ray image 204 from the X-ray CT image 202 by the post-digitally reconstructed radiograph (DRR) method described in the literature (A method to produce and validate a digitally reconstructed radiograph-based computer simulation for optimisation of chest radiographs acquired with a computed radiography imaging system, C.S. MOORE, The British Journal of Radiology, 84 (2011), 890-902).

[0045] FIG. 5 is a diagram illustrating the report generating unit 134. As shown in FIG.

[0046] The report generation unit 134 generates a second interpretation report 208 based on the input first interpretation report 206. The report generation unit 134 can generate the second interpretation report 208 from the first interpretation report 206 by various methods. For example, the report generation unit 134 is provided with a conversion list, and generates the second interpretation report 208 by converting the wording written in the first interpretation report 206 based on the conversion list. Specifically, the report generation unit 134 is provided with an organ label conversion list 205A (FIG. 6), and generates the second interpretation report from the first interpretation report by converting the organ labels used in the first interpretation report 206 into organ labels of the second interpretation report 208. The report generation unit 134 also has a disease label conversion list 205B (FIG. 8) and generates a second radiology report from the first radiology report 206 by converting the disease labels used in the first radiology report 206 into disease labels of the second radiology report 208. Note that the organ label conversion list 205A and the disease label conversion list 205B are specific examples, and the report generation unit 134 may have other conversion lists and generate the second radiology report 208 from the first radiology report 206 using those conversion lists.

[0047] Fig. 6 is a diagram showing an example of the organ label conversion list 205A provided in the report generating unit 134. Note that Fig. 6 shows the organ label conversion list for the right lung, and the organ label conversion list for the left lung is omitted from the illustration.

[0048] As shown in the organ label conversion list 205A, each of the three-dimensional organ labels for the right lung is converted into a two-dimensional organ label. Specifically, right lung sections S1 to S3 in the three-dimensional organ labels become upper right lung section T1 in the two-dimensional organ labels. Right sections S4 to S6 become lower right lung section T3 in the two-dimensional organ labels. Sections S7 to S10 become middle right lung section T2 in the two-dimensional organ labels. Here, the three-dimensional organ labels are divided into relatively fine sections based on the X-ray CT image 202 having three-dimensional information. On the other hand, the two-dimensional organ labels correspond to plain X-ray images having two-dimensional information and are divided into relatively rough sections. The correspondence between the three-dimensional organ labels and the two-dimensional organ labels will be explained below.

[0049] FIG. 7 is a diagram illustrating the correspondence between the three-dimensional organ labels and the two-dimensional organ labels.

[0050] Organ labels 220 are assigned based on anatomical structure information obtained from the X-ray CT image 202. Since the X-ray CT image 202 contains three-dimensional information about the organs, labels are assigned to ten sections (sections S1 to S10) for each of the left and right lungs as shown in the figure. Since the X-ray CT image 202 contains three-dimensional information about the lungs, it is possible to observe the front and back sides of the lungs, and therefore the lungs can be divided into smaller sections and labeled.

[0051] On the other hand, since plain X-ray images have two-dimensional information, organ labels 222 are attached. As shown in the figure, plain X-ray images have labels for each of the left and right lungs, dividing them into three regions (upper lung T1, middle lung T2, and lower lung T3). Since plain X-ray images do not have three-dimensional information about the lungs and therefore do not allow observation of the front and back sides of the lungs, the lungs can be divided into three regions and labeled. Note that the method of defining lung regions in the X-ray CT image 202 and plain X-ray images described above is merely an example, and lung regions may be defined in other ways. In this way, the report generation unit 134 generates the second interpretation report 208 from the first interpretation report 206 by using the organ label conversion list 205A.

[0052] FIG. 8 is a diagram illustrating the disease label conversion list 205B provided in the report generating unit 134. As shown in FIG.

[0053] As shown in the illustrated disease label conversion list 205B, each 3D disease label is converted into a 2D disease label. Specifically, the 3D disease labels spicules, serrated, and lobulated are converted into irregular 2D disease labels. The 3D disease label calcification is converted into "XX" in the 2D disease label. The 3D disease label cavity is converted into "XX" in the 2D disease label. The 3D disease labels are assigned relatively detailed disease labels based on the X-ray CT image 202 containing 3D information. On the other hand, the 2D disease labels correspond to plain X-ray images containing 2D information, and relatively rough disease labels are assigned. Note that the lung disease labels in the X-ray CT image 202 and plain X-ray images described above are merely examples, and lung disease labels may be assigned in other forms. In this way, the report generation unit 134 generates the second radiology report 208 from the first radiology report 206 using the disease label conversion list 205B.

[0054] FIG. 9 is a diagram illustrating conversion from the first report to the second report by the report generating unit 134 including the organ label conversion list 205A and the disease label conversion list 205B described above.

[0055] As shown in the figure, the report generation unit 134 converts "right segments S4 and S5" in the first interpretation report 206 to "lower right lung" based on the organ label conversion list 205A, and generates a second interpretation report 208. Also, the report generation unit 134 converts "serrated and spicules" in the first interpretation report 206 to "irregular" based on the disease label conversion list 205B, and generates the second interpretation report 208.

[0056] As described above, the report generation unit 134 has a conversion list and generates the second interpretation report 208 from the first interpretation report 206 based on the conversion list. Note that although the above describes an example in which the report generation unit 134 generates the second interpretation report 208 from the first interpretation report 206 using the conversion list, this aspect is not limited to this. For example, the report generation unit 134 may be configured with a trained model and generate the second interpretation report 208 from the first interpretation report 206.

[0057] FIG. 10 is a functional block diagram illustrating the learning model 126, the error acquisition unit 136, and the learning control unit 138.

[0058] The learning model 126 is configured as a convolutional neural network (CNN), which is one of deep learning models.

[0059] The learning model 126 has a multi-layer structure and holds a multiplicity of weight parameters. The learning model 126 can change from an unlearned model to a trained model by updating the weight parameters from their initial values ​​to optimal values. The initial values ​​of the weight parameters of the learning model 126 may be any values, or, for example, weight parameters of a trained model that outputs a known radiology report may be applied.

[0060] This learning model 126 comprises an input layer 126A, an intermediate layer 126B having multiple sets of convolutional layers and pooling layers, and an output layer 126C, and each layer has a structure in which multiple "nodes" are connected by "edges."

[0061] The pseudo plain X-ray image 204 from the training data set 200 is input to the input layer 126A.

[0062] The intermediate layer 126B includes a convolutional layer, a pooling layer, and the like, and extracts features from the image input from the input layer 126A. The convolutional layer filters nearby nodes in the previous layer (performing a convolution operation using a filter) to obtain a "feature map." The pooling layer reduces the feature map output from the convolutional layer to create a new feature map. The "convolutional layer" is responsible for feature extraction, such as edge extraction from the image, and the "pooling layer" is responsible for providing robustness to the extracted features so that they are not affected by translation, etc. Note that the intermediate layer 126B is not limited to cases where convolutional layers and pooling layers are alternately arranged, but also includes cases where convolutional layers are consecutive and normalization layers. The final convolutional layer, conv, outputs a feature map indicating events to be interpreted from the pseudo plain X-ray image 204.

[0063] The output layer 126C is a part that outputs the output result of the learning model 126 (the estimation report 210).

[0064] The error acquisition unit 136 acquires the output result (estimated report 210) output from the output layer 126C of the learning model 126 and the second radiology report 208 corresponding to the pseudo plain X-ray image 204, and calculates the error between them. The error may be calculated using, for example, the Jaccard coefficient or the Dice coefficient.

[0065] Based on the error calculated by the error acquisition unit 136, the learning control unit 138 adjusts the weight parameters of the learning model 126 using the error backpropagation method to minimize the distance in the feature space between the second radiology report 208 and the output of the learning model 126 or to maximize the similarity.

[0066] This parameter adjustment process is repeated, and learning is repeated until the error calculated by the error acquisition unit 136 converges.

[0067] In this way, the training data set is used to create a trained learning model 126 with optimized weight parameters.

[0068] Next, a learning method using learning device 100 will be described.

[0069] FIG. 11 is a diagram illustrating a learning method using the learning device 100 and each step executed by a processor according to a learning program.

[0070] First, the learning data acquisition unit 130 acquires the learning dataset (X-ray CT image 202 and first radiology report 206) 200 stored in the memory 114 (step S10). Then, the X-ray CT image 202 is sent to the pseudo image generation unit 132, which generates a pseudo plain X-ray image 204 based on the X-ray CT image 202 (step S11). Next, the report generation unit 134 converts the organ label 220 of the first radiology report 206 based on the organ label conversion list 205A (step S12). The report generation unit 134 also converts the disease label of the first radiology report 206 based on the disease label conversion list (step S13). By converting the labels, the report generation unit 134 generates a second radiology report 208. Next, the learning model 126 outputs an estimated report 210 based on the input pseudo plain X-ray image 204 (step S14). Thereafter, the error acquisition unit 136 acquires the error between the estimated report 210 and the second radiology report 208 (step S15), and the learning control unit 138 trains the learning model 126 based on the acquired error (step S16).

[0071] As described above, according to this embodiment, a pseudo plain X-ray image 204 and a second interpretation report 208 for the pseudo plain X-ray image 204 are generated from a training data set 200 of an X-ray CT image 202 having three-dimensional information and a first interpretation report 206 for the X-ray CT image 202, and learning is performed using the pseudo plain X-ray image 204 and the second interpretation report 208. This allows this aspect to be trained to output a highly accurate interpretation report. Furthermore, according to a trained model trained using the training method of this embodiment, a plain X-ray image can be input and a highly accurate interpretation report for the input plain X-ray image can be output.

[0072] <Second embodiment> In the above example, the pseudo plain X-ray image 204 in a standing position is generated from the X-ray CT image 202 in a standing position. However, in this embodiment, even if the X-ray CT image 202 in a lying position (first position) is stored in the memory 114, the pseudo plain X-ray image 204 in a standing position (second position) can be generated and input to the learning model 126.

[0073] 12 is a diagram illustrating the position changing unit 150 that converts an X-ray CT image in a lying position into an X-ray CT image in an upright position. The position changing unit 150 is provided in the learning data acquiring unit 130, for example.

[0074] The position conversion unit 150 converts the X-ray CT image 202A in a lying position stored in the memory 114 into an X-ray CT image in an upright position. The position conversion unit 150 can convert the X-ray CT image 202A in a lying position into an X-ray CT image 202B in an upright position by various methods. For example, the position conversion unit 150 may be configured with a trained model that has undergone machine learning, and may output an X-ray CT image 202B in an upright position from the input X-ray CT image 202A in a lying position.

[0075] In this manner, in this embodiment, the X-ray CT image 202A in a supine position is converted into the X-ray CT image 202B in an upright position. Then, the pseudo image generation unit 132 generates the pseudo plain X-ray image 204 from the converted X-ray CT image 202B in an upright position. Therefore, even an X-ray CT image taken in a supine position can be appropriately used in this embodiment.

[0076] <Third embodiment> In the example described above, the estimated report 210 is generated based on the pseudo plain X-ray image 204 of an AP (anterior to posterior) view or a PA (posterior to anterior) view based on the X-ray CT image 202. However, in this embodiment, a pseudo X-ray image is generated from an image taken in another direction, for example, a lateral view, and the estimated report 210 is generated based on the pseudo X-ray image.

[0077] FIG. 13 is a diagram illustrating how the pseudo image generating unit 132 generates pseudo X-ray images in two directions.

[0078] The pseudo image generation unit 132 generates a pseudo plain X-ray image 204a projected in the AP direction (first direction) and a pseudo plain X-ray image 204b projected in the LAT (Lateral) direction (second direction) based on the X-ray CT image 202. The pseudo image generation unit 132 can generate the pseudo plain X-ray image 204a in the AP direction and the pseudo plain X-ray image 204b in the LAT direction using a known technique. For example, the pseudo image generation unit 132 generates the pseudo plain X-ray image 204a in the AP direction and the pseudo plain X-ray image 204b in the LAT direction using the DRR method described above.

[0079] In this manner, in this embodiment, a pseudo plain X-ray image 204a projected in the AP direction and a pseudo plain X-ray image 204b projected in the LT direction are generated based on the X-ray CT image 202. Then, the pseudo plain X-ray image 204a projected in the AP direction and the pseudo plain X-ray image 204b projected in the LAT direction are input to the learning model 126, so that learning is performed to output a more accurate interpretation report.

[0080] <Fourth embodiment> In the above example, the report generation unit 134 includes the organ label conversion list 205A and the disease label conversion list 205B. In this embodiment, the report generation unit 134 converts the knowledge graph and generates the second interpretation report 208 from the first interpretation report 206 based on the conversion. Specifically, the report generation unit 134 converts the first knowledge graph corresponding to the first interpretation report 206 into a second knowledge graph corresponding to the second interpretation report 208, and generates the estimated report 210 based on the conversion. For example, the report generation unit 134 includes an anatomical knowledge graph for X-ray CT images (first knowledge graph) and a disease knowledge graph for X-ray CT images (first knowledge graph), and converts each knowledge graph into an anatomical knowledge graph for plain X-ray images (second knowledge graph) and a disease knowledge graph for plain X-ray images (second knowledge graph). Then, the report generation unit 134 generates the second interpretation report based on the conversion.

[0081] FIG. 14 is a diagram illustrating an example of conversion of the anatomical knowledge graph provided in the report generating unit 134. In FIG.

[0082] 14, an anatomical knowledge graph for an X-ray CT image is shown at 250. Since the X-ray CT image 202 has three-dimensional information, the lung regions can be divided into smaller sections.

[0083] 15 and 16 are diagrams conceptually showing an anatomical knowledge graph in an X-ray CT image 202. Fig. 15 is a diagram showing the regions of the lungs as viewed from the inside surface, and Fig. 16 is a diagram showing the regions of the lungs as viewed from the outside surface.

[0084] The right lung is indicated by reference numeral 260 in FIG. 15 and reference numeral 264 in FIG. 16. The right lung is divided into ten regions, S1 to S10. Note that the S4 region cannot be observed from the medial side, and is therefore only shown in FIG. 16. On the other hand, the left lung is indicated by reference numeral 262 in FIG. 15 and reference numeral 266 in FIG. 16. The left lung is divided into regions S1 to S10, just like the right lung, but S1 and S2 are the same region (denoted as S1+2), and therefore the left lung is divided into nine regions. In this way, the X-ray CT image 202 contains three-dimensional information, and therefore the right and left lungs can each be divided into regions S1 to S10, as described above.

[0085] 14, the anatomical knowledge graphs denoted by reference numerals 252 and 254 are for plain X-ray images (AP and lateral views). In the plain X-ray images, the right and left lungs are each divided into three segments in the AP view, and the lungs are divided into two segments in the lateral view.

[0086] FIG. 17 is a diagram conceptually showing an anatomical knowledge graph for a plain X-ray image.

[0087] The right lung in the AP plain X-ray image 268a is divided into the upper right lung U1, the middle right lung U2, and the lower right lung U3, while the left lung is divided into the upper left lung U4, the middle left lung U5, and the lower left lung U6. Furthermore, the lung in the lateral plain X-ray image 268b is divided into the upper U7 and the lower U8.

[0088] In the anatomical knowledge graph 250 for X-ray CT images shown in FIG. 14, the lungs are branched into a right lung and a left lung, and the left lung is branched into a left upper lobe and a left lower lobe. The left upper lobe is branched into a left S1+S2 region, a left S3 region, a left S4 region, and a left S5 region. The left lower lobe is branched into a left S6 region, a left S8 region, a left S9 region, and a left S10 region. The right lung is branched into a right right lobe, a right middle lobe, and a right lower lobe. The right right lobe is branched into a right S1 region, a right S2 region, and a right S3 region. The right middle lobe is branched into a right S4 region and a right S5 region. The right lower lobe is branched into a right S6 region, a right S8 region, a right S9 region, and a right S10 region.

[0089] The anatomical knowledge graph for plain X-ray images shown in FIG. 14 includes an anatomical knowledge graph for a plain X-ray image 268a of an AP image and an anatomical knowledge graph for a plain X-ray image 268b of a lateral image. In the anatomical knowledge graph for a plain X-ray image of an AP image, the lungs are branched into a left lung and a right lung. The left lung is branched into an upper left, a middle left, and a lower left portion. The right lung is branched into an upper right, a middle right, and a lower right portion. In the anatomical knowledge graph for a plain X-ray image in the lateral direction, the lungs are branched into an upper and a lower portion. Then, the report generation unit 134 converts the anatomical knowledge graph for X-ray CT images 250 into anatomical knowledge graphs for plain X-ray images 252 and 254 as shown by the arrows in FIG. 14, and generates first radiology reports 206 to second radiology reports 208 based on this conversion.

[0090] FIG. 18 is a diagram showing an example of conversion of a disease knowledge graph provided in the report generating unit 134.

[0091] The disease knowledge graph shown in Figure 18 is an example of a disease knowledge graph related to nodules. Note that in Figure 18, since expressing it as a knowledge graph would be cumbersome, it is shown as a table.

[0092] The disease knowledge graph 270 for X-ray CT images is divided into categories of absorption value, boundary, shape, marginal characteristics, internal characteristics, and relationship with surrounding tissue. The classification target (class) of absorption value is classified into solid, partially solid, and ground glass type. The boundary is classified into clear and unclear. The shape is classified into irregular and near-circular. The marginal characteristics are classified into irregular, smooth, serrated, spiculed, lobulated, and linear. The internal characteristics are classified into bronchial radiolucency, calcification, cavity, and fat. The relationship with surrounding tissue is classified into pleural indentation and pleural contact.

[0093] On the other hand, in the disease knowledge graph for plain X-ray images 272, absorption values ​​are not easily visible due to the absorption coefficient similar to that of lung tissue, so they are classified only as solid. Boundaries are classified as clear or unclear, as in the disease knowledge graph for X-ray CT images 270. Shapes are also classified as irregular or circular, as in the disease knowledge graph for X-ray CT images 270. Since only the overall shape is visible in plain X-ray images, marginal features are not described. Internal features are visible due to the absorption coefficient similar to that of bone, and are classified as calcification. Relationships with surrounding tissues are classified as pleural indentation or pleural contact depending on the imaging direction. The report generation unit 134 then converts the disease knowledge graph for X-ray CT images 270 into the disease knowledge graph for plain X-ray images 272, as indicated by the arrows in FIG. 18 , and generates first radiology reports 206 to second radiology reports 208 based on this conversion.

[0094] FIG. 19 is a diagram for explaining the conversion from the first report to the second report by the report generating unit 134 equipped with the above-mentioned anatomical knowledge graph and disease knowledge graph.

[0095] Based on the conversion of the anatomical knowledge graph, the report generation unit 134 converts "right segments S4 and S5" in the first interpretation report 280 to "lower right lung" to generate the second interpretation report 282. Also, based on the conversion of the disease knowledge graph, the report generation unit 134 deletes "the margins are saw-toothed with spicules" from the first interpretation report 280 to generate the second interpretation report 282.

[0096] As described above, in this embodiment, the report generation unit 134 converts the anatomical knowledge graph and disease knowledge graph from those for X-ray CT images to those for plain X-ray images, and generates the second interpretation report 282 from the first interpretation report 280 based on the conversion.

[0097] <Fifth embodiment> <First example> Next, another embodiment (first example) of the learning of the learning model 126 will be described. In the above-described embodiment, an example has been described in which a pseudo plain X-ray image 204 is input to the learning model 126, and learning is performed to minimize the error between the estimated report output from the learning model 126 and the second radiology report. In this example, in addition to the above-described learning, the learning model 126 is trained using actual X-ray images and disease labels of the actual X-ray images, which are an additional training data set.

[0098] FIG. 20 is a diagram illustrating an additional training data set used in this example.

[0099] The additional training dataset 300 is composed of actual plain X-ray images 302 and disease labels 304. Here, the actual plain X-ray images 302 are X-ray images of the chest actually captured, for example, in the AP direction. The disease labels 304 are labels assigned by a doctor when they interpret the actual plain X-ray images 302, and are labels indicating, for example, the presence or absence of nodules. Specifically, the additional training dataset is obtained from the NIH (National Institutes of Health) Chest X-ray Dataset or the like.

[0100] FIG. 21 is a diagram illustrating the learning of the learning model 126 in this example.

[0101] In this example, a pseudo plain X-ray image 204 and an actual plain X-ray image 302 are input to the learning model 126. Note that, for example, the pseudo plain X-ray image 204 and the actual plain X-ray image 302 are input alternately to the learning model 126. Then, the learning model 126 outputs an estimated report 210. Here, the pseudo plain X-ray image 204 and the actual plain X-ray image 302 are images of the same subject, but may be images of different subjects.

[0102] The learning model 126 is composed of a DenseNet (Densely Connected Convolutional Networks) 127A and a knowledge graph 127B. The DenseNet 127A includes multiple dense blocks and multiple transition layers before and after the dense blocks, and has a network structure that exhibits high performance in classification tasks (e.g., disease detection). Within the dense blocks, skip connections are imposed on all layers to reduce gradient vanishing. The transition layers include convolutional layers and / or pooling layers. The knowledge graph 127B outputs a radiology report using, for example, the technology described in the literature (Li, Christy Y., et al., "Knowledge-Driven Encode, Retrieve, Paraphrase for Medical Image Report Generation," AAAI, 2019). The knowledge graph 127B outputs an inferred report 210 based on the output from the DenseNet 127A. The knowledge graph 127B is composed of, for example, an anatomy knowledge graph 306 and a disease knowledge graph 308. Here, in the learning of the conversion from the pseudo X-ray image to the disease knowledge graph, assistance is provided using the actual X-ray image and the disease label. Specifically, the disease label (presence or absence of nodules) is added to the subspace of the knowledge graph 127B of the learning model 126, and the label of presence or absence of nodules is added to the error of the actual X-ray image. As a result, the learning model 126 outputs the estimated report 210 by referring to the disease label 304, and is trained to output a more accurate radiology report.

[0103] <Second example> Next, a description will be given of another embodiment (second example) of the learning of the learning model 126. In this example, in addition to the above-described learning, the learning model 126 is trained using actual X-ray images and disease labels of the actual X-ray images, which are additional training data sets.

[0104] FIG. 22 is a diagram illustrating an additional training data set used in this example.

[0105] The additional learning dataset 320 is composed of an actual plain X-ray image 302 and an interpretation report (third interpretation report) 322. Here, the interpretation report 322 is, for example, an interpretation report created by a doctor who actually interprets the actual plain X-ray image 302.

[0106] 23 is a diagram illustrating the learning of the learning model 126 in this example. Note that the same reference numerals are used to denote parts that have already been explained, and explanations thereof will be omitted.

[0107] In this example, a pseudo plain X-ray image 204 and an actual plain X-ray image 302 are input to the learning model 126. Note that, for example, the pseudo plain X-ray image 204 and the actual plain X-ray image 302 are input alternately to the learning model 126. Then, the learning model 126 outputs an estimated report 210 for the pseudo plain X-ray image 204 and an estimated report 324 for the actual plain X-ray image 302. Here, learning is performed using the same DenseNet 127A and knowledge graph 127B for the pseudo plain X-ray image 204 and the actual plain X-ray image 302. Specifically, when the pseudo plain X-ray image 204 is input, the estimated report 210 is output as described above, and the learning model 126 is trained based on the error between the estimated report 210 and the second radiology report. On the other hand, when the actual plain X-ray image 302 is input, the estimated report 324 is output similarly via the DenseNet 127A and knowledge graph 127B. Then, the error acquisition unit 136 acquires the error between the output estimated report 324 and a portion of the radiology report 322 of the additional learning data set 320, and the learning control unit 138 causes the learning model 126 to learn based on the error.

[0108] As described above, the learning model 126 is trained using the pseudo plain X-ray image 204 as well as the actual plain X-ray image 302. By such training, a trained model that outputs a more accurate radiological report can be generated.

[0109] <Other> In the above embodiment, the hardware structure of the processing unit that executes various processes is the following various processors: The various processors include a CPU (Central Processing Unit), which is a general-purpose processor that executes software (programs) and functions as various processing units, a programmable logic device (PLD), such as an FPGA (Field Programmable Gate Array), whose circuit configuration can be changed after manufacture, and a dedicated electrical circuit, such as an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing specific processes.

[0110] A single processing unit may be configured with one of these various processors, or may be configured with two or more processors of the same or different types (for example, multiple FPGAs, or a combination of a CPU and an FPGA). Also, multiple processing units may be configured with a single processor. Examples of multiple processing units configured with a single processor include, first, a configuration in which one processor is configured with a combination of one or more CPUs and software, as typified by computers such as client and server, and this processor functions as multiple processing units. Second, a configuration in which a processor is used to realize the functions of an entire system including multiple processing units on a single IC (Integrated Circuit) chip, as typified by a system-on-chip (SoC). In this way, the various processing units are configured with one or more of the above-mentioned various processors as a hardware structure.

[0111] Furthermore, the hardware structure of these various processors is, more specifically, an electric circuit made up of a combination of circuit elements such as semiconductor elements.

[0112] The above-described configurations and functions can be realized by any hardware, software, or a combination of both. For example, the present invention can be applied to a program that causes a computer to execute the above-described processing steps (processing procedures), a computer-readable recording medium (non-transitory recording medium) on which such a program is recorded, or a computer on which such a program can be installed.

[0113] Although examples of the present invention have been described above, it goes without saying that the present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the present invention. [Explanation of symbols]

[0114] 100: Learning device 112: Communications Department 114: Memory 116:Operation unit 118:CPU 120:RAM 122:ROM 124:Display section 126: Learning model 129: Processor 130: Learning data acquisition unit 132: Pseudo image generation unit 134: Report generation section 136:Error acquisition section 138: Learning control unit 200: Training dataset 202: X-ray CT image 204: Pseudo plain X-ray image 205A: Organ label conversion list 205B: Disease label conversion list 206: First interpretation report 208: Second interpretation report 210: Estimated Report

Claims

1. A learning device including a processor, a memory that stores a learning data set of an X-ray CT image having three-dimensional information and a first interpretation report for the X-ray CT image, and a learning model that generates an interpretation report from a plain X-ray image having two-dimensional information, The processor: a process of projecting the X-ray CT image to generate a pseudo plain X-ray image and inputting the pseudo plain X-ray image into the learning model; A process of converting the first radiology report to generate a second radiology report for the pseudo plain X-ray image; a process of acquiring an error between an estimated report for the pseudo plain X-ray image output by the learning model based on the pseudo plain X-ray image input and the second radiology report; training the learning model using the error; A learning device that performs the following:

2. The learning device described in claim 1, wherein the process of generating the second radiology report generates the second radiology report from the first radiology report by converting organ labels included in the first radiology report into organ labels of the second radiology report.

3. The learning device described in claim 1 or 2, wherein the process of generating the second radiology report generates the second radiology report from the first radiology report by converting a disease label contained in the first radiology report into a disease label of the second radiology report.

4. The process of generating the second radiology report includes: The learning device according to claim 1 , wherein a first knowledge graph corresponding to the first radiology report is converted into a second knowledge graph corresponding to the second radiology report, and the second radiology report is generated based on the conversion.

5. The memory stores the X-ray CT image obtained by capturing an object in a first posture, and the learning model generates an interpretation report from the plain X-ray image obtained by capturing an object in a second posture, A learning device described in any one of claims 1 to 4, wherein the process of inputting the pseudo-simple X-ray image generates the pseudo-simple X-ray image in the second posture from the X-ray CT image in the first posture, and inputs the pseudo-simple X-ray image in the second posture to the learning model.

6. A learning device described in any one of claims 1 to 5, wherein the process of inputting the pseudo-simple X-ray image generates the pseudo-simple X-ray image projected in a first direction from the X-ray CT image and the pseudo-simple X-ray image projected in a second direction, and inputs the pseudo-simple X-ray image projected in the first direction and the pseudo-simple X-ray image projected in the second direction to the learning model.

7. the memory stores an additional training data set of the plain X-ray images and disease labels of the plain X-ray images; A learning device described in any one of claims 1 to 6, wherein the process of acquiring the error acquires the error between the estimated report for the pseudo-simple X-ray image output by the learning model with reference to the disease label and the second interpretation report.

8. the memory stores an additional learning data set of the plain X-ray image and a third interpretation report for the plain X-ray image; A learning device described in any one of claims 1 to 6, wherein the process of acquiring the error acquires the error between the estimated report for the pseudo-simple X-ray image output based on the pseudo-simple X-ray image input by the learning model and the second interpretation report, and the error between the estimated report for the simple X-ray image output based on the simple X-ray image input by the learning model and the third interpretation report.

9. A learning method in which a processor uses a learning data set of X-ray CT images having three-dimensional information and a first interpretation report for the X-ray CT images stored in a memory to train a learning model that generates an interpretation report from a plain X-ray image having two-dimensional information, the method comprising: generating a pseudo plain X-ray image by projecting the X-ray CT image, and inputting the pseudo plain X-ray image into the learning model; converting the first radiology report to generate a second radiology report for the pseudo plain X-ray image; acquiring an error between an estimated report for the pseudo plain X-ray image output by the learning model based on the pseudo plain X-ray image input and the second radiology report; training the learning model using the error; Learning methods including.

10. The learning method described in claim 9, wherein the step of generating the second radiology report generates the second radiology report from the first radiology report by converting organ labels included in the first radiology report into organ labels of the second radiology report.

11. The learning method described in claim 9 or 10, wherein the step of generating the second radiology report generates the second radiology report from the first radiology report by converting a disease label included in the first radiology report into a disease label of the second radiology report.

12. The step of generating the second radiology report includes: The learning method according to claim 9, further comprising: converting a first knowledge graph corresponding to the first radiology report into a second knowledge graph corresponding to the second radiology report; and generating the second radiology report based on the conversion.

13. A learning program that causes the processor to execute the processing of each step in the learning method according to any one of claims 9 to 12.

14. A non-transitory computer-readable recording medium having the program according to claim 13 recorded thereon.

15. A trained model trained by the training method according to any one of claims 9 to 12.

Citation Information

Patent Citations

  • Head medical image auxiliary interpretation report generation method based on neural network

    CN111223085A

  • Medical image information identification method, device and system based on multiple neural networks

    CN112215845A

  • Device, method, and program for supporting preparation of medical document

    JP2019153250A

  • Medical information processing device and medical information processing program

    JP2020173614A

  • Systems, methods and media for automatically generating a bone age assessment from a radiograph

    US20200020097A1