Learning device, method and program, and information processing device, method and program
Patent Information
- Application Number
- JP2021124683
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-07-29
- Publication Date
- 2025-06-02
- Estimated Expiration
- 2041-07-29
AI Technical Summary
The limited number of medical images and sentences poses a challenge in constructing a trained model that accurately associates images with sentences, particularly in the field of medical image analysis, making it difficult to achieve high accuracy in image-sentence associations.
A learning device that utilizes two neural networks to derive feature amounts for images and sentences, adjusting the networks to minimize the distance between relevant attributes in a feature space, thereby enhancing the association accuracy by learning the relevance of object attributes in both modalities.
This approach allows for high-accuracy association of images and sentences, enabling precise identification of objects in medical images and accurate creation of interpretation reports.
Smart Images

Figure 00000022_0000 
Figure 00000023_0000 
Figure 00000023_0001
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a learning device, a method, and a program, and an information processing device, a method, and a program. [Background technology]
[0002] A method has been proposed for constructing a feature space to which features such as feature vectors extracted from images belong using a trained model that has undergone machine learning such as deep learning. For example, Non-Patent Document 1 proposes a method for training a trained model so that features of images belonging to the same class are close to each other in the feature space, and features of images belonging to different classes are far apart in the feature space. Also known is a technique for associating features extracted from images with features extracted from sentences based on the distance in the feature space (see Non-Patent Document 2). [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Deep metric learning using Triplet network, Elad Hoffer et al., 20 Dec 2014, arXiv:1412.6622 [Non-patent document 2] Learning Two-Branch Neural Networks for Image-Text Matching Tasks, Liwei Wang et al., 11 Apr 2017, arXiv:1704.03470 Summary of the Invention [Problem to be solved by the invention]
[0004] As described in Non-Patent Document 2, in order to accurately train a trained model that associates images with sentences, a large amount of training data that associates features contained in images with features described in sentences is required. However, since the number of images and sentences is limited, it may be difficult to prepare a large amount of training data. In particular, in the field of medical image analysis, since the number of medical images is limited, it is difficult to construct a trained model that can associate images with sentences with high accuracy.
[0005] The present disclosure has been made in consideration of the above circumstances, and aims to enable highly accurate correspondence between images and text. [Means for solving the problem]
[0006] A learning device according to the present disclosure includes at least one processor, The processor derives a first feature amount for an object included in the image using a first neural network; deriving second features for sentences including descriptions of objects using a second neural network; A first attribute is an attribute of an object included in an image, and a second attribute is an attribute of a sentence included in the image. The first neural network and the second neural network are trained so that, in a feature space to which the first feature and the second feature belong, the higher the correlation between the combination of the first attribute and the second attribute, the smaller the distance between the derived first feature and the derived second feature, compared to when the correlation between the combination of the first attribute and the second attribute is low. This constructs a first derivation model that derives features for an object included in an image and a second derivation model that derives features for a sentence that includes a description of the object.
[0007] In the learning device according to the present disclosure, the first attribute may be the position of an object in an image, and the second attribute may be the position of the object described in a sentence.
[0008] Furthermore, in the learning device according to the present disclosure, the first attribute may be the property of an object included in an image, and the second attribute may be the property of an object described in a sentence.
[0009] In the learning device according to the present disclosure, the processor may obtain the first attribute and the second attribute based on a co-occurrence relationship between texts expressing properties.
[0010] Furthermore, in the learning device according to the present disclosure, the processor may train the first neural network and the second neural network so that, when an object included in the image corresponds to an object described in the sentence, the distance between the derived first feature and the derived second feature in the feature space is small, and, when the object included in the image does not correspond to an object described in the sentence, the processor may train the first neural network and the second neural network so that the distance between the derived first feature and the derived second feature in the feature space is large.
[0011] In addition, in the learning device according to the present disclosure, the image is a medical image, and the object included in the image is a lesion included in the medical image, The sentence may be a finding sentence that describes findings about the lesion.
[0012] A first information processing device according to the present disclosure includes at least one processor, The processor derives first feature amounts for one or more objects included in the target image using a first derived model constructed by the learning device according to the present disclosure; deriving second features for one or more target sentences including a description related to the object using a second derived model constructed by the learning device according to the present disclosure; Identifying a first feature corresponding to the second feature based on a distance in a feature space between the derived first feature and the derived second feature; The object from which the specified first feature amount has been derived is displayed in the target image, distinguished from other areas.
[0013] A second information processing device according to the present disclosure includes at least one processor, The processor accepts an input of a target sentence including a description of an object; deriving second features for the input target sentence using a second derivation model constructed by the learning device according to the present disclosure; By referring to a database in which first feature amounts for one or more objects included in each of a plurality of reference images, which are derived by a first derivation model constructed by the learning device according to the present disclosure, are associated with each of the reference images, at least one first feature amount corresponding to the second feature amount is identified based on the distance in feature space between the first feature amount for the plurality of reference images and the derived second feature amount; A reference image associated with the identified first feature amount is identified.
[0014] The learning method according to the present disclosure includes: deriving a first feature amount for an object included in an image using a first neural network; deriving second features for sentences including descriptions of objects using a second neural network; A first attribute is an attribute of an object included in an image, and a second attribute is an attribute of a sentence included in the image. The first neural network and the second neural network are trained so that, in a feature space to which the first feature and the second feature belong, the higher the correlation between the combination of the first attribute and the second attribute, the smaller the distance between the derived first feature and the derived second feature, compared to when the correlation between the combination of the first attribute and the second attribute is low. This constructs a first derivation model that derives features for an object included in an image and a second derivation model that derives features for a sentence that includes a description of the object.
[0015] A first information processing method according to the present disclosure includes deriving first feature amounts for one or more objects included in a target image using a first derived model constructed by a learning method according to the present disclosure; deriving second features for one or more target sentences including a description related to the object using a second derived model constructed by the learning method according to the present disclosure; Identifying a first feature corresponding to the second feature based on a distance in a feature space between the derived first feature and the derived second feature; The object from which the specified first feature amount has been derived is displayed in the target image, distinguished from other areas.
[0016] A second information processing method according to the present disclosure includes receiving an input of a target sentence including a description related to an object; deriving second features for the input target sentence using a second derived model constructed by the learning method according to the present disclosure; By referring to a database in which first feature amounts for one or more objects included in each of a plurality of reference images, which are derived by a first derivation model constructed by the learning method according to the present disclosure, are associated with each of the reference images, at least one first feature amount corresponding to the second feature amount is identified based on the distance in feature space between the first feature amount for the plurality of reference images and the derived second feature amount; A reference image associated with the identified first feature amount is identified.
[0017] The learning method and the first and second information processing methods according to the present disclosure may be provided as a program for causing a computer to execute the method. [Effects of the Invention]
[0018] According to the present disclosure, images and text can be associated with high accuracy. [Brief explanation of the drawings]
[0019] [Figure 1] FIG. 1 is a diagram showing a schematic configuration of a medical information system to which a learning device and an information processing device according to a first embodiment of the present disclosure are applied. [Figure 2] FIG. 1 is a diagram showing a schematic configuration of a learning device according to a first embodiment; [Figure 3]FIG. 1 is a diagram showing a schematic configuration of an information processing apparatus according to a first embodiment; [Figure 4] Functional configuration diagram of a learning device according to a first embodiment [Figure 5] Diagram showing an example of a medical image and interpretation report [Figure 6] FIG. 1 is a diagram illustrating the processes performed by a first derivation unit, a second derivation unit, an attribute acquisition unit, and a learning unit in a first embodiment. [Figure 7] Schematic diagram of the second neural network [Figure 8] Diagram to explain the relative positions of the lung compartments [Figure 9] Diagram for explaining the derivation of loss [Figure 10] Functional configuration diagram of an information processing device according to a first embodiment [Figure 11] FIG. 1 is a diagram illustrating identification of a first feature amount. [Figure 12] Diagram showing the display screen [Figure 13] 1 is a flowchart showing a learning process performed in the first embodiment. [Figure 14] 1 is a flowchart showing information processing performed in the first embodiment. [Figure 15] Diagram showing examples of co-occurrence relationships [Figure 16] FIG. 10 is a diagram illustrating the processes performed by the first derivation unit, the second derivation unit, the attribute acquisition unit, and the learning unit in the second embodiment. [Figure 17] Diagram showing training data [Figure 18] FIG. 10 is a diagram illustrating plots of feature amounts in the third embodiment. [Figure 19] Functional configuration diagram of an information processing device according to a second embodiment [Figure 20] Diagram to explain search [Figure 21] Diagram showing the display screen [Figure 22] 10 is a flowchart showing information processing performed in the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0020] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. First, the configuration of a medical information system to which a learning device and an information processing device according to a first embodiment of the present disclosure are applied will be described. FIG. 1 is a diagram showing a schematic configuration of a medical information system 1. The medical information system 1 shown in FIG. 1 is a system for capturing an examination target region of a patient as a subject, storing the medical images acquired by capturing the images, having a radiologist interpret the medical images and create an interpretation report, and allowing the requesting doctor from the medical department to view the interpretation report and observe the details of the medical image to be interpreted, based on an examination order from a doctor from a medical department using a known ordering system.
[0021] As shown in Figure 1, the medical information system 1 is configured by connecting multiple imaging devices 2, multiple interpretation WSs (Workstations) 3 which are interpretation terminals, a medical treatment WS 4, an image server 5, an image DB (DataBase) 5A, a report server 6, a report DB 6A, and a learning device 7 in a state where they can communicate with each other via a wired or wireless network 10.
[0022] Each device is a computer installed with an application program that causes the device to function as a component of the medical information system 1. The application program is recorded on a recording medium such as a DVD (Digital Versatile Disc) or a CD-ROM (Compact Disc Read Only Memory) and distributed, and is installed on the computer from the recording medium. Alternatively, the application program is stored in a storage device of a server computer connected to the network 10 or in network storage in an externally accessible state, and is downloaded to the computer upon request and installed.
[0023] The imaging device 2 is a device (modality) that captures an image of a patient's diagnostic target area to generate a medical image representing the diagnostic target area. Specifically, it is a plain X-ray imaging device, a CT device, an MRI device, a PET (Positron Emission Tomography) device, etc. The medical image generated by the imaging device 2 is transmitted to the image server 5 and stored in the image DB 5A.
[0024] The image interpretation WS3 is a computer used by, for example, a radiologist to interpret medical images and create image interpretation reports, and includes an information processing device according to this embodiment (details of which will be described later). The image interpretation WS3 issues requests for viewing medical images to the image server 5, performs various image processing on medical images received from the image server 5, displays the medical images, and accepts input of findings related to the medical images. The image interpretation WS3 also performs analysis processing on medical images, supports the creation of image interpretation reports based on the analysis results, requests the report server 6 to register and view image interpretation reports, and displays image interpretation reports received from the report server 6. These processes are performed by the image interpretation WS3 executing software programs for each process.
[0025] The medical treatment WS4 is a computer used by, for example, a doctor in a medical department to observe images in detail, view radiology reports, and create electronic medical records, and is composed of a processing device, a display device such as a monitor, and input devices such as a keyboard and a mouse. The medical treatment WS4 sends image viewing requests to the image server 5, displays images received from the image server 5, requests radiology reports to the report server 6, and displays radiology reports received from the report server 6. These processes are performed by the medical treatment WS4 executing software programs for each process.
[0026] The image server 5 is a general-purpose computer installed with a software program that provides the functions of a database management system (DBMS). The image server 5 also has storage in which the image DB 5A is configured. This storage may be a hard disk device connected to the image server 5 via a data bus, or a disk device connected to a NAS (Network Attached Storage) or SAN (Storage Area Network) connected to the network 10. When the image server 5 receives a request to register a medical image from the imaging device 2, it formats the medical image in a database format and registers it in the image DB 5A.
[0027] The image DB 5A stores image data and supplementary information of medical images acquired by the imaging device 2. The supplementary information includes, for example, an image ID (identification) for identifying each medical image, a patient ID for identifying the patient, an examination ID for identifying the examination, a unique ID (UID) assigned to each medical image, the examination date and time when the medical image was generated, the type of imaging device used in the examination to acquire the medical image, patient information such as the patient's name, age, and gender, the examination site (imaging site), imaging information (imaging protocol, imaging sequence, imaging technique, imaging conditions, use of contrast agent, etc.), and a series number or collection number when multiple medical images are acquired in one examination. In this embodiment, a first feature value of the medical image derived in the interpretation WS 3 as described below is associated with the medical image and registered in the image DB 5A.
[0028] In addition, when the image server 5 receives a viewing request from the image interpretation WS3 and medical treatment WS4 via the network 10, it searches for medical images registered in the image DB 5A and transmits the searched medical images to the image interpretation WS3 and medical treatment WS4 that made the request.
[0029] A software program that provides a general-purpose computer with the functions of a database management system is installed in the report server 6. When the report server 6 receives a request to register an interpretation report from the interpretation WS 3, it converts the interpretation report into a database format and registers it in the report DB 6A.
[0030] The report DB 6A stores a large number of radiology reports, each containing a statement of findings created by a radiology doctor using the radiology WS 3. The radiology report may include information such as the medical image to be interpreted, an image ID for identifying the medical image, a radiology doctor ID for identifying the radiology doctor who performed the interpretation, the name of the lesion, location information of the lesion, and characteristics of the lesion. In this embodiment, the report DB 6A stores radiology reports in association with one or more medical images for which the radiology report was created.
[0031] In addition, when the report server 6 receives a request to view an interpretation report from the interpretation WS3 and medical treatment WS4 via the network 10, it searches for the interpretation report registered in the report DB6A and sends the retrieved interpretation report to the interpretation WS3 and medical treatment WS4 that made the request.
[0032] The network 10 is a wired or wireless local area network that connects various devices within the hospital. If the interpretation WS3 is installed in other hospitals or clinics, the network 10 may be configured by connecting the local area networks of each hospital via the Internet or a dedicated line.
[0033] Next, the learning device 7 will be described. First, the hardware configuration of the learning device 7 according to the first embodiment will be described with reference to FIG. 2. As shown in FIG. 2, the learning device 7 includes a CPU (Central Processing Unit) 11, non-volatile storage 13, and memory 16 as a temporary storage area. The learning device 7 also includes a display 14 such as an LCD display, an input device 15 including a keyboard, a pointing device such as a mouse, and a network I / F (Interface) 17 connected to a network 10. The CPU 11, storage 13, display 14, input device 15, memory 16, and network I / F 17 are connected to a bus 18. The CPU 11 is an example of a processor in the present disclosure.
[0034] The storage 13 is realized by a hard disk drive (HDD), a solid state drive (SSD), a flash memory, etc. The storage 13 as a storage medium stores a learning program 12. The CPU 11 reads the learning program 12 from the storage 13, expands it into the memory 16, and executes the expanded learning program 12.
[0035] Next, an information processing device 30 according to the first embodiment, which is included in the interpretation WS 3, will be described. First, the hardware configuration of the information processing device 30 according to this embodiment will be described with reference to FIG. 3. As shown in FIG. 3, the information processing device 30 includes a CPU 41, non-volatile storage 43, and memory 46 as a temporary storage area. The information processing device 30 also includes a display 44 such as a liquid crystal display, an input device 45 including a keyboard, a pointing device such as a mouse, and the like, and a network I / F 47 connected to the network 10. The CPU 41, storage 43, display 44, input device 45, memory 46, and network I / F 47 are connected to a bus 48. The CPU 41 is an example of a processor in the present disclosure.
[0036] Like the storage 13, the storage 43 is realized by an HDD, an SSD, a flash memory, or the like. The storage 43 as a storage medium stores an information processing program 42. The CPU 41 reads the information processing program 42 from the storage 43, loads it into the memory 46, and executes the loaded information processing program 42.
[0037] Next, the functional configuration of the learning device according to the first embodiment will be described. Fig. 4 is a diagram showing the functional configuration of the learning device according to the first embodiment. As shown in Fig. 4, the learning device 7 includes an information acquisition unit 21, a first derivation unit 22, a second derivation unit 23, an attribute acquisition unit 24, and a learning unit 25. When the CPU 11 executes the learning program 12, the CPU 11 functions as the information acquisition unit 21, the first derivation unit 22, the second derivation unit 23, the attribute acquisition unit 24, and the learning unit 25.
[0038] The information acquisition unit 21 acquires medical images and interpretation reports from the image server 5 and the report server 6, respectively, via the network I / F 17. The medical images and interpretation reports are used for training the neural network, which will be described later. FIG. 5 shows examples of medical images and interpretation reports. As shown in FIG. 5, the medical image 51 is a three-dimensional image made up of multiple tomographic images. In this embodiment, the medical image 51 is a CT image of the chest of a human body. Furthermore, as shown in FIG. 5, the multiple tomographic images include a tomographic image 51A that includes a lesion.
[0039] 5, the image interpretation report 52 includes a finding statement 53. The content of the finding statement 53 relates to a lesion contained in the medical image 51, for example, "There is a 12 mm solid nodule in the right lower lobe S6. There is also a micronodule in the left lung S9."
[0040] Next, we will explain the first derivation unit 22, the second derivation unit 23, the attribute acquisition unit 24, and the learning unit 25. Fig. 6 is a diagram schematically showing the processes performed by the first derivation unit 22, the second derivation unit 23, the attribute acquisition unit 24, and the learning unit 25 in the first embodiment.
[0041] The first derivation unit 22 derives first feature amounts for one or more objects included in the medical image using a first neural network 61 (NN) to construct a first derivation model for deriving feature amounts for objects included in the medical image. In this embodiment, the first neural network 61 is a convolutional neural network (CNN), but is not limited to this. As shown in FIG. 6 , the first derivation unit 22 inputs an image 55, such as a medical image containing an object such as a lesion, to the first neural network 61. The first neural network 61 extracts an object 55A, such as a lesion, included in the image 55 and derives a feature vector of the object 55A as a first feature amount V1.
[0042] The second derivation unit 23 derives second feature quantities for sentences including descriptions related to objects using a second neural network (NN) 62 to construct a second derivation model for deriving feature quantities for sentences including descriptions related to objects. FIG. 7 is a diagram schematically illustrating the second neural network 62. As shown in FIG. 7, the second neural network 62 has an embedding layer 62A, a recurrent neural network layer (hereinafter referred to as an RNN layer) 62B, and a fully connected layer 62C. The second derivation unit 23 performs morphological analysis on the input sentence to divide the sentence into words, and inputs the words into the embedding layer 62A. The embedding layer 62A outputs feature vectors of the words included in the input sentence. For example, when sentence 56, "There is a 12 mm solid nodule in the right lower lobe S6," is input to second neural network 62, sentence 56 is divided into the words "There is a 12 mm solid nodule in the right lower lobe S6." Each word is then input to an element of embedding layer 62A.
[0043] The RNN layer 62B outputs a feature vector 66 that takes into account the context of the feature vector 65 of each word. The fully connected layer 62C integrates the feature vectors 66 output by the RNN layer 62B and outputs the feature vector of the sentence 56 input to the second neural network 62 as a second feature V2. Note that when the feature vectors 66 output by the RNN 62B are input to the fully connected layer 62C, the weighting of the feature vectors 66 for important words may be increased.
[0044] Here, in the first embodiment, the second feature V2-1 is acquired by sentence 56, "There is a 12 mm solid nodule in the right lower lobe S6," and the second feature V2-2 is acquired by sentence 57, "There is also a micronodule in the left lung S9."
[0045] The attribute acquisition unit 24 acquires a first attribute P1, which is an attribute of an object 55A included in an image 55, and second attributes P2-1 and P2-2, which are attributes of sentences 56 and 57. Hereinafter, the second attributes P2-1 and P2-2 may be represented by the second attribute P2. In the first embodiment, the first attribute P1 is the position of the object 55A included in the image 55, and the second attribute P2 is the position of the object described in sentences 56 and 57. The attribute acquisition unit 24 acquires the first attribute P1 based on one or more elements representing the position included in the first feature V1. Furthermore, the attribute acquisition unit 24 acquires the second attributes P2-1 and P2-2 based on one or more elements representing the position included in the second feature V2-1 and V2-2. In this embodiment, it is assumed that "right lower lobe S7" is acquired as the first attribute P1, "right lower lobe S6" is acquired as the second attribute P2-1, and "left lung S9" is acquired as the second attribute P2-2.
[0046] The attribute acquisition unit 24 has a derivation model that has undergone machine learning to derive attributes of objects included in an image, and a derivation model that has undergone machine learning to derive attributes of objects described in a sentence. Attributes output by each derivation model include the location of a lesion, as well as the type, size, and characteristics of the lesion. The attribute acquisition unit 24 derives a positive or negative determination result for multiple types of characteristic items as the characteristics of the object. The characteristic items include, for example, the shape of the margin (lobulated, spicules), the absorption value (solid, ground-glass), the clarity of the boundary, the presence or absence of calcification, the presence or absence of pleural indentation, etc. for a lesion included in the lung.
[0047] The learning unit 25 learns the first neural network 61 and the second neural network 62 so that the distance between the derived first feature amount V1 and second feature amount V2 in the feature space to which the first feature amount V1 and the second feature amount V2 belong becomes smaller when the relevance of the combination of the first attribute P1 and the second attribute P2 is higher than when the relevance of the combination of the first attribute P1 and the second attribute P2 is low. For this purpose, the learning unit 25 derives the relevance of the combination of the first attribute P1 and the second attribute P2.
[0048] FIG. 8 is a diagram for explaining the positional relationship of the regions that make up the lungs. As shown in FIG. 8, there are right and left lungs, and the right lung is divided into upper, middle, and lower lobes, while the left lung is divided into upper and lower lobes. Furthermore, the right upper lobe is divided into regions S1 to S3, the right middle lobe is divided into regions S4 and S5, and the right lower lobe is divided into regions S6 to S10. The left upper lobe is divided into regions S1+2 to S5, and the left lower lobe is divided into regions S6 to S10. The right lower lobe is synonymous with the right lower lobe.
[0049] The learning unit 25 determines the relevance of the combination of the first attribute P1 and the second attribute P2 based on the distance between the position represented by the first attribute P1 and the position represented by the second attribute P2. For example, if the position represented by the first attribute P1 and the position represented by the second attribute P2 are in the same lobe and in the same or adjacent regions, the learning unit 25 determines the relevance of the combination to be "high." Otherwise, the learning unit 25 determines the relevance of the combination to be "low." Note that the relevance of the combination is not limited to two types, "high" and "low," and three or more types of relevance may be determined. For example, if the position represented by the first attribute P1 and the position represented by the second attribute P2 are the same, the relevance of the combination may be determined to be "high." If the regions are adjacent, the relevance of the combination may be determined to be "slightly high." Otherwise, the relevance of the combination may be determined to be "low."
[0050] The relevance of the combination may be determined based on a rule base, or may be determined using a trained model that has been trained to output the relevance of the combination when the first attribute P1 and the second attribute P2 are input.
[0051] The learning unit 25 plots the first feature amount V1 and the second feature amount V2 in a feature space defined by the first feature amount V1 and the second feature amount V2. Then, the learning unit 25 derives the distance between the first feature amount V1 and the second feature amount V2 in the feature space. Here, since the first feature amount V1 and the second feature amount V2 are n-dimensional vectors, the feature space is also n-dimensional. Note that, for the sake of explanation, FIG. 6 shows the first feature amount V1 and the second feature amount V2 as two-dimensional, and plots the first feature amount V1 and the second feature amount V2 (V2-1, V2-2) in the two-dimensional feature space.
[0052] When the relevance of the combination of the first attribute P1 and the second attribute P2 is "high," the learning unit 25 trains the first neural network 61 and the second neural network 62 so that the first feature amount V1 and the second feature amount V2 become closer in the feature space. On the other hand, when the relevance of the combination of the first attribute P1 and the second attribute P2 is "low," the learning unit 25 trains the first neural network 61 and the second neural network 62 so that the first feature amount V1 and the second feature amount V2 become farther apart in the feature space.
[0053] Here, the first attribute P1 is the right lower lobe S7, and the second attribute P2-1 is the right lower lobe S6, so the relevance of the combination of the first attribute P1 and the second attribute P2-1 is "high." On the other hand, the second attribute P2-2 is the left lung S9, so the relevance of the combination of the first attribute P1 and the second attribute P2-2 is "low." For this reason, the learning unit 25 trains the first neural network 61 and the second neural network 62 so that the derived first feature amount V1 and the derived second feature amount V2-1 are closer to each other and the derived first feature amount V1 and the derived second feature amount V2-2 are farther apart in the feature space shown in FIG.
[0054] To this end, the learning unit 25 derives the distance between the first feature amount V1 and the second feature amount V2 in the feature space. Any distance, such as Euclidean distance or Mahalanobis distance, can be used as the distance. Then, a loss to be used during learning is derived based on the distance. FIG. 9 is a diagram for explaining the derivation of the loss. First, for the first feature amount V1 and the second feature amount V2-1 whose combination has a "high" relevance, the learning unit 25 calculates the distance d1 in the feature space. Then, the learning unit 25 compares the distance d1 with a predetermined threshold value α0 and derives the loss L1 based on the following formula (1).
[0055] That is, if the distance d1 between the first feature V1 and the second feature V2-1 is greater than the threshold value α0, the loss L1 for training the first and second neural networks 61, 62 is calculated as d1-α0 so that the distance of the second feature V2-1 from the first feature V1 is smaller than the threshold value α0. On the other hand, if the distance d1 between the first feature V1 and the second feature V2-1 is equal to or smaller than the threshold value α0, there is no need to reduce the distance d1 between the first feature V1 and the second feature V2-1, so L1 is set to 0. L1=d1-α0(d1>α0) L1=0(d1≦α0) (1)
[0056] On the other hand, for the first feature amount V1 and the second feature amount V2-2 whose attribute relevance is "low", the learning unit 25 calculates the distance d2 in the feature space. Then, the learning unit 25 compares the distance d2 with a predetermined threshold value β0 and derives the loss L2 based on the following formula (2).
[0057] That is, if the distance d2 between the first feature V1 and the second feature V2-2 is smaller than the threshold value β0, the loss L2 for training the first and second neural networks 61, 62 is calculated from β0-d2 so that the distance of the second feature V2-2 from the first feature V1 is greater than the threshold value β0. On the other hand, if the distance d2 between the first feature V1 and the second feature V2-2 is equal to or greater than the threshold value β0, there is no need to increase the distance d2 between the first feature V1 and the second feature V2-2, so L2 is set to 0. L2=β0-d2(d2<β0) L2=0(d2≧β0) (2)
[0058] The learning unit 25 learns the first neural network 61 and the second neural network 62 based on the derived losses L1 and L2. That is, when d1>α0 and when d2<β0, the learning unit 25 learns the weights of the connections between the layers constituting the first neural network 61 and the second neural network 62 and the kernel coefficients used for convolution so that the losses L1 and L2 become small.
[0059] The learning unit 25 then repeatedly performs learning until the loss L1 becomes equal to or less than a predetermined threshold value α0 and the loss L2 becomes equal to or greater than a threshold value β0. Preferably, the learning unit 25 repeatedly performs learning until the loss L1 becomes equal to or less than the threshold value α0 a predetermined number of times in a row and the loss L2 becomes equal to or greater than the threshold value β0 a predetermined number of times in a row. This allows the first and second derivation models to be constructed that derive the first feature value V1 and the second feature value V2 such that the distance in the feature space is small when the relevance between the combination of the attribute of an object included in the image and the attribute of an object described in the sentence is high, and the distance in the feature space is large when the relevance between the combination is low. The learning unit 25 may repeatedly perform learning a predetermined number of times.
[0060] The first and second derived models thus constructed are transmitted to the interpretation WS 3 and used in the information processing device according to the first embodiment.
[0061] In the first embodiment, the medical images and interpretation reports acquired by the information acquisition unit 21 are associated with each other, i.e., even if the findings contained in the interpretation report relate to findings in the medical image, this relationship is not taken into consideration during the learning process described below, and the first and second neural networks 61, 62 are trained based only on the relevance of the combination of the first attribute P1 and the second attribute P2.
[0062] Next, the functional configuration of the information processing device according to the first embodiment will be described. Fig. 10 is a diagram showing the functional configuration of the information processing device according to the first embodiment. As shown in Fig. 10, the information processing device 30 includes an information acquisition unit 31, a first analysis unit 32, a second analysis unit 33, an identification unit 34, and a display control unit 35. When the CPU 41 executes the information processing program 42, the CPU 41 functions as the information acquisition unit 31, the first analysis unit 32, the second analysis unit 33, the identification unit 34, and the display control unit 35.
[0063] The information acquisition unit 31 acquires a target medical image G0 to be read from the image server 5 according to an instruction from the input device 45 by the radiologist who is the operator.
[0064] The first analysis unit 32 analyzes the target medical image G0 using the first derivation model 32A constructed by the learning device 7 described above, and derives a first feature quantity V1 for an object such as a lesion included in the target medical image G0. In the present embodiment, it is assumed that the target medical image G0 includes two objects, and first feature quantities V1-1 and V1-2 are derived for each of the two objects respectively.
[0065] <s Here, in the information processing apparatus 30 according to the first embodiment, in the reading WS3, the radiologist reads the target medical image G0, and a reading report is generated by inputting a findings text including the reading result using the input device 45. The second analysis unit 33 analyzes the input findings text using the second derivation model 33A constructed by the learning device 7 described above, and derives a second feature quantity V2 for the input findings text.
[0066] The specifying unit 34 derives the distance in the feature space between the first feature quantity V1 derived by the first analysis unit 32 and the second feature quantity V2 derived by the second analysis unit 33. Then, based on the derived distance, it specifies the first feature quantity V1 corresponding to the second feature quantity V2. FIG. 11 is a diagram for explaining the specification of the first feature quantity. In FIG. 11, the feature space is shown two-dimensionally for the sake of explanation. As shown in FIG. 11, in the feature space, when comparing the distance d3 between the first feature quantity V1-1 and the second feature quantity V2 and the distance d4 between the first feature quantity V1-2 and the second feature quantity V2, d3 < d4. Therefore, the specifying unit 34 specifies the first feature quantity corresponding to the second feature quantity V2 as the first feature quantity V1-1.
[0067] The display control unit 35 displays the object from which the identified first feature amount has been derived, distinguishing it from other areas in the target medical image G0. FIG. 12 is a diagram showing a screen for creating an interpretation report displayed on the interpretation WS3. As shown in FIG. 12, the interpretation report creation screen 70 has an image display area 71 and a text display area 72. The target medical image G0 is displayed in the image display area 71. In FIG. 12, the target medical image G0 is one tomographic image that constitutes a three-dimensional image of the chest. The text display area 72 displays a finding entered by the radiologist. In FIG. 12, the finding reads, "There is a 10 mm solid nodule in the right lung S6." Note that the right lung S6 is synonymous with the right lower lobe S6.
[0068] The target medical image G0 shown in FIG. 12 includes a lesion 73 in the right lung and a lesion 74 in the left lung. Comparing the first feature value V1-1 derived for the lesion 73 in the right lung with the first feature value V1-2 derived for the lesion 74 in the left lung, the distance between the first feature value V1-1 and the second feature value V2 derived for the finding, "A 10 mm solid nodule is present in the right lung S6," is smaller. Therefore, the display control unit 35 displays the lesion 73 in the right lung in a manner that distinguishes it from other areas in the target medical image G0. In FIG. 12, the lesion 73 in the right lung is surrounded by a rectangular mark 75 to distinguish it from other areas, but this is not limiting. Marks of any shape, such as arrows, can be used.
[0069] Next, the processing performed in the first embodiment will be described. Fig. 13 is a flowchart of the learning processing according to the first embodiment. It is assumed that the images and radiology reports used for learning are acquired by the information acquisition unit 21 from the image server 5 and the report server 6, respectively, and stored in the storage 13. It is also assumed that the condition for ending learning is that learning has been performed a predetermined number of times.
[0070] First, the first derivation unit 22 derives a first feature amount V1 for an object included in an image using the first neural network 61 (step ST1). The second derivation unit 23 derives a second feature amount V2 for a sentence including a description related to the object using the second neural network 62 (step ST2). Note that the processing of step ST2 may be performed first, or the processing of steps ST1 and ST2 may be performed in parallel.
[0071] Next, the attribute acquisition unit 24 acquires a first attribute P1, which is an attribute of an object included in the image, and a second attribute P2, which is an attribute of a sentence (attribute acquisition: step ST3).Then, the learning unit 25 trains the first neural network and the second neural network so that the higher the correlation between the combination of the first attribute P1 and the second attribute P2, the smaller the distance between the derived first feature amount V1 and second feature amount V2 (step ST4).The learning unit 25 further determines whether learning has been performed a predetermined number of times (predetermined number of times learning: step ST5).If step ST5 is negative, the process returns to step ST1 and repeats the processes of steps ST1 to ST5.If step ST5 is positive, the process ends.
[0072] Next, information processing according to the first embodiment will be described. Fig. 14 is a flowchart of information processing according to the first embodiment. It is assumed that a target medical image G0 to be processed has been acquired by the information acquisition unit 31 and stored in the storage 43. First, the first analysis unit 32 analyzes the target medical image G0 using the first derivation model 32A to derive a first feature value V1 for an object such as a lesion contained in the target medical image G0 (step ST11).
[0073] Next, the information acquisition unit 31 acquires the finding statement input by the radiologist using the input device 45 (step ST12), and the second analysis unit 33 analyzes the input finding statement using the second derivation model 33A to derive a second feature value V2 for the input finding statement (step ST13).
[0074] Next, the identification unit 34 derives the distance in feature space between the first feature amount V1 derived by the first analysis unit 32 and the second feature amount V2 derived by the second analysis unit 33, and identifies the first feature amount V1 corresponding to the second feature amount V2 based on the derived distance (step ST14).Then, the display control unit 35 displays the object from which the identified first feature amount V1 was derived in the target medical image G0, distinguishing it from other areas (step ST15), and the process ends.
[0075] In this way, in the learning device according to the first embodiment, the first derived model 32A and the second derived model 33A are constructed by training the first neural network 61 and the second neural network 62 so that the higher the correlation between the combination of the first attribute P1 and the second attribute P2 in the feature space to which the first feature V1 and the second feature V2 belong, the smaller the distance between the derived first feature V1 and the second feature V2.
[0076] Therefore, by applying the first derived model 32A and the second derived model 33A constructed by learning to the information processing device 30 according to the first embodiment, the first feature V1 and the second feature V2 are derived so that an image including an object with a high relevance in the combination of attributes and a sentence including a description of the object are associated with each other, and a medical image including an object with a low relevance in the combination of attributes and a sentence including a description of the object are not associated with each other. Therefore, by using the derived first feature V1 and second feature V2, it is possible to associate images and sentences with high accuracy.
[0077] In addition, since it is possible to accurately match images with text, when creating an interpretation report for a medical image, it is possible to accurately identify objects described in the input findings text in the medical image.
[0078] In the learning device according to the first embodiment, the attribute acquisition unit 24 acquires the position of an object as the first attribute P1 and the second attribute P2, but this is not limited to this. The property of an object included in an image may be acquired as the first attribute P1, and the property of an object described in a sentence may be acquired as the second attribute P2. This will be described below as the second embodiment. The configuration of the learning device according to the second embodiment is the same as the configuration of the learning device according to the first embodiment, and only the processing performed is different, so a description of the device configuration will be omitted.
[0079] In the second embodiment, the learning unit 25 of the learning device 7 determines the relevance of a combination of a first attribute P1 and a second attribute P2 based on co-occurrence relationships in texts expressing characteristics. In the field of natural language processing, co-occurrence refers to the simultaneous appearance of two character strings in a given document or sentence. For example, the word "irregular" is often found in a finding together with the word "sawtooth-shaped," but is rarely found in a finding together with the word "near-circular." Therefore, "irregular" and "sawtooth-shaped" co-occur more frequently, while "irregular" and "near-circular" rarely co-occur.
[0080] In the second embodiment, co-occurrence relationships between words included in multiple finding sentences are derived in advance and stored in the storage 43. The co-occurrence relationships may be derived manually or by analyzing multiple finding sentences. In this case, a trained model that has been trained to derive co-occurrence relationships from finding sentences may be used, or the co-occurrence relationships may be derived based on a rule base.
[0081] Fig. 15 is a diagram showing examples of co-occurrence relationships. In Fig. 15, the magnitude of the co-occurrence relationship is expressed as a numerical value between 0 and 1. As shown in Fig. 15, the co-occurrence relationship between "irregular" and "sawtooth" is 0.96, and the co-occurrence relationship between "irregular" and "near-circular" is 0.10.
[0082] In the second embodiment, the attribute acquisition unit 24 acquires the property of an object 55A included in an image 55 as a first attribute P1, and acquires the property of the object described in a sentence 56 as a second attribute P2. In the second embodiment, the learning unit 25 trains the first and second neural networks 61, 62 so that the distance in the feature space between the first feature amount V1 and the second feature amount V2 decreases as the correlation between the combination of the first attribute P1, which is the property of the object, and the second attribute increases.
[0083] 16 is a diagram schematically illustrating the processing performed by the first derivation unit 22, the second derivation unit 23, the attribute acquisition unit 24, and the learning unit 25 in the second embodiment. In FIG. 16, the same components as those in FIG. 6 are given the same reference numerals, and detailed descriptions thereof will be omitted. In the second embodiment, the same image 55 as in the first embodiment is used for training the first neural network 61. Meanwhile, the second neural network 62 is trained using a sentence 58 stating, "A nodule with an unclear boundary and saw-toothed margins is present in the right lower lobe S6," and a sentence 59 stating, "A micronodule with a near-circular shape is also present in the left lung S9."
[0084] In the second embodiment, the first derivation unit 22, similar to the first embodiment, extracts an object 55A such as a lesion contained in an image 55 using a first neural network 61, and derives the feature of the object 55A as a first feature V1.
[0085] The second derivation unit 23 uses the second neural network 62 to derive features for the sentences 58 and 59 as second features V2-3 and V2-4, respectively.
[0086] In the second embodiment, the attribute acquisition unit 24 acquires the property of an object 55A included in an image 55 as a first attribute P1-1, and acquires the properties of the objects described in sentences 58 and 59 as second attributes P2-3 and P2-4, respectively. In the second embodiment, it is assumed that "indistinct boundary" and "irregular shape" are acquired as the first attribute P1-1, "indistinct boundary" and "saw-tooth" are acquired as the second attribute P2-3, and "near-circular" is acquired as the second attribute P2-4.
[0087] In the second embodiment, the learning unit 25 derives the association between the combination of the first attribute P1-1 and each of the second attributes P2-3 and P2-4 based on the co-occurrence relationship. In the second embodiment, the first attribute P1-1 is “unclear boundary” and “irregular,” and the second attribute P2-3 is “unclear boundary” and “sawtooth.” The first attribute P1-1 and the second attribute P2-3 share the property of “unclear boundary.” When multiple properties are included as attributes, the learning unit 25 derives the association between the combination using different properties other than the common property. Therefore, when deriving the association between the combination of the first attribute P1-1 and the second attribute P2-3, the learning unit 25 derives the association between the combination of “irregular” and “sawtooth.” Here, as shown in FIG. 15 , the co-occurrence relationship between “irregular” and “sawtooth” is 0.96.
[0088] On the other hand, when deriving the association between the combination of the first attribute P1-1 and the second attribute P2-4, the learning unit 25 derives the association between the combination based on the smaller of the co-occurrence relationship between "unclear boundary" and "nearly circular" and the co-occurrence relationship between "irregular" and "nearly circular." As shown in FIG. 15 , the value of the co-occurrence relationship between "irregular" and "nearly circular" is 0.10, and the value of the co-occurrence relationship between "unclear boundary" and "nearly circular" is 0.30. In this case, the learning unit 25 derives the association between the combination of the first attribute P1-1 and the second attribute P2-4 based on the value of the co-occurrence relationship between "irregular" and "nearly circular." Note that the association between the combination may also be derived based on the larger of the co-occurrence relationship between "unclear boundary" and "nearly circular" and the co-occurrence relationship between "irregular" and "nearly circular." Furthermore, the relevance of the combination may be derived based on the average values of the co-occurrence relationship between "indistinct boundary" and "nearly circular" and the co-occurrence relationship between "irregular shape" and "nearly circular".
[0089] In the second embodiment, the learning unit 25 determines the relevance of the combination of the first attribute P1 and the second attribute P2 to be "high" when the co-occurrence relationship is equal to or greater than a predetermined threshold. If the co-occurrence relationship is less than the threshold, the learning unit 25 determines the relevance of the combination of the first attribute P1 and the second attribute P2 to be "low." For example, 0.7 can be used as the threshold. Therefore, the learning unit 25 determines the relevance of the combination of the first attribute P1-1 and the second attribute P2-3 to be "high," and the relevance of the combination of the first attribute P1-1 and the second attribute P2-4 to be "low."
[0090] The relevance is not limited to two types, "high" and "low," and three or more types of relevance may be determined. For example, the relevance may be determined as "high," "slightly high," or "low" depending on the value of the position represented by the first attribute P1 and the co-occurrence relationship represented by the second attribute P2.
[0091] In the second embodiment, the learning unit 25 plots the first feature amount V1 and the second feature amounts V2-3 and V2-4 in a feature space. If the relevance of the combination of the first attribute P1 and the second attribute P2 is "high," the learning unit 25 trains the first neural network 61 and the second neural network 62 so that the first feature amount V1 and the second feature amount V2 become closer to each other in the feature space. On the other hand, if the relevance of the combination of the first attribute P1 and the second attribute P2 is "low," the learning unit 25 trains the first neural network 61 and the second neural network 62 so that the first feature amount V1 and the second feature amount V2 become farther apart in the feature space.
[0092] Here, the relevance of the combination of the first attribute P1-1 and the second attribute P2-3 is "high." On the other hand, the relevance of the combination of the first attribute P1 and the second attribute P2-4 is "low." Therefore, in the feature space shown in FIG. 16, the learning unit 25 trains the first neural network 61 and the second neural network 62 so that the first feature amount V1 and the second feature amount V2-3 become closer to each other and the first feature amount V1 and the second feature amount V2-4 become farther apart. The derivation of the loss used in the learning is the same as in the first embodiment, and therefore a detailed description thereof will be omitted here.
[0093] In the first and second embodiments, the first and second neural networks 61, 62 are trained based on the association between the combination of the first attribute P1 and the second attribute P2, but this is not limited to this. In addition to the association between the combination of the first attribute P1 and the second attribute P2, the first and second neural networks 61, 62 may be trained using a combination of a medical image and a radiology report corresponding to the medical image as training data. This will be described below as a third embodiment.
[0094] In the third embodiment, a combination of a medical image and a radiology report corresponding to the medical image is used as training data to train the first and second neural networks 61 and 62. FIG. 17 is a diagram showing the training data used in the third embodiment. As shown in FIG. 17, the training data 80 consists of a training image 81 and a radiology report 82 that corresponds to the training image 81 and contains a statement of findings about the training image 81. The training image 81 is a tomographic image of the lung and includes a lesion as an object in the right lower lobe S6.
[0095] The radiology report 82 includes three findings 82A to 82C. The finding 82A is "A 12 mm solid nodule is present in the right lower lobe S6." The finding 82B is "A 5 mm GGN (ground-glass nodule) is present in the right lung S7." The finding 82C is "A micronodule is present in the left lung S9." Here, since the training image 81 includes an object in the right lower lobe S6, among the three findings 82A to 82C, the finding 82A corresponds to the training image 81. Note that the training image 81 is one of the multiple tomographic images that make up the three-dimensional image. The findings 82B and 82C were generated as a result of interpreting tomographic images other than the training image 81. Therefore, the training image 81 does not correspond to the findings 82B and 82C.
[0096] In the third embodiment, the first derivation unit 22 derives the feature amount of the teacher image 81 as the first feature amount V1-3 using the first neural network 61. Furthermore, the second derivation unit 23 derives the feature amounts of the finding statements 82A to 82C as the second feature amounts V2-5, V2-6, and V2-7 using the second neural network 62.
[0097] In the third embodiment, the attribute acquisition unit 24 acquires an attribute of an object included in the teacher image 81 as a first attribute P1-3. Also, attributes of objects described in the finding statements 82A to 82C are acquired as second attributes P2-5, P2-6, and P2-7. In the third embodiment, the first attribute P1-3 represents the position of an object included in the teacher image 81, and the second attributes P2-5, P2-6, and P2-7 represent the positions of objects described in the finding statements 82A to 82C. Therefore, the first attribute P1-3 is "right lower lobe S6," the second attribute P2-5 is "right lower lobe S6," the second attribute P2-6 is "right lung S7," and the second attribute P2-7 is "left lung S9." The first attribute P1-3 may be the property of the object included in the teacher image 81, and the second attributes P2-5, P2-6, and P2-7 may be the property of the object described in the observation statements 82A to 82C.
[0098] The learning unit 25 plots the first feature amount V1-3 and the second feature amounts V2-5, V2-6, and V2-7 in the feature space. FIG. 18 is a diagram for explaining the plotting of the feature amounts in the third embodiment. Note that, for the sake of explanation, the feature space is also shown in two dimensions in FIG. 18. Then, as in the first and second embodiments, the learning unit 25 trains the first and second neural networks 61 and 62 so that the distance between the first feature amount and the second feature amount in the feature space decreases as the correlation between the combination of the first attribute P1 and the second attribute P2 increases.
[0099] In the third embodiment, the first attribute P1-3 is "right lower lobe S6," the second attribute P2-5 is "right lower lobe S6," the second attribute P2-6 is "right lung S7," and the second attribute P2-7 is "left lung S9." Therefore, the first attribute P1-3 and the second attribute P2-5 are highly related, and the first attribute P1-3 and the second attributes P2-6 and P2-7 are less related.
[0100] Therefore, the learning unit 25 learns the first neural network 61 and the second neural network 62 so that the first feature amount V1-3 and the second feature amount V2-5 are closer to each other and the first feature amount V1-3 and the second feature amounts V2-6 and V2-7 are farther apart in the feature space.
[0101] Furthermore, in the third embodiment, the first and second neural networks 61 and 62 are trained based on the correspondence between the teacher image 81 and the findings in the teacher data 80. That is, the training unit 25 further trains the first and second neural networks 61 and 62 so that, in the feature space, the first feature V1-3 and the second feature V2-5 become closer to each other and the first feature V1-3 and the second feature V2-6 and V2-7 become farther apart. In this case, the training unit 25 may train the first and second neural networks 61 and 62 so that the degree of separation between the first feature V1-3 and the second feature V2-6 becomes smaller than the degree of separation between the first feature V1-3 and the second feature V2-7. Note that either the learning based on the association between the combination of the first attribute P1 and the second attribute P2 or the learning based on the correspondence between the teacher image and the findings in the teacher data may be performed first.
[0102] In this way, in the third embodiment, learning is further performed based on the correspondence between the training images and the findings in the training data, which allows the first and second neural networks 61 and 62 to be trained with higher accuracy, and as a result, the first derived model 32A and the second derived model 33A can be constructed, which can derive features with higher accuracy.
[0103] Next, a second embodiment of the information processing device will be described. Fig. 19 is a functional configuration diagram of the information processing device according to the second embodiment. In Fig. 19, the same components as those in Fig. 10 are given the same reference numerals, and detailed description will be omitted. As shown in Fig. 19, the information processing device 30A according to the second embodiment differs from the information processing device according to the first embodiment in that it includes a search unit 36 instead of the identification unit 34.
[0104] In the information processing device 30A according to the second embodiment, the information acquisition unit 31 acquires a large number of medical images stored in the image server 5. Then, the first analysis unit 32 derives a first feature V1 for each of the medical images. The information acquisition unit 31 transmits the first feature V1 to the image server 5. In the image server 5, the medical images are associated with the first feature V1 and stored in the image DB 5A. The medical images associated with the first feature V1 and registered in the image DB 5A will be referred to as reference images in the following description.
[0105] In the information processing device 30A according to the second embodiment, an interpretation doctor interprets the target medical image G0 in the interpretation WS3 and inputs a finding statement including the interpretation result using the input device 45, thereby generating an interpretation report. The second analysis unit 33 analyzes the input finding statement using the second derivation model 33A constructed by the learning device 7 described above, thereby deriving a second feature value V2 for the input finding statement.
[0106] The search unit 36 refers to the image DB 5 and searches in the feature space for a reference image associated with a first feature V1 that is close to the second feature V2 derived by the second analysis unit 33. Fig. 20 is a diagram for explaining the search performed in the information processing device 30A according to the second embodiment. Note that, for the sake of explanation, Fig. 20 also shows the feature space in two dimensions. For the sake of explanation, five first feature values V1-11 to V1-15 are plotted in the feature space.
[0107] The search unit 36 identifies a first feature value in the feature space whose distance from the second feature value V2 is within a predetermined threshold value. In Fig. 20, a circle 85 having a radius d5 and centered on the second feature value V2 is shown. The search unit 36 identifies a first feature value included in the circle 85 in the feature space. In Fig. 20, three first feature values V1-11 to V1-13 are identified.
[0108] The search unit 36 searches the image DB 5A for reference images associated with the identified first feature amounts V1-11 to V1-13, and acquires the searched reference images from the image server 5.
[0109] The display control unit 35 displays the acquired reference image on the display 44. FIG. 21 is a diagram showing a display screen in the information processing device 30A according to the second embodiment. As shown in FIG. 21, the display screen 90 has an image display area 91, a text display area 92, and a result display area 93. The target medical image G0 is displayed in the image display area 91. In FIG. 21, the target medical image G0 is one of the tomographic images that constitutes a three-dimensional image of the chest. The text display area 92 displays the finding entered by the radiologist. In FIG. 21, the finding reads, "There is a 10 mm solid nodule in the right lung S6."
[0110] The result display area 93 displays the reference images searched by the search unit 36. In FIG.
[0111] Next, information processing according to the second embodiment will be described. FIG. 22 is a flowchart of information processing according to the second embodiment. It is assumed that the first feature amount for the reference image is derived by the first analysis unit 32 and is associated with the reference image and registered in large numbers in the image DB 5A. It is also assumed that the target medical image G0 is displayed on the display 44 by the display control unit 35. In the second embodiment, the information acquisition unit 31 acquires a finding statement input by the radiologist using the input device 45 (step ST21), and the second analysis unit 33 analyzes the input finding statement using the second derivation model 33A to derive a second feature amount V2 for the object described in the input finding statement (step ST22).
[0112] Next, the search unit 36 refers to the image DB 5A to search for a reference image associated with the first feature V1 that is close to the second feature V2 (step ST23).Then, the display control unit 35 displays the searched reference image on the display 44 (step ST24), and the process ends.
[0113] The reference images R1 to R3 retrieved in the second embodiment are medical images with similar features to the findings entered by the radiologist. Since the findings relate to the target medical image G0, the reference images R1 to R3 are similar in case to the target medical image G0. Therefore, according to the second embodiment, the target medical image G0 can be interpreted by referring to the reference images with similar cases. Furthermore, the radiology report for the reference images can be obtained from the report server 6 and used to create the radiology report for the target medical image G0.
[0114] In the first embodiment, the first attribute P1 and the second attribute P2 represent the position of an object, and in the second embodiment, the first attribute P1 and the second attribute P2 represent the characteristics of the object. However, this is not limiting. The first attribute P1 may be both the position and the characteristics of the object 55A included in the image 55, and the second attribute P2 may be both the position and the characteristics of the object described in the sentence 56. In this case, the first and second neural networks 61 and 62 are trained so that, in the feature space to which the first and second features belong, the closer the correlation between the combination of the first attribute P1 and the second attribute P2 with respect to the position, the smaller the distance between the derived first feature and the second feature. Furthermore, the first and second neural networks 61 and 62 are trained so that, in the feature space to which the first and second features belong, the closer the correlation between the combination of the first attribute P1 and the second attribute P2 with respect to the characteristics, the smaller the distance between the derived first feature and the second feature. This makes it possible to construct the first derived model 32A and the second derived model 33A with higher accuracy.
[0115] In addition, both the position and characteristics of object 55A contained in image 55 may be used as first attribute P1, and both descriptions representing the position and characteristics of object 55A contained in sentence 56 may be used as second attribute P2. In addition, as in the third embodiment, the first and second neural networks 61, 62 may be trained using a combination of a medical image and a radiology report corresponding to the medical image as training data.
[0116] In the above embodiment, the first and second neural networks 61 and 62 are trained so that the first feature V1 and the second feature V2 are closer to or farther apart in the feature space depending on the degree of correlation between the combination of the first attribute P1 and the second attribute P2. However, this is not limiting. For example, suppose the first attribute P1 is the right lung S6 and the second attribute P1 is the right lung S7. The right lung S6 and the right lung S7 are not located at the same position, but are anatomically adjacent to each other. In such a case, the first and second neural networks 61 and 62 may be trained so that the first feature V1 and the second feature V2 are not too close to each other, but are not too far apart.
[0117] In this case, the learning unit 25 trains the first and second neural networks 61, 62 so that the distance between the first feature amount V1 and the second feature amount V2 in the feature space is equal to or greater than the threshold value α0 and equal to or less than the threshold value β0 shown in FIG. 9. That is, the learning unit 25 calculates the loss L3 using the following formula (3) for each of three cases: when the distance d6 between the first feature amount V1 and the second feature amount V2 in the feature space is less than the threshold value α0, when it is greater than the threshold value β0, and when it is equal to or greater than α0 and equal to or less than β0. The first and second neural networks 61, 62 are then trained so that the calculated loss L3 is small. This makes it possible to construct the first derivation model 32A and the second derivation model 33A so that the first feature amount V1 and the second feature amount V2 can be derived while maintaining an appropriate distance in the feature space, even when the relevance of the attribute combination is neither too high nor too low. L3=α0-d6(d6<α0) L3=d6-β0(d6>β0) (3) L3=0(α0≦d6≦β0)
[0118] In the above embodiment, a derivation model is constructed to derive feature quantities of medical images and commentary on medical images, but the present disclosure is not limited to this. For example, the technology of the present disclosure can also be applied to constructing a derivation model to derive feature quantities of photographic images and commentary on photographic images.
[0119] Furthermore, in the above embodiment, the following various processors can be used as the hardware structure of processing units that perform various processes, such as the information acquisition unit 21, first derivation unit 22, second derivation unit 23, attribute acquisition unit 24, and learning unit 25 in the learning device 7, and the information acquisition unit 31, first analysis unit 32, second analysis unit 33, identification unit 34, display control unit 35, and search unit 36 in the information processing devices 30 and 30A. As described above, the various processors include a CPU, which is a general-purpose processor that executes software (programs) and functions as various processing units, as well as dedicated electrical circuits that are processors having a circuit configuration specifically designed to perform specific processes, such as a programmable logic device (PLD), a processor whose circuit configuration can be changed after manufacture, such as an FPGA (Field Programmable Gate Array), and an application-specific integrated circuit (ASIC).
[0120] A single processing unit may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs or a combination of a CPU and an FPGA). Also, multiple processing units may be configured with a single processor. Examples of multiple processing units configured with a single processor include, first, a configuration in which one processor is configured with a combination of one or more CPUs and software, as typified by client and server computers, and this processor functions as multiple processing units. Second, a configuration in which a processor is used to realize the functions of an entire system including multiple processing units on a single IC (Integrated Circuit) chip, as typified by a System on Chip (SoC). In this way, the various processing units are configured with one or more of the above-mentioned various processors as a hardware structure.
[0121] Furthermore, more specifically, the hardware structure of these various processors can be an electric circuit that combines circuit elements such as semiconductor elements. [Explanation of symbols]
[0122] 1 Medical Information System 2. Imaging equipment 3 Image Reading Workshop 4. Medical Workshop 5 Image Server 5A Image DB 6 Report Server 6A Report DB 7 Learning Device 10 Network 11,41 CPU 12 Study Programs 13,43 Storage 14,44 display 15,45 Input Devices 16,46 memory 17,47 Network I / F 18,48 bus 21 Information Acquisition Department 22 First derivation part 23 Second derived part 24 Attribute acquisition part 25 Learning Department 30, 30A Information processing equipment 31 Information Acquisition Department 32 1st Analysis Department 32A First Derived Model 33 2nd Analysis Department 33A Second Derived Model 34 Specific part 35 Display control unit 36 Search Section 42 Information Processing Program 51 Medical Imaging 51A Tomographic Image 52 Image interpretation report 53 Observations 55 images 55A Object 56~59 sentences 61 First Neural Network 62 Second Neural Network 62A Buried Layer 62B RNN layer 62C fully connected layer 65,66 feature vector 70 Image interpretation report creation screen 71,91 Image display area 72,92 Text display area 73,74 Lesions 75 marks 80 training data 81 Teacher Images 82,82A,82B,82C Observations 90 display screen 93 Reference image display area d3,d4 distance G0 Target Medical Images P1,P1-1,P1-3 1st attribute P2-1,P2-2,P2-3,P2-4,P2-5,P2-6,P2-7 2nd attribute R1, R2, R3 Reference Image V1,V1-1,V1-2,V1-3,V1-11~V1-15 1st special quantity V2,V2-1,V2-2,V2-3,V2-4,V2-5,V2-6,V2-7 2nd special quantity
Claims
1. at least one processor; The processor: deriving a first feature amount for an object included in an image using a first neural network; deriving second features for sentences including descriptions relating to the object using a second neural network; acquiring a first attribute that is an attribute of an object included in the image and a second attribute that is an attribute of the sentence; A learning device that constructs a first derivation model that derives features for an object included in an image and a second derivation model that derives features for a sentence that includes a description of the object, by training the first neural network and the second neural network so that the distance between the derived first feature and the second feature is smaller when the correlation between the combination of the first attribute and the second attribute is higher than when the correlation between the combination of the first attribute and the second attribute is low in a feature space to which the first feature and the second feature belong.
2. The learning device according to claim 1 , wherein the first attribute is a position of an object in the image, and the second attribute is a position of an object described in the sentence.
3. 3. The learning device according to claim 1, wherein the first attribute is a property of an object included in the image, and the second attribute is a property of an object described in the sentence.
4. The learning device according to claim 3 , wherein the processor acquires the first attribute and the second attribute based on a co-occurrence relationship in text that expresses a property.
5. 5. The learning device of claim 1, wherein the processor trains the first neural network and the second neural network so that, when an object included in the image corresponds to an object described in the sentence, the distance between the derived first feature and the derived second feature in the feature space becomes small, and when the object included in the image does not correspond to an object described in the sentence, the processor trains the first neural network and the second neural network so that the distance between the derived first feature and the derived second feature in the feature space becomes large.
6. the image is a medical image, and the object included in the image is a lesion included in the medical image; The learning device according to claim 1 , wherein the sentence is a finding sentence that describes a finding about a lesion.
7. at least one processor; The processor: deriving first feature amounts for one or more objects included in a target image using a first derived model constructed by the learning device according to any one of claims 1 to 6; deriving second feature quantities for one or more target sentences including descriptions relating to an object using a second derived model constructed by the learning device according to any one of claims 1 to 6; identifying the first feature amount corresponding to the second feature amount based on the distance in a feature space between the derived first feature amount and the derived second feature amount; an information processing device that displays the object from which the specified first feature amount is derived in the target image in a manner that distinguishes it from other areas;
8. at least one processor; The processor: Accepting input of a target sentence including a description of an object; deriving a second feature amount for the input target sentence using a second derived model constructed by the learning device according to any one of claims 1 to 6; a database in which first feature quantities for one or more objects included in each of a plurality of reference images, derived by a first derivation model constructed by the learning device according to any one of claims 1 to 6, are associated with each of the reference images, and at least one of the first feature quantities corresponding to the second feature quantities is identified based on a distance in a feature space between the first feature quantities for the plurality of reference images and the derived second feature quantities; An information processing device that identifies a reference image associated with the identified first feature amount.
9. deriving a first feature amount for an object included in an image using a first neural network; deriving second features for sentences including descriptions relating to the object using a second neural network; acquiring a first attribute that is an attribute of an object included in the image and a second attribute that is an attribute of the sentence; A learning method for constructing a first derivation model that derives features for an object included in an image and a second derivation model that derives features for a sentence that includes a description of the object, by training the first neural network and the second neural network so that the distance between the derived first feature and the derived second feature is smaller when the correlation between the combination of the first attribute and the second attribute is higher than when the correlation between the combination of the first attribute and the second attribute is low in a feature space to which the first feature and the second feature belong.
10. deriving first feature amounts for one or more objects included in a target image using a first derived model constructed by the learning method according to claim 9; deriving second feature amounts for one or more target sentences including descriptions relating to the object using a second derived model constructed by the learning method according to claim 9; identifying the first feature amount corresponding to the second feature amount based on the distance in a feature space between the derived first feature amount and the derived second feature amount; An information processing method for displaying the object from which the specified first feature amount has been derived in the target image in a manner that distinguishes it from other areas.
11. Accepting input of a target sentence including a description of an object; deriving second features for the input target sentence using a second derived model constructed by the learning method according to claim 9; a database in which first feature quantities for one or more objects included in each of a plurality of reference images, derived by a first derivation model constructed by the learning method according to claim 9, are associated with each of the reference images, and at least one of the first feature quantities corresponding to the second feature quantities is identified based on a distance in a feature space between the first feature quantities for the plurality of reference images and the derived second feature quantities; An information processing method for identifying a reference image associated with the identified first feature amount.
12. deriving a first feature amount for an object included in an image by a first neural network; deriving second features for sentences including descriptions related to the object by a second neural network; a step of acquiring a first attribute that is an attribute of an object included in the image and a second attribute that is an attribute of the sentence; a learning program that causes a computer to execute a procedure for constructing a first derivation model that derives features for an object included in an image and a second derivation model that derives features for a sentence that includes a description of the object, by training the first neural network and the second neural network so that the higher the correlation between the combination of the first attribute and the second attribute in a feature space to which the first feature and the second feature belong, the smaller the distance between the derived first feature and the derived second feature is compared to when the correlation between the combination of the first attribute and the second attribute is low.
13. a step of deriving a first feature amount for one or more objects included in a target image using a first derived model constructed by the learning program according to claim 12; deriving second feature quantities for one or more target sentences including descriptions relating to the object using a second derivation model constructed by the learning program according to claim 12; a step of identifying the first feature amount corresponding to the second feature amount based on a distance in a feature space between the derived first feature amount and the derived second feature amount; and displaying the object from which the specified first feature amount has been derived in the target image in a manner that distinguishes it from other areas.
14. receiving an input of a target sentence including a description of an object; a step of deriving a second feature amount for the input target sentence using a second derivation model constructed by the learning program according to claim 12; a step of identifying at least one first feature quantity corresponding to a second feature quantity based on a distance in a feature space between the first feature quantity for the plurality of reference images and the derived second feature quantity, by referring to a database in which the first feature quantity for one or more objects included in each of the plurality of reference images, derived by a first derivation model constructed by the learning program according to claim 12, is associated with each of the reference images; and specifying a reference image associated with the specified first feature amount.