Learning device, method, and program, and information processing device, method, and program
The learning device uses neural networks to accurately associate medical images with text by deriving feature amounts, identifying candidate objects, and training networks to minimize attribute differences, effectively linking medical images with text descriptions for precise identification and display.
Patent Information
- Application Number
- JP2021140166
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-08-30
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2041-08-30
AI Technical Summary
The limited availability of medical images and sentences makes it challenging to construct a trained model that accurately associates images with text, particularly in the field of medical image analysis.
A learning device that uses neural networks to derive feature amounts for images and text, identifies candidate objects for pairing, estimates attributes, and trains the networks to minimize the difference between estimated and actual attributes, thereby constructing models that accurately associate medical images with text descriptions.
Enables highly accurate correspondence between medical images and text, allowing for precise identification and display of image features and their corresponding text descriptions.
Smart Images

Figure 0007718915000001 
Figure 0007718915000002 
Figure 0007718915000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a learning device, a method, and a program, and an information processing device, a method, and a program. [Background technology]
[0002] A method has been proposed for constructing a feature space to which features such as feature vectors extracted from images belong using a trained model that has undergone machine learning such as deep learning. For example, Non-Patent Document 1 proposes a method for training a trained model so that features of images belonging to the same class are close to each other in the feature space, and features of images belonging to different classes are far apart in the feature space. Also known is a technique for associating features extracted from images with features extracted from sentences based on the distance in the feature space (see Non-Patent Document 2). [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Deep metric learning using Triplet network, Elad Hoffer et al., 20 Dec 2014, arXiv:1412.6622 [Non-patent document 2] Learning Two-Branch Neural Networks for Image-Text Matching Tasks, Liwei Wang et al., 11 Apr 2017, arXiv:1704.03470 Summary of the Invention [Problem to be solved by the invention]
[0004] As described in Non-Patent Document 2, in order to accurately train a trained model that associates images with sentences, a large amount of training data that associates features contained in images with features described in sentences is required. However, since the number of images and sentences is limited, it may be difficult to prepare a large amount of training data. In particular, in the field of medical image analysis, since the number of medical images is limited, it is difficult to construct a trained model that can associate images with sentences with high accuracy.
[0005] The present disclosure has been made in consideration of the above circumstances, and aims to enable highly accurate correspondence between images and text. [Means for solving the problem]
[0006] A learning device according to the present disclosure includes at least one processor, The processor derives, using a first neural network, a first feature amount for each of a plurality of first objects included in the first data (e.g., image data); deriving second features for second data (e.g., text data) including one or more second objects using a second neural network; Identifying a first object candidate to be paired with the second object from among the plurality of first objects; Estimating attributes of a second object paired with the first object candidate based on the first object candidate; At least one of the first neural network and the second neural network is trained so that the difference between the estimated attribute of the paired second object and the attribute of the paired second object derived from the second data is reduced, thereby constructing at least one of a first derived model that derives features for the object included in the first data and a second derived model that derives features for the second data including the object.
[0007] In addition, in the learning device according to the present disclosure, the processor may identify the first object candidate based on the distance between the first feature amount and the second feature amount in a feature space to which the first feature amount and the second feature amount belong.
[0008] In the learning device according to the present disclosure, the processor may identify the first object candidate based on the degree of association between the first feature amount and the second feature amount.
[0009] In the learning device according to the present disclosure, the processor may estimate an attribute of the paired second object from the first feature amount for the first object candidate.
[0010] In the learning device according to the present disclosure, the processor may derive an attribute of the paired second object from the first data.
[0011] In addition, in the learning device according to the present disclosure, when multiple first object candidates are identified, the processor may estimate the attributes of the paired second object from the sum or weighted sum of the first features for the multiple first object candidates.
[0012] In addition, in the learning device according to the present disclosure, when multiple first object candidates are identified, the processor may estimate the attributes of the paired second object from the sum or weighted sum of the multiple first object candidates in the first data.
[0013] In addition, in the learning device according to the present disclosure, the processor may further train at least one of the first neural network and the second neural network so that, when the first object and the second object correspond to each other, the distance between the derived first feature and the derived second feature becomes smaller in the feature space to which the first feature and the second feature belong, and, when the first object and the second object do not correspond to each other, the processor may further train at least one of the first neural network and the second neural network so that the distance between the derived first feature and the derived second feature becomes larger in the feature space.
[0014] In addition, in the learning device according to the present disclosure, the first data is image data representing an image including a first object, The second data may be text data representing a sentence including a description of the second object.
[0015] In addition, in the learning device according to the present disclosure, the image is a medical image, and the first object included in the image is a lesion included in the medical image, The sentence may be a finding sentence about a medical image, and the second object may be a finding about a lesion in the sentence.
[0016] A first information processing device according to the present disclosure includes at least one processor, The processor derives first feature amounts for one or more first objects included in the target image using a first derived model constructed by the learning device according to the present disclosure; deriving second features for one or more target sentences including a description related to a second object using a second derived model constructed by the learning device according to the present disclosure; Identifying a first feature corresponding to the second feature based on a distance in a feature space between the derived first feature and the derived second feature; The first object from which the specified first feature amount has been derived is displayed in the target image in a manner distinguished from other areas.
[0017] In the first information processing device according to the present disclosure, the processor estimates an attribute of a second object paired with the identified first object based on the first object from which the identified first feature amount was derived; The estimated attributes may also be displayed.
[0018] In addition, in the first information processing device according to the present disclosure, the processor derives an attribute of a second object described in the target sentence; The target sentence may be displayed with descriptions regarding derived attributes that differ from estimated attributes in the target sentence distinguished from other descriptions.
[0019] A second information processing device according to the present disclosure includes at least one processor, The processor accepts an input of a target sentence including a description of a second object; deriving second features for the input target sentence using a second derivation model constructed by the learning device according to the present disclosure; By referring to a database in which first feature amounts for one or more first objects included in each of a plurality of reference images, which are derived by a first derivation model constructed by the learning device according to the present disclosure, are associated with each of the reference images, at least one first feature amount corresponding to the second feature amount is identified based on the distance in feature space between the first feature amount for the plurality of reference images and the derived second feature amount; A reference image associated with the identified first feature amount is identified.
[0020] A learning method according to the present disclosure includes: deriving, by a first neural network, a first feature amount for each of a plurality of first objects included in first data; deriving second features for second data including one or more second objects using a second neural network; Identifying a first object candidate to be paired with the second object from among the plurality of first objects; Estimating attributes of a second object paired with the first object candidate based on the first object candidate; At least one of the first neural network and the second neural network is trained so that the difference between the estimated attribute of the paired second object and the attribute of the paired second object derived from the second data is reduced, thereby constructing at least one of a first derived model that derives features for the object included in the first data and a second derived model that derives features for the second data including the object.
[0021] A first information processing method according to the present disclosure includes deriving first feature amounts for one or more first objects included in a target image using a first derived model constructed by a learning device according to the present disclosure; deriving second features for one or more target sentences including a description related to a second object using a second derived model constructed by the learning device according to the present disclosure; Identifying a first feature corresponding to the second feature based on a distance in a feature space between the derived first feature and the derived second feature; The first object from which the specified first feature amount has been derived is displayed in the target image in a manner distinguished from other areas.
[0022] A second information processing method according to the present disclosure includes accepting input of a target sentence including a description related to a second object; deriving second features for the input target sentence using a second derivation model constructed by the learning device according to the present disclosure; By referring to a database in which first feature amounts for one or more first objects included in each of a plurality of reference images, which are derived by a first derivation model constructed by the learning device according to the present disclosure, are associated with each of the reference images, at least one first feature amount corresponding to the second feature amount is identified based on the distance in feature space between the first feature amount for the plurality of reference images and the derived second feature amount; A reference image associated with the identified first feature amount is identified.
[0023] The learning method and the first and second information processing methods according to the present disclosure may be provided as a program for causing a computer to execute the method. [Effects of the Invention]
[0024] According to the present disclosure, images and text can be associated with high accuracy. [Brief explanation of the drawings]
[0025] [Figure 1] FIG. 1 is a diagram showing a schematic configuration of a medical information system to which a learning device and an information processing device according to a first embodiment of the present disclosure are applied. [Figure 2] FIG. 1 is a diagram showing a schematic configuration of a learning device according to a first embodiment; [Figure 3] FIG. 1 is a diagram showing a schematic configuration of an information processing apparatus according to a first embodiment; [Figure 4] Functional configuration diagram of a learning device according to a first embodiment [Figure 5] Diagram showing an example of a medical image and interpretation report [Figure 6] FIG. 1 is a diagram illustrating the processes performed by a first derivation unit, a second derivation unit, an attribute acquisition unit, and a learning unit in a first embodiment. [Figure 7] Schematic diagram of the second neural network [Figure 8] Functional configuration diagram of an information processing device according to a first embodiment [Figure 9] FIG. 1 is a diagram illustrating identification of a first feature amount. [Figure 10] Image showing the image interpretation report creation screen [Figure 11] 1 is a flowchart showing a learning process performed in the first embodiment. [Figure 12] 1 is a flowchart showing information processing performed in the first embodiment. [Figure 13] FIG. 10 is a diagram showing teacher data used in the learning device of the second embodiment. [Figure 14] FIG. 10 is a diagram illustrating plots of feature amounts in the second embodiment. [Figure 15]Functional configuration diagram of an information processing device according to a second embodiment [Figure 16] Diagram to explain search [Figure 17] Diagram showing the display screen [Figure 18] 10 is a flowchart showing information processing performed in the second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0026] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. First, the configuration of a medical information system to which a learning device and an information processing device according to a first embodiment of the present disclosure are applied will be described. FIG. 1 is a diagram showing a schematic configuration of a medical information system 1. The medical information system 1 shown in FIG. 1 is a system for capturing an examination target region of a patient as a subject, storing the medical images acquired by capturing the images, having a radiologist interpret the medical images and create an interpretation report, and allowing the requesting doctor from the medical department to view the interpretation report and observe the details of the medical image to be interpreted, based on an examination order from a doctor from a medical department using a known ordering system.
[0027] As shown in Figure 1, the medical information system 1 is configured by connecting multiple imaging devices 2, multiple interpretation WSs (Workstations) 3 which are interpretation terminals, a medical treatment WS 4, an image server 5, an image DB (DataBase) 5A, a report server 6, a report DB 6A, and a learning device 7 in a state where they can communicate with each other via a wired or wireless network 10.
[0028] Each device is a computer installed with an application program that causes the device to function as a component of the medical information system 1. The application program is stored in an externally accessible state in a storage device of a server computer connected to the network 10 or in network storage, and is downloaded and installed into the computer upon request. Alternatively, the application program is recorded on a recording medium such as a DVD (Digital Versatile Disc) or CD-ROM (Compact Disc Read Only Memory) and distributed, and then installed into the computer from the recording medium.
[0029] The imaging device 2 is a device (modality) that captures an image of a patient's diagnostic target area to generate a medical image representing the diagnostic target area. Specifically, it is a plain X-ray imaging device, a CT (Computed Tomography) device, an MRI (Magnetic Resonance Imaging) device, a PET (Positron Emission Tomography) device, etc. The medical image generated by the imaging device 2 is transmitted to the image server 5 and stored in the image DB 5A.
[0030] The image interpretation WS3 is a computer used by, for example, a radiologist to interpret medical images and create image interpretation reports, and includes an information processing device according to this embodiment (details of which will be described later). The image interpretation WS3 issues requests for viewing medical images to the image server 5, performs various image processing on medical images received from the image server 5, displays the medical images, and accepts input of findings related to the medical images. The image interpretation WS3 also performs analysis processing on medical images, supports the creation of image interpretation reports based on the analysis results, requests the report server 6 to register and view image interpretation reports, and displays image interpretation reports received from the report server 6. These processes are performed by the image interpretation WS3 executing software programs for each process.
[0031] The medical treatment WS4 is a computer used by, for example, a doctor in a medical department to observe images in detail, view radiology reports, and create electronic medical records, and is composed of a processing device, a display device such as a monitor, and input devices such as a keyboard and a mouse. The medical treatment WS4 sends image viewing requests to the image server 5, displays images received from the image server 5, requests radiology reports to the report server 6, and displays radiology reports received from the report server 6. These processes are performed by the medical treatment WS4 executing software programs for each process.
[0032] The image server 5 is a general-purpose computer installed with a software program that provides the functions of a database management system (DBMS). The image server 5 also has storage in which the image DB 5A is configured. This storage may be a hard disk device connected to the image server 5 via a data bus, or a disk device connected to a NAS (Network Attached Storage) or SAN (Storage Area Network) connected to the network 10. When the image server 5 receives a request to register a medical image from the imaging device 2, it formats the medical image in a database format and registers it in the image DB 5A.
[0033] The image DB 5A stores image data and supplementary information of medical images acquired by the imaging device 2. The supplementary information includes, for example, an image ID (identification) for identifying each medical image, a patient ID for identifying the patient, an examination ID for identifying the examination, a unique ID (UID) assigned to each medical image, the examination date and time when the medical image was generated, the type of imaging device used in the examination to acquire the medical image, patient information such as the patient's name, age, and gender, the examination site (imaging site), imaging information (imaging protocol, imaging sequence, imaging technique, imaging conditions, use of contrast agent, etc.), and a series number or collection number when multiple medical images are acquired in one examination. In this embodiment, a first feature value of the medical image derived in the interpretation WS 3 as described below is associated with the medical image and registered in the image DB 5A.
[0034] In addition, when the image server 5 receives a viewing request from the image interpretation WS3 and medical treatment WS4 via the network 10, it searches for medical images registered in the image DB 5A and transmits the searched medical images to the image interpretation WS3 and medical treatment WS4 that made the request.
[0035] A software program that provides a general-purpose computer with the functions of a database management system is installed in the report server 6. When the report server 6 receives a request to register an interpretation report from the interpretation WS 3, it converts the interpretation report into a database format and registers it in the report DB 6A.
[0036] The report DB 6A stores a large number of radiology reports in a predetermined data format, each containing a finding created by a radiology physician using the radiology WS 3. The data of the radiology report includes text data representing the finding. The radiology report may include information such as the medical image to be interpreted, an image ID for identifying the medical image, a radiology physician ID for identifying the radiology physician who performed the interpretation, the name of the lesion, location information of the lesion, and characteristics of the lesion. In this embodiment, the report DB 6A stores radiology reports in association with one or more medical images for which the radiology reports were created.
[0037] In addition, when the report server 6 receives a request to view an interpretation report from the interpretation WS3 and medical treatment WS4 via the network 10, it searches for the interpretation report registered in the report DB6A and sends the retrieved interpretation report to the interpretation WS3 and medical treatment WS4 that made the request.
[0038] The network 10 is a wired or wireless local area network that connects various devices within the hospital. If the interpretation WS3 is installed in other hospitals or clinics, the network 10 may be configured by connecting the local area networks of each hospital via the Internet or a dedicated line.
[0039] Next, the learning device 7 will be described. First, the hardware configuration of the learning device 7 according to the first embodiment will be described with reference to FIG. 2. As shown in FIG. 2, the learning device 7 includes a CPU (Central Processing Unit) 11, non-volatile storage 13, and memory 16 as a temporary storage area. The learning device 7 also includes a display 14 such as an LCD display, an input device 15 including a keyboard, a pointing device such as a mouse, and a network I / F (Interface) 17 connected to a network 10. The CPU 11, storage 13, display 14, input device 15, memory 16, and network I / F 17 are connected to a bus 18. The CPU 11 is an example of a processor in the present disclosure.
[0040] The storage 13 is realized by a hard disk drive (HDD), a solid state drive (SSD), a flash memory, etc. The storage 13 as a storage medium stores a learning program 12. The CPU 11 reads the learning program 12 from the storage 13, expands it into the memory 16, and executes the expanded learning program 12.
[0041] Next, an information processing device 30 according to the first embodiment, which is included in the interpretation WS 3, will be described. First, the hardware configuration of the information processing device 30 according to this embodiment will be described with reference to FIG. 3. As shown in FIG. 3, the information processing device 30 includes a CPU 41, non-volatile storage 43, and memory 46 as a temporary storage area. The information processing device 30 also includes a display 44 such as a liquid crystal display, an input device 45 including a keyboard, a pointing device such as a mouse, and the like, and a network I / F 47 connected to the network 10. The CPU 41, storage 43, display 44, input device 45, memory 46, and network I / F 47 are connected to a bus 48. The CPU 41 is an example of a processor in the present disclosure.
[0042] Like the storage 13, the storage 43 is realized by an HDD, an SSD, a flash memory, or the like. The storage 43 as a storage medium stores an information processing program 42. The CPU 41 reads the information processing program 42 from the storage 43, loads it into the memory 46, and executes the loaded information processing program 42.
[0043] Next, the functional configuration of the learning device according to the first embodiment will be described. Fig. 4 is a diagram showing the functional configuration of the learning device according to the first embodiment. As shown in Fig. 4, the learning device 7 includes a first derivation unit 22, a second derivation unit 23, a candidate identification unit 24, an attribute estimation unit 25, and a learning unit 26. When the CPU 11 executes the learning program 12, the CPU 11 functions as the information acquisition unit 21, the first derivation unit 22, the second derivation unit 23, the candidate identification unit 24, the attribute estimation unit 25, and the learning unit 26.
[0044] The information acquisition unit 21 acquires medical images and interpretation reports from the image server 5 and the report server 6, respectively, via the network I / F 17. The medical images and interpretation reports are used for training the neural network, which will be described later. FIG. 5 shows examples of medical images and interpretation reports. As shown in FIG. 5, the medical image 51 is a three-dimensional image made up of multiple tomographic images. In this embodiment, the medical image 51 is a CT image of the chest of a human body. Furthermore, as shown in FIG. 5, the multiple tomographic images include a tomographic image 51A that includes multiple lesions.
[0045] 5, the image interpretation report 52 includes a finding statement 53. The finding statement 53 relates to a lesion contained in the medical image 51 and states, "A 12 mm partially solid nodule is observed in the right S3. A near-circular ground-glass nodule is also observed in the left S8." Note that while the medical image 51 and the image interpretation report 52 are associated with each other, the lesions contained in the medical image 51 are not associated with the individual findings contained in the image interpretation report 52.
[0046] Next, we will explain the first derivation unit 22, the second derivation unit 23, the candidate identification unit 24, the attribute estimation unit 25, and the learning unit 26. Fig. 6 is a diagram schematically showing the processes performed by the first derivation unit 22, the second derivation unit 23, the candidate identification unit 24, the attribute estimation unit 25, and the learning unit 26 in the first embodiment.
[0047] The first derivation unit 22 derives first feature amounts for multiple objects included in the medical image using a first neural network 61 (NN) to construct a first derivation model for deriving feature amounts for objects included in the medical image. In this embodiment, the first neural network 61 is a convolutional neural network (CNN), but is not limited to this. As shown in FIG. 6 , the first derivation unit 22 inputs an image 55, such as a medical image including an object such as a lesion, to the first neural network 61. Image data of the image 55 is an example of first data. The first neural network 61 extracts lesions 55A and 55B included in the image 55 as objects and derives feature vectors of the lesions 55A and 55B as first feature amounts V1-1 and V1-2, respectively. Note that in the following description, the first feature amounts V1-1 and V1-2 may be represented by the reference symbol V1.
[0048] The first neural network 61 may include two neural networks: one for extracting an object included in a medical image, and the other for deriving a feature amount for the extracted object.
[0049] The second derivation unit 23 derives second feature quantities for sentences including descriptions related to objects using a second neural network (NN) 62 to construct a second derivation model for deriving feature quantities for sentences including descriptions related to objects. FIG. 7 is a diagram schematically illustrating the second neural network 62. As shown in FIG. 7, the second neural network 62 has an embedding layer 62A, a recurrent neural network layer (hereinafter referred to as an RNN layer) 62B, and a fully connected layer 62C. The second derivation unit 23 performs morphological analysis on the input sentence to divide the sentence into words, and inputs the words into the embedding layer 62A. The embedding layer 62A outputs feature vectors of the words included in the input sentence. For example, when sentence 56, which reads "A 12mm partially solid nodule is found in the right S3," is input to second neural network 62, sentence 56 is divided into the words "A 12mm partially solid nodule is found in the right S3." Each word is then input to an element of embedding layer 62A.
[0050] The RNN layer 62B outputs a feature vector 67 that takes into account the context of the feature vector 66 of each word. The fully connected layer 62C integrates the feature vectors 67 output by the RNN layer 62B and outputs the feature vector of the sentence 56 input to the second neural network 62 as a second feature V2. Note that when the feature vectors 67 output by the RNN 62B are input to the fully connected layer 62C, the weighting of the feature vectors 67 for important words may be increased.
[0051] In the first embodiment, it is assumed that the second feature V2-1 is acquired from sentence 56, "A 12 mm partially solid nodule is observed in the right S3," and the second feature V2-2 is acquired from sentence 57, "A nearly circular ground-glass nodule is also observed in the left S8." In the following description, the reference symbols for the second features V2-1 and V2-2 may be represented by V2. The text data representing sentences 56 and 57 is an example of second data.
[0052] The second derivation unit 23 may also be configured to structure the input sentence and input the named entities obtained by structuring into the second neural network 62 to derive the second feature V2. In this embodiment, structuring means extracting named entities such as the position, findings, and size of objects contained in the sentence, and further assigning to the named entities a factual determination result indicating whether the named entity is positive, negative, or suspicious. For example, by structuring the sentence 56 "A 12mm partial solid nodule was observed in the right S3," the named entities "right S3," "12mm," and "partial solid nodule (finding +)" are obtained. Here, "finding (+)" indicates that the factuality is positive.
[0053] The candidate identification unit 24 identifies, from among a plurality of objects included in the image, candidate objects that will pair with the object described in the sentence. Hereinafter, an object included in the image will be referred to as a first object, and an object described in the sentence will be referred to as a second object. Accordingly, the candidate identification unit 24 identifies, from among a plurality of first objects, candidate first objects that will pair with the second object. Note that lesions 55A and 55B included in image 55 are an example of a first object. Furthermore, at least a portion of the description in sentence 56, "A 12 mm partially solid nodule was observed in the right S3," is an example of a second object, and at least a portion of the description in sentence 57, "A nearly circular ground-glass nodule was also observed in the left S8," is an example of a second object.
[0054] In this embodiment, the candidate identification unit 24 identifies a first object candidate based on the distance between the first feature amount and the second feature amount in a feature space to which the first feature amount and the second feature amount belong. To this end, the candidate identification unit 24 plots the first feature amount V1 and the second feature amount V2 in a feature space defined by the first feature amount V1 and the second feature amount V2. Here, since the first feature amount V1 and the second feature amount V2 are n-dimensional vectors, the feature space is also n-dimensional. Note that, for the sake of explanation, in FIG. 6, the first feature amount V1 and the second feature amount V2 are two-dimensional, and the first feature amount V1 (V1-1, V1-2) is plotted as a black circle and the second feature amount V2 (V2-1, V2-2) is plotted as a white circle in the two-dimensional feature space.
[0055] The candidate identification unit 24 derives the distance between each of the second feature amounts V2-1 and V2-2 and the first feature amounts V1-1 and V1-2 in the feature space. Then, the candidate identification unit 24 identifies the first object from which the first feature amount V1 is derived, for which the distance is equal to or less than a predetermined threshold value Th1, as a first object candidate. Specifically, for the second feature amount V2-1 derived from sentence 56, the candidate identification unit 24 identifies the lesion 55A from which the first feature amount V1-1 is derived as a first object candidate. Furthermore, for the second feature amount V2-2 derived from sentence 57, the candidate identification unit 24 identifies the lesion 55B from which the first feature amount V1-2 is derived as a first object candidate.
[0056] The candidate identification unit 24 may identify a first object candidate based on the degree of association between the first feature amount V1 and the second feature amount V2. The degree of association may be the cosine similarity (inner product) between the first feature amount V1 and the second feature amount V2. In this case, the degree of association takes a value between -1 and +1. The candidate identification unit 24 identifies, as the first object candidate, a first object from which a first feature amount V1 has been derived, the degree of association of which is equal to or greater than a predetermined threshold value Th2.
[0057] The attribute estimation unit 25 estimates the attributes of a paired second object to be paired with the first object candidate, based on the first object candidate identified by the candidate identification unit 24. In this embodiment, the attribute estimation unit 25 estimates the attributes of a paired second object to be paired with the first object candidate, from the first feature amount V1 of the first object candidate. Specifically, the attribute estimation unit 25 estimates the attributes of a paired second object to be paired with lesion 55A, which is described in sentence 56, from the first feature amount V1-1 for lesion 55A included in image 55, and estimates the attributes of a paired second object to be paired with lesion 55B, which is described in sentence 57, from the first feature amount V1-2 for lesion 55B.
[0058] For this purpose, the attribute estimation unit 25 has a derivation model 25A that has been subjected to machine learning to derive attributes from the first feature value V1 of the first object candidate. Attributes output by the derivation model 25A include, for example, the position, size, and properties of the lesion. It should be noted that, as the properties of the object, a determination result of whether the result is positive or negative for a plurality of types of property items is derived. The property items include, for example, for a lesion contained in the lung, the shape of the margin (lobulated, spicules), the absorption value (solid, ground-glass), the clarity of the boundary, the presence or absence of calcification, the presence or absence of pleural indentation, etc.
[0059] The derived model 25A is, for example, a convolutional neural network, and is constructed by machine learning so that when a first feature amount V1 is input, the derived model 25A outputs an attribute represented by the first feature amount V1. The attribute represented by the first feature amount V1 output by the attribute estimation unit 25 becomes the estimated attribute of the paired second object that pairs with the estimated first object candidate.
[0060] That is, since the first feature V1-1 represents the attribute of the lesion 55A included in the image 55, the attribute output by the derived model 25A of the attribute estimation unit 25 using the first feature V1-1 is an estimated attribute of the second object described in the sentence 56 that is paired with the lesion 55A. Furthermore, since the first feature V1-2 represents the attribute of the lesion 55B included in the image 55, the attribute output by the derived model 25A using the first feature V1-2 is an estimated attribute of the second object described in the sentence 57 that is paired with the lesion 55B. Hereinafter, the estimated attribute will be referred to as the estimated attribute.
[0061] In FIG. 6, "Right S1," "14 mm," and "Partially Solid" are shown as inferred attributes 64A based on the first feature V1-1. "Right S1," "14 mm," and "Partially Solid" are attributes of position, size, and property, respectively. Also, "Left S9," "8 mm," and "Frosted Glass" are shown as inferred attributes 64B based on the first feature V1-2. "Left S9," "8 mm," and "Frosted Glass" are attributes of position, size, and property, respectively.
[0062] The attribute estimation unit 25 may estimate the attribute of a second object described in a sentence that is paired with the lesions 55A and 55B from data (pixel values) of the regions of the lesions 55A and 55B in the image 55, instead of the first feature V1. In this case, the derivation model 25A of the attribute estimation unit 25 is constructed so as to output the attribute from the data of the regions of the lesions in the image.
[0063] Incidentally, when an image contains multiple lesions, there may be only one sentence corresponding to the multiple lesions. For example, when an image contains two lesions, the finding sentence may be "Two 12 mm partially solid nodules are observed in the right S3." In this case, a first feature V1 is derived for each of the two lesions, but there are two candidates corresponding to one sentence in the feature space. In this case, the two first feature V1 may be added, and the attribute may be estimated using the added first feature V1. In this case, the two first feature V1 may be weighted and added. The weight coefficient may be determined so that it increases as the distance between the second feature V2 and each of the two first feature V1 in the feature space decreases.
[0064] Furthermore, when an image contains multiple lesions and there is only one sentence corresponding to the multiple lesions, the attributes of the paired second object paired with lesions 55A and 55B may be estimated from the data (pixel values) of the regions of lesions 55A and 55B in image 55. In this case, the data of the regions of lesions 55A and 55B may be added, and the attribute may be estimated using the added data. In this case, the two pieces of data may be weighted and added. The weighting coefficient may be determined so that it increases as the distance between the second feature value and each of the two first feature values V1 in the feature space decreases.
[0065] The learning unit 26 learns at least one of the first neural network 61 and the second neural network 62 so that the difference between the attribute of the estimated paired second object and the attribute of the paired second object that pairs with the first object candidate derived from the sentence becomes small, thereby constructing a first derived model that derives features for objects included in an image and a second derived model that derives features for objects included in a sentence.
[0066] That is, the learning unit 26 learns at least one of the first neural network 61 and the second neural network 62 so as to reduce the difference between the estimated attribute 64A estimated by the attribute estimation unit 25 and the attribute 65A of the paired second object derived from the sentence 56, and the difference between the estimated attribute 64B estimated by the attribute estimation unit 25 and the attribute 65B of the paired second object derived from the sentence 57. Note that in this embodiment, both the first neural network 61 and the second neural network 62 are trained, but this is not limitative. Only one of the first neural network 61 and the second neural network 62 may be trained.
[0067] Specifically, the learning unit 26 derives the difference between an estimated attribute 64A of sentence 56 estimated by the attribute estimation unit 25 and an attribute 65A of sentence 56 as a loss L1. In addition, the learning unit 26 derives the difference between an estimated attribute 64B of sentence 57 estimated by the attribute estimation unit 25 and an attribute 65B of sentence 57 as a loss L2. In FIG. 6, "right S3," "12mm," and "partially solid" are shown as attributes 65A, and "left S8" and "frosted glass" are shown as attributes 65B.
[0068] When the second derivation unit 23 structures the sentences 56 and 57, the structured named entities may be used as attributes of the paired second objects described in the sentences 56 and 57. Furthermore, the attributes 65A and 65B may be derived using a derivation model (not shown) that has been machine-learned to derive attributes of the paired second objects described in the sentences from the sentences.
[0069] Alternatively, a dictionary may be prepared in which keywords for various positions, various sizes, and various characteristics are registered, and the attributes of the paired second object described in a sentence may be derived by referring to the keywords registered in the dictionary. For example, position keywords may include "right S1, right S2...", size keywords may include "number + mm, number + cm, large", and characteristics keywords may include "partially solid, ground-glass type, nodule, cyst...". Here, since sentence 56 is "A 12mm partially solid nodule was observed in right S3," attributes 65A of "right S3", "12mm", and "partially solid" can be derived by referring to the dictionary.
[0070] The learning unit 26 learns the first neural network 61 and the second neural network 62 based on the derived losses L1 and L2. That is, the learning unit 26 learns the weights of the connections between the layers constituting the first neural network 61 and the second neural network 62 and the kernel coefficients used for convolution so as to reduce the losses L1 and L2.
[0071] Specifically, the learning unit 26 repeatedly performs learning until the losses L1 and L2 become equal to or less than a predetermined threshold. Note that the learning unit 26 may repeat learning a predetermined number of times. As a result, a first derivation model and a second derivation model are constructed that derive the first feature amount V1 and the second feature amount V2 such that the distance in the feature space is small when an object included in an image corresponds to an object described in a sentence, and the distance in the feature space is large when they do not correspond.
[0072] The first and second derived models thus constructed are transmitted to the interpretation WS 3 and used in the information processing device according to the first embodiment.
[0073] Next, the functional configuration of the information processing device according to the first embodiment will be described. Fig. 8 is a diagram showing the functional configuration of the information processing device according to the first embodiment. As shown in Fig. 8, the information processing device 30 includes an information acquisition unit 31, a first analysis unit 32, a second analysis unit 33, an identification unit 34, an attribute derivation unit 35, and a display control unit 36. When the CPU 41 executes an information processing program 42, the CPU 41 functions as the information acquisition unit 31, the first analysis unit 32, the second analysis unit 33, the identification unit 34, the attribute derivation unit 35, and the display control unit 36.
[0074] The information acquisition unit 31 acquires a target medical image G0 to be interpreted from the image server 5 in response to an instruction from the input device 45 by an operator, that is, a radiologist.
[0075] The first analysis unit 32 derives a first feature value V1 for an object such as a lesion included in the target medical image G0 by analyzing the target medical image G0 using the first derived model 32A constructed by the above-described learning device 7. In this embodiment, it is assumed that the target medical image G0 includes two objects, and first feature values V1-1 and V1-2 are derived for each of the two objects.
[0076] In the information processing device 30 according to the first embodiment, an interpretation doctor interprets the target medical image G0 in the interpretation WS3 and inputs a finding statement including the interpretation result using the input device 45, thereby generating an interpretation report. The second analysis unit 33 analyzes the input finding statement using the second derivation model 33A constructed by the learning device 7 described above, thereby deriving a second feature value V2 for the input finding statement.
[0077] The specific part 34 derives the distance in the feature space between the first feature quantity V1 derived by the first analysis part 32 and the second feature quantity V2 derived by the second analysis part 33. Then, based on the derived distance, it specifies the first feature quantity V1 corresponding to the second feature quantity V2. FIG. 9 is a diagram for explaining the specification of the first feature quantity. In FIG. 9, for the sake of explanation, the feature space is shown two-dimensionally. As shown in FIG. 9, in the feature space, when comparing the distance d3 between the first feature quantity V1-1 and the second feature quantity V2 and the distance d4 between the first feature quantity V1-2 and the second feature quantity V2, d3 < d4. Therefore, the specific part 34 specifies the first feature quantity corresponding to the second feature quantity V2 as the first feature quantity V1-1.
[0078] Based on the specified first feature quantity V1-1, the attribute derivation part 35 derives the attribute of the object from which the first feature quantity V1-1 was derived. For this purpose, the attribute derivation part 35 has a derivation model 35A similar to the derivation model 25A of the attribute estimation part 25 of the learning device 7. The attribute derivation part 35 derives the attribute of the object from which the first feature quantity V1-1 was derived by means of the derivation model 35A. For example, the derivation model 35A derives "location: lower lobe of the right lung S6, size: 23 mm, property: partial filling" as the attribute of the lesion 73 included in the target medical image G0.
[0079] The display control part 36 displays the object from which the specified first feature quantity was derived, distinguishing it from other regions in the target medical image G0. FIG. 10 is a diagram showing the creation screen of the reading report displayed on the reading WS3. As shown in FIG. 10, the creation screen 70 of the reading report has an image display area 71 and a text display area 72. The target medical image G0 is displayed in the image display area 71. In FIG. 10, the target medical image G0 is one tomographic image constituting a three-dimensional image of the chest. Findings texts input by the radiologist are displayed in the text display area 72. In FIG. 10, the findings text "A 23-mm partial filling type nodule is recognized in S8 of the lower lobe of the right lung. The margin is unclear," is displayed. The second findings text is in the middle of being described.
[0080] The target medical image G0 shown in FIG. 10 includes a lesion 73 in the right lung and a lesion 74 in the left lung. Comparing the first feature value V1-1 derived for the lesion 73 in the right lung with the first feature value V1-2 derived for the lesion 74 in the left lung, the distance between the first feature value V1-1 and the second feature value V2 derived for the finding, "A 23 mm partially solid nodule was observed in the lower lobe S8 of the right lung," is smaller. Therefore, the display control unit 36 displays the lesion 73 in the right lung distinct from other areas in the target medical image G0. In FIG. 10, the lesion 73 in the right lung is surrounded by a rectangular mark 75 to distinguish it from other areas, but this is not limiting. Marks of any shape, such as arrows, can be used.
[0081] An annotation 76 containing attributes is displayed to the right of the lesion 73 in the target medical image G0. The annotation 76 includes the following description: "Location: right lower lobe S6 of the lung, size: 23 mm, characteristics: partially solid." Here, the location information contained in the finding displayed in the text display area 72 is "right lower lobe S8," while the location information contained in the attribute derived from the first feature V1-1 is "right lower lobe S6 of the lung." Therefore, the display control unit 36 underlines the portion of the finding displayed in the text display area 72 that differs from the attribute contained in the annotation 76, that is, "S8." Note that the method of highlighting the differences is not limited to underlining, and may involve changing the color or thickness of the characters, for example.
[0082] Next, the processing performed in the first embodiment will be described. Fig. 11 is a flowchart of the learning processing according to the first embodiment. It is assumed that the images and radiology reports used for learning are acquired by the information acquisition unit 21 from the image server 5 and the report server 6, respectively, and stored in the storage 13. It is also assumed that the condition for terminating learning is that the loss is equal to or less than a predetermined threshold value.
[0083] First, the first derivation unit 22 derives a first feature amount V1 for an object included in an image using the first neural network 61 (step ST1). The second derivation unit 23 derives a second feature amount V2 for a sentence including a description related to the object using the second neural network 62 (step ST2). Note that the processing of step ST2 may be performed first, or the processing of steps ST1 and ST2 may be performed in parallel.
[0084] Next, the candidate identification unit 24 identifies an object candidate that will pair with the second object described in the sentence from among a plurality of first objects included in the image (candidate identification: step ST3).Then, the attribute estimation unit 25 estimates the attribute of a paired second object that will pair with the first object candidate based on the first object candidate identified by the candidate identification unit 24 (step ST4).
[0085] Next, the learning unit 26 derives a loss, which is the difference between the attribute of the estimated paired second object and the attribute of the paired second object derived from the sentence (step ST5). Then, it is determined whether the loss is equal to or less than a predetermined threshold (step ST6). If the result of step ST6 is negative, at least one of the first neural network 61 and the second neural network 62 is trained so as to reduce the loss (step ST7), and the process returns to step ST1, and the processes of steps ST1 to ST6 are repeated. If the result of step ST6 is positive, the process ends.
[0086] Next, information processing according to the first embodiment will be described. Fig. 12 is a flowchart of information processing according to the first embodiment. It is assumed that a target medical image G0 to be processed is acquired by the information acquisition unit 31 and stored in the storage 43. First, the first analysis unit 32 analyzes the target medical image G0 using the first derivation model 32A to derive a first feature value V1 for an object such as a lesion contained in the target medical image G0 (step ST11).
[0087] Next, the information acquisition unit 31 acquires the finding statement input by the radiologist using the input device 45 (step ST12), and the second analysis unit 33 analyzes the input finding statement using the second derivation model 33A to derive a second feature value V2 for the input finding statement (step ST13).
[0088] Next, the identification unit 34 derives the distance in feature space between the first feature amount V1 derived by the first analysis unit 32 and the second feature amount V2 derived by the second analysis unit 33, and identifies the first feature amount V1 corresponding to the second feature amount V2 based on the derived distance (step ST14). Furthermore, the attribute derivation unit 35 derives the attribute of the object from which the first feature amount V1 was derived (step ST15). Then, the display control unit 36 displays the identified object from which the first feature amount V1 was derived in the target medical image G0, distinguishing it from other areas, and displays the annotation 76 describing the attribute (step ST16), thereby ending the process.
[0089] In this way, in the learning device according to the first embodiment, a first object candidate that pairs with a second object described in a sentence is identified from among a plurality of first objects included in an image, the attributes of a paired second object that pairs with the first object candidate are estimated based on the first object candidate, and at least one of the first neural network 61 and the second neural network 62 is trained so that the difference between the estimated attributes of the paired second object and the attributes of the paired second object derived from the sentence is small, thereby constructing a first derived model 32A and a second derived model 33A.
[0090] Therefore, by applying the first derived model 32A and the second derived model 33A constructed by learning to the information processing device 30 according to the first embodiment, even if an object included in an image and an object described in a sentence are not associated one-to-one, the first feature V1 and the second feature V2 are derived so that an image including an object and a sentence including a description of the object included in the image are associated. Therefore, by using the derived first feature V1 and second feature V2, it is possible to associate an image with a sentence with high accuracy.
[0091] In addition, since it is possible to accurately match images with text, when creating an interpretation report for a medical image, it is possible to accurately identify objects described in the input findings text in the medical image.
[0092] Although the information processing device according to the first embodiment is described as including the attribute derivation unit 35, the present invention is not limited to this. The information processing device may not include the attribute derivation unit 35. In this case, annotations containing attributes will not be displayed.
[0093] In the first embodiment, the first neural network 61 and the second neural network 62 are trained so as to reduce the difference between the attribute of the estimated paired second object and the attribute of the paired second object derived from the sentence, but in addition to this, the first and second neural networks 61, 62 may be trained using a combination of a medical image and a radiology report corresponding to the medical image as training data. This will be described below as a second embodiment of the learning device.
[0094] In the learning device according to the second embodiment, a combination of a medical image and an interpretation report corresponding to the medical image is used as training data to train the first and second neural networks 61 and 62. FIG. 13 is a diagram showing the training data used in the second embodiment. As shown in FIG. 13, training data 80 consists of a training image 81 and an interpretation report 82 corresponding to the training image 81. The training image 81 is a tomographic image of the lung, and includes lesions 81A and 81B as first objects in the right lower lobe S6 and the left lower lobe S8, respectively.
[0095] The radiology report 82 includes three findings 82A to 82C. The finding 82A is "A 12 mm solid nodule is present in the right lower lobe S6." The finding 82B is "A 5 mm GGN (ground-glass nodule) is present in the right lung S7." The finding 82C is "A micronodule is present in the left lung S9." Here, since the training image 81 includes the lesion 81A in the right lower lobe S6, among the three findings 82A to 82C, the finding 82A corresponds to the lesion 81A. Note that the training image 81 is one of the multiple tomographic images that make up the three-dimensional image. The findings 82B and 82C were generated as a result of interpreting tomographic images other than the training image 81. Therefore, the training image 81 does not correspond to the findings 82B and 82C. Furthermore, the training image 81 includes a lesion 81B in the left lower lobe S8, but this does not correspond to any of the three findings 82A to 82C.
[0096] In the second embodiment, the first derivation unit 22 derives the feature quantities of lesions 81A and 81B included in the teacher image 81 as first feature quantities V1-3 and V1-4 using the first neural network 61. The second derivation unit 23 derives the feature quantities of each of the findings 82A to 82C as second feature quantities V2-3, V2-4, and V2-5 using the second neural network 62.
[0097] In the second embodiment, the candidate identifying unit 24 identifies a first object candidate that pairs with a second object described in the corresponding finding statement 82A from among a plurality of first objects (lesions 81A and 81B) included in the teacher image 81. That is, the candidate identifying unit 24 identifies the lesion 81A as a first object candidate.
[0098] Furthermore, based on the lesion 81A that is the first object candidate identified by the candidate identification unit 24, the attribute estimation unit 25 estimates the attribute of the paired second object that is described in the finding statement 82A and pairs with the lesion 81A.
[0099] The learning unit 26 learns the first neural network 61 and the second neural network 62 so that the difference between the estimated attribute of the paired second object and the attribute of the paired second object derived from the observation sentence 82A becomes small.
[0100] Furthermore, in the second embodiment, the learning unit 26 plots the first feature amounts V1-3 and V1-4 and the second feature amounts V2-3, V2-4, and V2-5 in a feature space. Fig. 14 is a diagram for explaining the plotting of the feature amounts in the second embodiment. Note that, for the sake of explanation, Fig. 14 also shows the feature space in two dimensions.
[0101] In the second embodiment, the learning unit 26 trains the first and second neural networks 61, 62 based on the correspondence between the lesions 81A, 81B and the findings included in the teacher image 81 in the teacher data 80. That is, the learning unit 26 further trains the first and second neural networks 61, 62 so that, in the feature space, the first feature V1-3 and the second feature V2-4 become closer to each other and the first feature V1-3 and the second feature V2-5, V2-6 become farther apart. In this case, the learning unit 26 may train the first and second neural networks 61, 62 so that the degree of separation between the first feature V1-3 and the second feature V2-5 becomes smaller than the degree of separation between the first feature V1-3 and the second feature V2-6. Furthermore, the learning unit 26 trains the first and second neural networks 61, 62 so that the first feature V1-4 is separated from the second feature V2-4, V2-5, and V2-6 in the feature space. Either the learning based on the difference in attributes or the learning based on the correspondence between the objects and the observation statements included in the teacher images in the teacher data may be performed first.
[0102] In this way, the learning device according to the second embodiment further performs learning based on the correspondence between the objects included in the training images in the training data and the findings, which allows the first and second neural networks 61 and 62 to be trained with even greater accuracy, thereby constructing the first derived model 32A and the second derived model 33A that can derive features with even greater accuracy.
[0103] Next, a second embodiment of an information processing device will be described. Fig. 15 is a functional configuration diagram of an information processing device according to the second embodiment. In Fig. 15, the same components as those in Fig. 8 are given the same reference numerals, and detailed description thereof will be omitted. As shown in Fig. 15, an information processing device 30A according to the second embodiment differs from the information processing device according to the first embodiment in that it does not include an attribute derivation unit 35, and includes a search unit 37 instead of the identification unit 34.
[0104] In the information processing device 30A according to the second embodiment, the information acquisition unit 31 acquires a large number of medical images stored in the image server 5. Then, the first analysis unit 32 derives a first feature V1 for each of the medical images. The information acquisition unit 31 transmits the first feature V1 to the image server 5. In the image server 5, the medical images are associated with the first feature V1 and stored in the image DB 5A. The medical images associated with the first feature V1 and registered in the image DB 5A will be referred to as reference images in the following description.
[0105] In the information processing device 30A according to the second embodiment, an interpretation doctor interprets the target medical image G0 in the interpretation WS3 and inputs a finding statement including the interpretation result using the input device 45, thereby generating an interpretation report. The second analysis unit 33 analyzes the input finding statement using the second derivation model 33A constructed by the learning device 7 described above, thereby deriving a second feature value V2 for the input finding statement.
[0106] The search unit 37 refers to the image DB 5 and searches in the feature space for a reference image associated with a first feature V1 that is close to the second feature V2 derived by the second analysis unit 33. Fig. 16 is a diagram for explaining the search performed in the information processing device 30A according to the second embodiment. Note that, for the sake of explanation, Fig. 16 also shows the feature space in two dimensions. For the sake of explanation, five first feature values V1-11 to V1-15 are plotted in the feature space.
[0107] The search unit 37 identifies a first feature value in the feature space whose distance from the second feature value V2 is within a predetermined threshold value. In Fig. 16, a circle 85 having a radius d5 and centered on the second feature value V2 is shown. The search unit 37 identifies a first feature value included in the circle 85 in the feature space. In Fig. 16, three first feature values V1-11 to V1-13 are identified.
[0108] The search unit 37 searches the image DB 5A for reference images associated with the identified first feature amounts V1-11 to V1-13, and acquires the searched reference images from the image server 5.
[0109] The display control unit 36 displays the acquired reference image on the display 44. FIG. 17 is a diagram showing a display screen in the information processing device 30A according to the second embodiment. As shown in FIG. 17, the display screen 90 has an image display area 91, a text display area 92, and a result display area 93. The target medical image G0 is displayed in the image display area 91. In FIG. 17, the target medical image G0 is one of the tomographic images that constitutes a three-dimensional image of the chest. The text display area 92 displays the finding entered by the radiologist. In FIG. 17, the finding reads, "There is a 10 mm solid nodule in the right lung S6."
[0110] The result display area 93 displays the reference images searched by the search unit 37. In FIG.
[0111] Next, information processing according to the second embodiment will be described. Fig. 18 is a flowchart of information processing according to the second embodiment. It is assumed that the first feature amount for the reference image is derived by the first analysis unit 32 and is associated with the reference image and registered in the image DB 5A. It is also assumed that the target medical image G0 is displayed on the display 44 by the display control unit 36. In the second embodiment, the information acquisition unit 31 acquires a finding statement input by the radiologist using the input device 45 (step ST21), and the second analysis unit 33 analyzes the input finding statement using the second derivation model 33A to derive a second feature amount V2 for the object described in the input finding statement (step ST22).
[0112] Next, the search unit 37 refers to the image DB 5A to search for a reference image associated with the first feature V1 that is close to the second feature V2 (step ST23).Then, the display control unit 36 displays the searched reference image on the display 44 (step ST25), and the process ends.
[0113] The reference images R1 to R3 retrieved in the second embodiment are medical images with similar features to the findings entered by the radiologist. Since the findings relate to the target medical image G0, the reference images R1 to R3 are similar in case to the target medical image G0. Therefore, according to the second embodiment, the target medical image G0 can be interpreted by referring to the reference images with similar cases. Furthermore, the radiology report for the reference images can be obtained from the report server 6 and used to create the radiology report for the target medical image G0.
[0114] In each of the above embodiments, at least one of the first neural network 61 and the second neural network 62 is trained so as to reduce the difference between the attribute of the estimated paired second object and the attribute of the paired second object derived from the sentence, but the derived model 25A of the attribute estimation unit 25 may be further trained so as to reduce the difference.
[0115] In addition, in the above-described embodiments, a derivation model is constructed to derive feature quantities of medical images and commentary on medical images, but the present disclosure is not limited to this. For example, the technology of the present disclosure can also be applied to constructing a derivation model to derive feature quantities of photographic images and commentary on photographic images.
[0116] Furthermore, in the above embodiments, images and text are used as the first data and second data in the present disclosure, but the present disclosure is not limited to these. Data such as moving images and audio may also be used as the first data and second data in the present disclosure.
[0117] Furthermore, in each of the above embodiments, the following various processors can be used as the hardware structure of processing units that perform various processes, such as the information acquisition unit 21, first derivation unit 22, second derivation unit 23, candidate identification unit 24, attribute estimation unit 25, and learning unit 26 in the learning device 7, and the information acquisition unit 31, first analysis unit 32, second analysis unit 33, identification unit 34, attribute derivation unit 35, display control unit 36, and search unit 37 in the information processing device 30, 30A. As described above, the various processors include a CPU, which is a general-purpose processor that executes software (programs) and functions as various processing units, as well as dedicated electrical circuits that are processors having a circuit configuration specifically designed to perform specific processes, such as a programmable logic device (PLD) whose circuit configuration can be changed after manufacture, such as an FPGA (Field Programmable Gate Array), and an application-specific integrated circuit (ASIC).
[0118] A single processing unit may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs or a combination of a CPU and an FPGA). Also, multiple processing units may be configured with a single processor. Examples of multiple processing units configured with a single processor include, first, a configuration in which one processor is configured with a combination of one or more CPUs and software, as typified by client and server computers, and this processor functions as multiple processing units. Second, a configuration in which a processor is used to realize the functions of an entire system including multiple processing units on a single IC (Integrated Circuit) chip, as typified by a System on Chip (SoC). In this way, the various processing units are configured with one or more of the above-mentioned various processors as a hardware structure.
[0119] Furthermore, more specifically, the hardware structure of these various processors can be an electric circuit that combines circuit elements such as semiconductor elements. [Explanation of symbols]
[0120] 1 Medical Information System 2. Imaging equipment 3 Image Reading Workshop 4. Medical Workshop 5 Image Server 5A Image DB 6 Report Server 6A Report DB 7 Learning Device 10 Network 11,41 CPU 12 Study Programs 13,43 Storage 14,44 display 15,45 Input Devices 16,46 memory 17,47 Network I / F 18,48 bus 21 Information Acquisition Department 22 First derivation part 23 Second derived part 24 Candidate identification part 25 Attribute estimation part 25A Derived Model 26 Learning Department 30, 30A Information processing equipment 31 Information Acquisition Department 32 1st Analysis Department 32A First Derived Model 33 2nd Analysis Department 33A Second Derived Model 34 Specific part 35 Attribute derivation part 35A Derived Model 36 Display control unit 37 Search Section 42 Information Processing Program 51 Medical Imaging 51A Tomographic Image 52 Image interpretation report 53 Observations 55 images 55A, 55B Objects 56,57 sentences 61 First Neural Network 62 Second Neural Network 62A Buried Layer 62B RNN layer 62C fully connected layer 64A,64B Estimated attributes 65A,65B attributes 66,67 Feature vector 70 Image interpretation report creation screen 71,91 Image display area 72,92 Text display area 73,74 Lesions 75 marks 76 Annotations 80 training data 81 Teacher Images 81A, 81B Lesions 82,82A,82B,82C Observations 85 yen 90 display screen 93 Reference image display area d3,d4 distance d5 radius G0 Target Medical Images L1,L2 loss R1, R2, R3 Reference Image V1, V1-1, V1-2, V1-3, V1-11 to V1-15 First feature V2,V2-1,V2-2,V2-3,V2-4,V2-5,V2-6 2nd feature amount
Claims
1. at least one processor; The processor: deriving a first feature amount for each of a plurality of first objects included in the first data by a first neural network; deriving second features for second data including one or more second objects using a second neural network; Identifying a first object candidate to be paired with the second object from among the plurality of first objects; estimating an attribute of a paired second object that is paired with the first object candidate from the first feature amount of the first object candidate; A learning device that constructs a first derived model that derives features for an object included in the first data and a second derived model that derives features for second data that includes the object by training the first neural network and the second neural network so that the difference between the estimated attributes of the paired second object and the attributes of the paired second object derived from the second data is small.
2. The learning device according to claim 1 , wherein the processor identifies the first object candidate based on a distance between the first feature amount and the second feature amount in a feature space to which the first feature amount and the second feature amount belong.
3. The learning device according to claim 1 , wherein the processor identifies the first object candidate based on a degree of association between the first feature amount and the second feature amount.
4. The learning device according to any one of claims 1 to 3, wherein when multiple first object candidates are identified, the processor estimates the attributes of the paired second object from the sum or weighted sum of the first features for the multiple first object candidates.
5. 5. The learning device according to claim 1, wherein, when the first object and the second object correspond to each other, the processor further trains the first neural network and the second neural network so that a distance between the derived first feature and the derived second feature becomes smaller in a feature space to which the first feature and the second feature belong, and when the first object and the second object do not correspond to each other, the processor further trains the first neural network and the second neural network so that a distance between the derived first feature and the derived second feature becomes larger in the feature space.
6. the first data is image data representing an image including a first object; The learning device according to claim 1 , wherein the second data is text data representing a sentence including a description related to the second object.
7. the image is a medical image, and a first object included in the image is a lesion included in the medical image; The learning device according to claim 6 , wherein the sentence is a finding sentence about the medical image, and the second object is a finding about a lesion in the sentence.
8. A system comprising at least one processor, The processor: deriving first feature amounts for one or more first objects included in a target image using a first derived model constructed by the learning device according to claim 6 or 7; deriving second feature quantities for one or more target sentences including a description related to a second object using a second derived model constructed by the learning device according to claim 6 or 7; identifying the first feature amount corresponding to the second feature amount based on the distance in a feature space between the derived first feature amount and the derived second feature amount; displaying the first object from which the specified first feature amount has been derived in the target image in a manner distinguishable from other areas; estimating attributes of the second object paired with the identified first object based on the identified first object from which the first feature amount was derived; The information processing device further displays the estimated attributes.
9. A system comprising at least one processor, The processor: deriving first feature amounts for one or more first objects included in a target image using a first derived model constructed by the learning device according to claim 6 or 7; deriving second feature amounts for one or more target sentences including a description related to a second object using a second derived model constructed by the learning device according to claim 6 or 7; identifying the first feature amount corresponding to the second feature amount based on the distance in a feature space between the derived first feature amount and the derived second feature amount; displaying the first object from which the specified first feature amount has been derived in the target image in a manner distinguishable from other areas; Deriving an attribute of a second object described in the target sentence; An information processing device that displays the target sentence while distinguishing descriptions in the target sentence regarding the derived attribute that differs from the estimated attribute from other descriptions.
10. the computer derives, by a first neural network, a first feature amount for each of a plurality of first objects included in the first data; deriving second features for second data including one or more second objects using a second neural network; Identifying a first object candidate to be paired with the second object from among the plurality of first objects; estimating an attribute of a paired second object that is paired with the first object candidate from the first feature amount of the first object candidate; A learning method for constructing a first derived model that derives features for an object included in first data and a second derived model that derives features for second data including the object, by training the first neural network and the second neural network so that the difference between the estimated attributes of the paired second object and the attributes of the paired second object derived from the second data is small.
11. a computer deriving first feature amounts for one or more first objects included in a target image using a first derived model constructed by the learning device according to claim 6 or 7; deriving second feature amounts for one or more target sentences including a description related to a second object using a second derived model constructed by the learning device according to claim 6 or 7; identifying the first feature amount corresponding to the second feature amount based on the distance in a feature space between the derived first feature amount and the derived second feature amount; displaying the first object from which the specified first feature amount has been derived in the target image in a manner distinguishable from other areas; estimating attributes of the second object paired with the identified first object based on the identified first object from which the first feature amount was derived; An information processing method further displaying the estimated attributes.
12. a computer deriving first feature amounts for one or more first objects included in a target image using a first derived model constructed by the learning device according to claim 6 or 7; deriving second feature amounts for one or more target sentences including a description related to a second object using a second derived model constructed by the learning device according to claim 6 or 7; identifying the first feature amount corresponding to the second feature amount based on the distance in a feature space between the derived first feature amount and the derived second feature amount; displaying the first object from which the specified first feature amount has been derived in the target image in a manner distinguishable from other areas; Deriving an attribute of a second object described in the target sentence; An information processing method for displaying the target sentence by distinguishing descriptions in the target sentence regarding the derived attribute that differs from the estimated attribute from other descriptions.
13. deriving, by a first neural network, a first feature amount for each of a plurality of first objects included in the first data; deriving second features for second data including one or more second objects using a second neural network; a step of identifying a first object candidate to be paired with the second object from among the plurality of first objects; a step of estimating an attribute of a paired second object paired with the first object candidate from the first feature amount of the first object candidate; a learning program that causes a computer to execute a procedure for constructing a first derived model that derives features for an object included in first data and a second derived model that derives features for second data that includes the object, by training the first neural network and the second neural network so that a difference between an estimated attribute of the paired second object and an attribute of the paired second object derived from the second data is small.
14. a step of deriving first feature amounts for one or more first objects included in a target image using a first derived model constructed by the learning device according to claim 6 or 7; deriving second features for one or more target sentences including a description related to a second object using a second derivation model constructed by the learning device according to claim 6 or 7; a step of identifying the first feature amount corresponding to the second feature amount based on a distance in a feature space between the derived first feature amount and the derived second feature amount; a step of displaying the first object from which the specified first feature amount has been derived in the target image in a manner that distinguishes it from other areas; a step of estimating attributes of the second object paired with the identified first object based on the identified first object from which the first feature amount has been derived; and a step of displaying the estimated attribute.
15. A method for deriving first feature quantities for one or more first objects included in a target image using a first derived model constructed by the learning device according to claim 6 or 7; deriving second features for one or more target sentences including a description related to a second object using a second derivation model constructed by the learning device according to claim 6 or 7; a step of identifying the first feature amount corresponding to the second feature amount based on a distance in a feature space between the derived first feature amount and the derived second feature amount; a step of displaying the first object from which the specified first feature amount has been derived in the target image in a manner that distinguishes it from other areas; deriving an attribute of a second object described in the target sentence; and a procedure for displaying the target sentence in such a way that descriptions relating to the derived attributes that differ from the estimated attributes in the target sentence are distinguished from other descriptions.
Citation Information
Patent Citations
Mapping learning method, information compression method, device, and program
JP2016197375A
Diagnosis support device, information processing method, diagnosis support system and program
JP2019074868A
Medical information processing device, method and program
JP2019212296A
Method for automatic learning from sensor and program
JP2021117967A
Surgical video retrieval based on preoperative images
JP2021516810A