Learning device, method and program, as well as information processing apparatus, method and program

The learning device uses neural networks to align image and sentence feature amounts, addressing the challenge of varying medical image descriptions with limited data, achieving accurate image-sentence associations.

JP2025111698AActive Publication Date: 2025-07-30FUJIFILM CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025074452
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-30
Estimated Expiration
2041-08-17

AI Technical Summary

Technical Problem

The challenge of accurately associating medical images with varying sentence expressions due to differences in how doctors describe findings, requiring large amounts of training data that are often limited, hinders effective image-sentence association models.

Method used

A learning device that uses first and second neural networks to derive feature amounts from images and structured sentence information, adjusting the networks to minimize distance in a feature space when the image and sentence correspond, and maximize it when they do not, enabling accurate association.

Benefits of technology

Enables high-accuracy association between medical images and sentences by constructing models that can handle diverse expressions, even with limited training data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025111698000001_ABST
    Figure 2025111698000001_ABST
Patent Text Reader

Abstract

To provide a learning device, a method and a program, as well as an information processing method and a program that can accurately associate an image with a sentence.SOLUTION: In a medical information system, a learning device 7 comprises: a first derivation unit 22 that derives a first feature value for an object included in an image via a first neural network; a structured information derivation unit 23 that derives structured information for a sentence by structuring a sentence including description related to the object included in the image; and a second derivation unit 24 that derives a second feature value for a sentence from the structured information via a second neural network. When the object included in the image corresponds to the object described in the sentence, learning of the first neural network and the second neural network is done such that a distance between the first feature value and the second feature value derived is small in a feature space to which the first feature value and the second feature value belong.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a learning device, method and program, and an information processing device, method and program.

Background Art

[0002] A method has been proposed for constructing a feature space to which feature quantities such as feature vectors extracted from images using a learned model obtained by machine learning such as deep learning belong. For example, Non-Patent Document 1 proposes a method of extracting feature quantities from each of an image and text, and estimating the relationship between the image and the text based on the feature quantities.

[0003] In addition, a method has also been proposed for analyzing text data to obtain word data and identifying an object in an image based on the word data (see Patent Document 1).

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Non-Patent Documents

[0005]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] Incidentally, when the content of an image is described as a sentence, even if the content is the same, the way of expression varies depending on the person who describes it. Therefore, for the findings sentence about a medical image, even if the findings are the same, the way of expression varies depending on the doctor. For example, for a medical image presenting findings that there is a solid nodule in region S6 of the right lung, its size is 10 mm, and the boundary is unclear, the findings sentence may be, depending on the doctor who describes it, "A solid nodule is recognized in the right lung S6. The size is 10 mm. The boundary is somewhat unclear.", "A solid nodule of 10 mm in size is recognized in the right lung S6. The margin is relatively unclear.", and "A solid nodule of φ10 mm in the lower right lobe S6. The boundary is somewhat unclear.", and the ways of expression are different like this. Thus, sentences such as findings sentences vary greatly because even if the content is the same, the way of expression is different. In order to construct a model that can accurately derive feature amounts from sentences with such diverse expressions, a large amount of training data is required.

[0007] However, since the number of sentences is limited, it is difficult to prepare a large amount of training data. Therefore, it is difficult to construct a learned model that can accurately associate an image with a sentence.

[0008] The present disclosure has been made in view of the above circumstances, and an object thereof is to enable accurate association between an image and a sentence.

Means for Solving the Problem

[0009] The learning device according to the present disclosure includes at least one processor, The processor derives a first feature amount about an object included in the image by a first neural network, By structuring a sentence including a description about an object included in the image, structured information about the sentence is derived, The processor derives a second feature amount about the sentence from the structured information by a second neural network, When the object included in the image corresponds to the object described in the sentence, in the feature space to which the first feature amount and the second feature amount belong, the distance between the first feature amount and the second feature amount derived is smaller than when the object included in the image does not correspond to the object described in the sentence. By learning the first neural network and the second neural network, a first derivation model for deriving a feature amount for the object included in the image and a second derivation model for deriving a feature amount for the sentence including the description about the object are constructed.

[0010] Note that in the learning device according to the present disclosure, when the object included in the image does not correspond to the object described in the sentence, the processor may learn the first neural network and the second neural network so that the distance between the first feature amount and the second feature amount derived is larger in the feature space than when the object included in the image corresponds to the object described in the sentence.

[0011] Further, in the learning device according to the present disclosure, the processor may extract one or more unique expressions related to the object from the sentence, and derive the unique expression and the determination result of the factuality as structured information by determining the factuality of the unique expression.

[0012] Further, in the learning device according to the present disclosure, the unique expression represents at least one of the position, the finding, and the size of the object. The determination result of the factuality may represent any one of positive, negative, and doubtful for the finding.

[0013] Further, in the learning device according to the present disclosure, when a plurality of unique expressions are extracted, the processor may further derive the relationship between the unique expressions as structured information.

[0014] Further, in the learning device according to the present disclosure, the relationship may represent whether or not each of the plurality of unique expressions is related.

[0015] Also, in the learning device according to the present disclosure, the processor may derive normalized structured information by normalizing the unique expression and factuality.

[0016] Also, in the learning device according to the present disclosure, the image is a medical image, and the object included in the image is a lesion included in the medical image, and the sentence may be a finding sentence in which findings about the lesion are described.

[0017] The first information processing device according to the present disclosure includes at least one processor, and the processor derives a first feature amount about one or more objects included in the target image by the first derivation model constructed by the learning device according to the present disclosure, derives structured information about the target sentence by structuring one or more target sentences including a description about the object, derives a second feature amount about the target sentence from the structured information about the target sentence by the second derivation model constructed by the learning device according to the present disclosure, identifies the first feature amount corresponding to the second feature amount based on the distance in the feature space between the derived first feature amount and second feature amount, and displays, in the target image, the object from which the identified first feature amount is derived separately from other regions.

[0018] The second information processing device according to the present disclosure includes at least one processor, and the processor receives an input of a target sentence including a description about the object, derives structured information about the target sentence by structuring the target sentence, derives a second feature amount about the input target sentence from the structured information about the target sentence by the second derivation model constructed by the learning device according to the present disclosure, By referring to a database in which first feature amounts for one or more objects included in each of a plurality of reference images, derived by a first derivation model constructed by a learning device according to the present disclosure, are associated with each of the reference images, at least one first feature amount corresponding to a second feature amount is specified based on a distance in a feature space between the first feature amounts for the plurality of reference images and the derived second feature amount, Specify the reference image associated with the specified first feature amount.

[0019] Note that in the first and second information processing apparatuses according to the present disclosure, the processor may notify a unique expression that contributed to the association with the first feature amount.

[0020] The learning method according to the present disclosure derives a first feature amount for an object included in an image by a first neural network, Derive structured information about a sentence by structuring a sentence including a description about an object included in the image, Derive a second feature amount for the sentence from the structured information by a second neural network, When the object included in the image corresponds to the object described in the sentence, in the feature space to which the first feature amount and the second feature amount belong, the distance between the first feature amount and the second feature amount derived is smaller than when the object included in the image does not correspond to the object described in the sentence. By learning the first neural network and the second neural network, a first derivation model for deriving a feature amount for an object included in an image and a second derivation model for deriving a feature amount for a sentence including a description about the object are constructed.

[0021] The first information processing method according to the present disclosure derives a first feature amount for one or more objects included in a target image by a first derivation model constructed by a learning device according to the present disclosure, Derive structured information about the target sentence by structuring one or more target sentences including a description about the object, The second derivation model constructed by the learning device according to the present disclosure derives the second feature quantity for the target sentence from the structured information about the target sentence, specify the first feature quantity corresponding to the second feature quantity based on the distance in the feature space between the derived first feature quantity and the second feature quantity, display the object that derived the specified first feature quantity in the target image separately from other regions.

[0022] The second information processing method according to the present disclosure receives an input of a target sentence including a description about an object, derive structured information about the target sentence by structuring the target sentence, The second feature quantity for the input target sentence is derived from the structured information about the target sentence by the second derivation model constructed by the learning device according to the present disclosure, By referring to a database in which the first feature quantity for one or more objects included in each of a plurality of reference images, derived by the first derivation model constructed by the learning device according to the present disclosure, is associated with each of the reference images, at least one first feature quantity corresponding to the second feature quantity is specified based on the distance in the feature space between the first feature quantity for the plurality of reference images and the derived second feature quantity, Specify the reference image associated with the specified first feature quantity.

[0023] Note that the learning method according to the present disclosure, as well as the first and second information processing methods, may be provided as programs for causing a computer to execute them.

Advantages of the Invention

[0024] According to the present disclosure, an image and a sentence can be associated with high accuracy.

Brief Description of the Drawings

[0025]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Embodiments for Carrying Out the Invention

[0026] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. First, the configuration of a medical information system to which a learning device and an information processing device according to the first embodiment of the present disclosure are applied will be described. FIG. 1 is a diagram showing a schematic configuration of a medical information system 1. The medical information system 1 shown in FIG. 1 is based on an examination order from a doctor in a medical department using a known ordering system, and performs imaging of an examination target site of a patient as a subject, storage of medical images acquired by the imaging, reading of the medical images by a radiologist and creation of a radiology report, and viewing of the radiology report by a doctor in the medical department as the requester and detailed observation of the medical images to be read.

[0027] As shown in FIG. 1, the medical information system 1 includes a plurality of imaging devices 2, a plurality of radiology WS (WorkStation) 3 which are radiology terminals, a diagnostic WS 4, an image server 5, an image DB (DataBase) 5A, a report server 6, a report DB 6A, and a learning device 7, which are connected to be communicable with each other via a wired or wireless network 10.

[0028] Each device is a computer installed with an application program for functioning as a component of the medical information system 1. The application program is recorded and distributed on a recording medium such as a DVD (Digital Versatile Disc) and a CD-ROM (Compact Disc Read Only Memory), and is installed on the computer from the recording medium. Alternatively, it is stored in a storage device of a server computer connected to the network 10 or in a network storage in a state accessible from the outside, and is downloaded and installed on the computer in response to a request.

[0029] The imaging device 2 is a device (modality) that generates a medical image representing a diagnostic target site by imaging the site to be diagnosed in a patient. Specifically, it includes a simple X-ray imaging device, a CT device, an MRI device, a PET (Positron Emission Tomography) device, etc. The medical image generated by the imaging device 2 is transmitted to the image server 5 and stored in the image DB 5A.

[0030] The reading WS 3 is a computer used, for example, by a radiologist to read medical images and create reading reports, etc., and includes an information processing device (details will be described later) according to this embodiment. In the reading WS 3, a request to view a medical image to the image server 5, various image processes on the medical image received from the image server 5, display of the medical image, and reception of input of findings sentences regarding the medical image are performed. Also, in the reading WS 3, analysis processing on the medical image, support for creating a reading report based on the analysis result, registration request and viewing request for the reading report to the report server 6, and display of the reading report received from the report server 6 are performed. These processes are performed by the reading WS 3 executing software programs for each process.

[0031] The medical treatment WS 4 is a computer used, for example, by a doctor in a medical department for detailed observation of images, viewing of reading reports, creation of electronic medical records, etc., and is composed of a processing device, a display device such as a display, and an input device such as a keyboard and a mouse. In the medical treatment WS 4, a request to view an image to the image server 5, display of the image received from the image server 5, a request to view a reading report to the report server 6, and display of the reading report received from the report server 6 are performed. These processes are performed by the medical treatment WS 4 executing software programs for each process.

[0032] The image server 5 is a general-purpose computer installed with a software program that provides the functions of a database management system (DBMS). The image server 5 also includes a storage device in which an image database 5A is configured. This storage device may be a hard disk device connected to the image server 5 and a data bus, or may be a disk device connected to a network-attached storage (NAS) and a storage area network (SAN) connected to the network 10. When the image server 5 receives a registration request for a medical image from the imaging device 2, it formats the medical image into a database format and registers it in the image database 5A.

[0033] In the image database 5A, the image data and the attached information of the medical images acquired by the imaging device 2 are registered. The attached information includes, for example, an image ID (identification) for identifying each medical image, a patient ID for identifying a patient, an examination ID for identifying an examination, a unique ID (UID: unique identification) assigned to each medical image, the examination date when the medical image was generated, the examination time, the type of imaging device used in the examination for acquiring the medical image, patient information such as the patient's name, age, and gender, the examination site (imaging site), imaging information (imaging protocol, imaging sequence, imaging method, imaging conditions, use of contrast agent, etc.), and information such as a series number or a collection number when a plurality of medical images are acquired in one examination. In the present embodiment, the first feature amount of the medical image derived as described later in the reading WS3 is associated with the medical image and registered in the image database 5A.

[0034] When the image server 5 receives a browsing request from the reading WS3 and the medical treatment WS4 via the network 10, it searches for the medical images registered in the image database 5A and transmits the searched medical images to the original request sources, the reading WS3 and the medical treatment WS4.

[0035] The report server 6 incorporates a software program that provides the functions of a database management system to a general-purpose computer. When the report server 6 receives a registration request for a reading report from the reading WS3, it formats the reading report into a database format and registers it in the report DB6A.

[0036] A large number of reading reports including the findings text created by the reading doctor using the reading WS3 are registered in the report DB6A. The reading report may include, for example, medical images to be read, image IDs for identifying the medical images, reading doctor IDs for identifying the reading doctor who performed the reading, lesion names, lesion location information, and information such as the nature of the lesion. In the present embodiment, in the report DB6A, a reading report and one or more medical images created for the reading report are registered in an associated manner.

[0037] Also, when the report server 6 receives a browsing request for a reading report from the reading WS3 and the medical treatment WS4 via the network 10, it searches for the reading report registered in the report DB6A and transmits the retrieved reading report to the requesting reading WS3 and medical treatment WS4.

[0038] The network 10 is a wired or wireless local area network that connects various devices within the hospital. When the reading WS3 is installed in another hospital or clinic, the network 10 may be configured to connect the local area networks of each hospital via the Internet or a dedicated line.

[0039] Next, the learning device 7 will be described. First, with reference to FIG. 2, the hardware configuration of the learning device 7 according to the first embodiment will be described. As shown in FIG. 2, the learning device 7 includes a CPU (Central Processing Unit) 11, a non-volatile storage 13, and a memory 16 as a temporary storage area. Further, the learning device 7 includes a display 14 such as a liquid crystal display, an input device 15 including a pointing device such as a keyboard and a mouse, and a network I / F (InterFace) 17 connected to the network 10. The CPU 11, the storage 13, the display 14, the input device 15, the memory 16, and the network I / F 17 are connected to a bus 18. Note that the CPU 11 is an example of the processor in the present disclosure.

[0040] The storage 13 is realized by an HDD (Hard Disk Drive), an SSD (Solid State Drive), a flash memory, or the like. A learning program 12 is stored in the storage 13 as a storage medium. The CPU 11 reads the learning program 12 from the storage 13, expands it in the memory 16, and executes the expanded learning program 12.

[0041] Next, the information processing device 30 according to the first embodiment included in the reading WS3 will be described. First, with reference to FIG. 3, the hardware configuration of the information processing device 30 according to the present embodiment will be described. As shown in FIG. 3, the information processing device 30 includes a CPU 41, a non-volatile storage 43, and a memory 46 as a temporary storage area. Further, the information processing device 30 includes a display 44 such as a liquid crystal display, an input device 45 including a pointing device such as a keyboard and a mouse, and a network I / F 47 connected to the network 10. The CPU 41, the storage 43, the display 44, the input device 45, the memory 46, and the network I / F 47 are connected to a bus 48. Note that the CPU 41 is an example of the processor in the present disclosure.

[0042] Similar to storage 13, storage 43 is implemented by a HDD, SSD, flash memory, or the like. An information processing program 42 is stored in storage 43 as a storage medium. The CPU 41 reads the information processing program 42 from the storage 43, expands it in the memory 46, and executes the expanded information processing program 42.

[0043] Next, the functional configuration of the learning device according to the first embodiment will be described. FIG. 4 is a diagram showing the functional configuration of the learning device according to the first embodiment. As shown in FIG. 4, the learning device 7 includes an information acquisition unit 21, a first derivation unit 22, a structured information derivation unit 23, a second derivation unit 24, and a learning unit 25. Then, when the CPU 11 executes the learning program 12, the CPU 11 functions as the information acquisition unit 21, the first derivation unit 22, the structured information derivation unit 23, the second derivation unit 24, and the learning unit 25.

[0044] The information acquisition unit 21 acquires a medical image and a radiology report related to the medical image from the image server 5 and the report server 6 via the network I / F 17, respectively. The medical image and the radiology report are used for the learning of the first and second neural networks described later. FIG. 5 is a diagram showing an example of a medical image and a radiology report. As shown in FIG. 5, the medical image 51 is a three-dimensional image composed of a plurality of tomographic images. In the present embodiment, the medical image 51 is a CT image of the chest of a human body. Also, as shown in FIG. 5, the plurality of tomographic images are assumed to include, for example, a tomographic image 55 including a lesion as an object 55A in the right lung S6.

[0045] Also, as shown in FIG. 5, the radiology report 52 includes findings texts 53 and 54. The findings texts 53 and 54 include descriptions regarding the objects, that is, lesions, included in the medical image 51. The findings text 53 shown in FIG. 5 includes the description "A solid nodule is recognized in the right lung S6. The size is 10 mm. The boundary is somewhat unclear." Also, the findings text 54 includes the description "There are also micronodules in the left lung S9."

[0046] Of the two findings 53 and 54 shown in FIG. 5, the finding 53 is generated as a result of reading the tomographic image 55 included in the medical image 51. Therefore, the tomographic image 55 corresponds to the finding 53. The finding 54 is generated as a result of reading tomographic images other than the tomographic image 55 in the medical image 51. Therefore, the tomographic image 55 and the finding 54 do not correspond to each other.

[0047] The first derivation unit 22 derives a first feature quantity of one or more objects included in the medical image by the first neural network 61 (NN) in order to construct a first derivation model that derives a feature quantity about the object included in the medical image. In the present embodiment, the first neural network 61 is a convolutional neural network (CNN (Convolutional Neural Network)), but is not limited thereto. As shown in FIG. 6, the first derivation unit 22 inputs an image such as a medical image including an object such as a lesion into the first neural network 61. The first neural network 61 extracts an object such as a lesion included in the image and derives a feature vector of the object as the first feature quantity V1.

[0048] The structured information derivation unit 23 derives structured information about the findings texts 53 and 54 by structuring the findings texts 53 and 54. FIG. 7 is a diagram showing the processing performed by the structured information derivation unit 23. Hereinafter, the structuring of the findings text 53 will be described, but the structured information may be derived in the same manner for the findings text 54. First, the structured information derivation unit 23 derives unique expressions related to the object from the findings text 53. The unique expression is an example of structured information. The unique expression represents at least one of the position, finding, and size of the object included in the findings text 53. In the present embodiment, all of the position, finding, and size of the object included in the findings text 53 are derived as unique expressions. Therefore, the structured information derivation unit 23 derives "right lung S6", "solid nodule", "10 mm", "margin", and "obscure" as unique expressions. Note that "right lung S6" is a unique expression representing the position, "solid nodule" is a unique expression representing the finding, "10 mm" is a unique expression representing the size, "margin" is a unique expression representing the position, and "obscure" is a unique expression representing the finding. Hereinafter, the unique expressions derived from the findings text 53 are shown as "right lung S6 (position)", "solid nodule (finding)", "10 mm (size)", "margin (position)", and "obscure (finding)".

[0049] Also, the structured information derivation unit 23 determines the factuality of the derived unique expressions. Specifically, the structured information derivation unit 23 determines whether the unique expression of the finding represents negative, positive, or doubtful, and derives the determination result. In the present embodiment, the unique expressions of the findings are "solid nodule" and "obscure", both of which are positive. Therefore, the structured information derivation unit 23 determines the factuality of "solid nodule" and "obscure" as positive, respectively. In FIG. 7, being positive is shown by attaching a + sign. Also, a - sign may be attached when it is negative, and a ± sign may be attached when there is doubt. The determination result of the factuality is an example of structured information.

[0050] In addition, the structured information derivation unit 23 derives the relationships between a plurality of specific expressions. The relationships are an example of structured information. Although the relationships are not used in the processes described later in the first embodiment, since the relationships are one type of structured information, the relationships will also be described here. The relationships indicate whether specific expressions are related to each other. For example, the specific expression "solid nodule (finding +)" representing a finding about a typical lesion among specific expressions is related to the specific expression "10 mm (size)" representing size, the specific expression "right lung S6 (location)" representing location, and the specific expression "obscure (finding +)" representing a finding, but is not related to the specific expression "margin (location)" representing location. Also, the specific expression "margin (location)" is related to the specific expression "obscure (finding +)".

[0051] Note that the derivation of the relationships may be performed by referring to a table in which the presence or absence of relationships between a large number of specific expressions is predefined. Alternatively, the relationships may be derived using a derivation model constructed by performing machine learning so as to output the presence or absence of relationships between specific expressions. Also, the lesions described in the specific expressions may be specified as keywords, and all specific expressions that modify the keywords may be specified as related specific expressions.

[0052] Furthermore, the structured information derivation unit 23 derives normalized structured information by normalizing specific expressions and factuality. Normalization in this embodiment means converting expressions that are synonymous but have variations into one standard expression. For example, "right lung S6" and "right lower lobe of the lung S6" are synonymous but have different expressions. Also, "10 mm" and "10.0 mm" are synonymous but have different expressions. Also, the combination of "margin" and "obscure (+)" is synonymous with the expression "clear margin (-)" where the factuality of "clear margin" is negative but has a different expression.

[0053] For example, in the present embodiment, a list associating synonymous expressions with normalized expressions for a large number of unique expressions and factuality is prepared in advance and stored in the storage 13. FIG. 8 is a diagram showing an example of a list associating synonymous expressions with normalized expressions. Then, the structured information derivation unit 23 refers to the list 59 and normalizes the unique expressions and factuality. As a result, the structured information derivation unit 23 derives the normalized structured information of "right lower lobe S6 (position)", "solid nodule (finding +)", "10.0 mm (size)", and "well-defined (finding -)" from the finding sentence 53. On the other hand, the structured information derivation unit 23 derives the normalized structured information of "left lower lobe S9 (position)", "micro (size)", and "nodule (finding +)" from the finding sentence 54.

[0054] The second derivation unit 24 derives second feature amounts for sentences including descriptions about an object from the structured information derived by the structured information derivation unit 23 by means of a second neural network (NN) 62 in order to construct a second derivation model for deriving feature amounts for sentences including descriptions about an object. FIG. 9 is a diagram schematically showing the second neural network 62. As shown in FIG. 9, the second neural network 62 includes an embedding layer 62A, an addition mechanism 62B, and a Transformer 62C. The second derivation unit 24 divides the input structured information into a unique expression, a type of unique expression, and a determination result of factuality, and inputs them to the embedding layer 62A. The embedding layer 62A outputs feature vectors 65 for the unique expression, the type of unique expression, and the determination result of factuality.

[0055] The addition mechanism 62B adds the feature vectors 65 for each individual structured information to derive feature vectors 66 for each structured information.

[0056] The Transformer is proposed, for example, in "Vaswani, Ashish, et al. 'Attention is all you need.' Advances in neural information processing systems. 2017." The Transformer 62C integrates the feature vectors 66 by repeatedly deriving the similarity between the feature vectors 66 and adding the feature vectors 66 with weights according to the derived similarity, and outputs the feature vector of the structured information input to the second neural network 62, that is, the feature vector of the finding sentence 53 input to the structured information derivation unit 23, as the second feature quantity V2.

[0057] Note that as a mechanism subsequent to the addition mechanism 62B, instead of the Transformer 62C, a network structure combining an RNN and an attention mechanism may be used. FIG. 10 is a diagram schematically showing a network structure combining an RNN and an attention mechanism. The network structure 67 shown in FIG. 10 has a recurrent neural network layer (hereinafter referred to as an RNN layer) 67A and an attention mechanism 67B.

[0058] The RNN layer 67A outputs a feature vector 68 that takes into account the context of the feature vector 66 output by the addition mechanism 62B. The attention mechanism 67B derives the inner product between the vector uw derived by pre-learning and each feature vector 68 as the weight coefficient w. The vector uw is learned so that a larger weight is assigned to the eigenrepresentation that contributes more significantly to the derivation of the output second feature quantity V2. Then, the attention mechanism 67B derives the second feature quantity V2 by weighted addition of the feature vector 68 with the derived weight coefficient w.

[0059] Here, in the first embodiment, as shown in FIG. 6, structured information 53A is derived from the finding sentence 53 of "A solid nodule is observed in the right lung S6. The size is 10 mm. The boundary is slightly unclear.", and it is assumed that the second feature amount V2-1 is obtained from the structured information 53A by the second neural network 62. Also, structured information 54A is derived from the finding sentence 54 of "There are also micronodules in the left lung S9.", and it is assumed that the second feature amount V2-2 is obtained from the structured information 54A by the second neural network 62.

[0060] When the object included in the image corresponds to the object described in the sentence, the learning unit 25 learns the first neural network 61 and the second neural network 62 so that the distance between the derived first feature amount V1 and the second feature amount V2 becomes small in the feature space to which the first feature amount V1 and the second feature amount V2 belong.

[0061] For this purpose, the learning unit 25 plots the first feature amount V1 and the second feature amount V2 in the feature space defined by the first feature amount V1 and the second feature amount V2. Then, the learning unit 25 derives the distance between the first feature amount V1 and the second feature amount V2 in the feature space. Here, since the first feature amount V1 and the second feature amount V2 are n-dimensional vectors, the feature space is also n-dimensional. In FIG. 6, for the sake of explanation, the first feature amount V1 and the second feature amount V2 are two-dimensional, and a state in which the first feature amount V1 and the second feature amount V2 (V2-1, V2-2) are plotted in the two-dimensional feature space is shown.

[0062] Here, the tomographic image 55 shown in FIG. 6 corresponds to the finding sentence 53 but does not correspond to the finding sentence 54. Therefore, the learning unit 25 learns the first neural network 61 and the second neural network 62 so that the first feature amount V1 approaches the second feature amount V2-1 and the first feature amount V1 moves away from the second feature amount V2-2 in the feature space.

[0063] For this purpose, the learning unit 25 derives the distance between the first feature amount V1 and the second feature amount V2 in the feature space. As the distance, any distance such as the Euclidean distance and the Mahalanobis distance can be used. Then, a loss used in learning is derived based on the distance. FIG. 11 is a diagram for explaining the derivation of the loss. First, regarding the corresponding first feature amount V1 and second feature amount V2-1, the learning unit 25 calculates the distance d1 in the feature space. Then, the distance d1 is compared with a predetermined threshold α0, and a loss L1 is derived based on the following formula (1).

[0064] That is, when the distance d1 between the first feature amount V1 and the second feature amount V2-1 is greater than the threshold α0, the loss L1 for learning the first and second neural networks 61 and 62 is calculated by d1 - α0 so that the distance of the second feature amount V2-1 from the first feature amount V1 becomes smaller than the threshold α0. On the other hand, when the distance d1 between the first feature amount V1 and the second feature amount V2-1 is equal to or less than the threshold α0, since it is not necessary to reduce the distance d1 between the first feature amount V1 and the second feature amount V2-1, L1 is set to 0. L1 = d1 - α0 (d1 > α0) L1 = 0 (d1 ≤ α0) (1)

[0065] On the other hand, regarding the non-corresponding first feature amount V1 and second feature amount V2-2, the learning unit 25 calculates the distance d2 in the feature space. Then, the predetermined threshold β0 is compared with the distance d2, and a loss L2 is derived based on the following formula (2).

[0066] That is, when the distance d2 between the first feature amount V1 and the second feature amount V2-2 is smaller than the threshold β0, the loss L2 for learning the first and second neural networks 61 and 62 is calculated by β0 - d2 so that the distance of the second feature amount V2-2 from the first feature amount V1 becomes larger than the threshold β0. On the other hand, when the distance d2 between the first feature amount V1 and the second feature amount V2-2 is equal to or greater than the threshold β0, since it is not necessary to increase the distance d2 between the first feature amount V1 and the second feature amount V2-2, L2 is set to 0. L2 = β0 - d2 (d2 < β0) L2 = 0 (when d2 ≥ β0) (2)

[0067] The learning unit 25 learns the first neural network 61 and the second neural network 62 based on the derived losses L1 and L2. That is, when d1 > α0 and when d2 < β0, the weights of the connections between the layers constituting each of the first neural network 61 and the second neural network 62 and the coefficients of the kernels used for convolution are learned so that the losses L1 and L2 become smaller.

[0068] Then, the learning unit 25 repeats the learning until the loss L1 becomes less than or equal to a predetermined threshold value and the loss L2 becomes less than or equal to the threshold value. Note that it is preferable that the learning unit 25 repeats the learning until the loss L1 becomes less than or equal to the threshold value for a predetermined number of consecutive times and the loss L2 becomes less than or equal to the threshold value for a predetermined number of consecutive times. Thereby, when the image and the text correspond to each other, the distance in the feature space becomes smaller compared to the case where the image and the text do not correspond to each other, and when the image and the text do not correspond to each other, the distance in the feature space becomes larger compared to the case where the image and the text correspond to each other, and a first derivation model and a second derivation model for deriving the first feature amount V1 and the second feature amount V2 are constructed. Note that the learning unit 25 may repeat the learning a predetermined number of times.

[0069] The first derivation model and the second derivation model constructed in this way are transmitted to the radiography WS3 and used in the information processing apparatus according to the first embodiment.

[0070] Next, the functional configuration of the information processing apparatus according to the first embodiment will be described. FIG. 12 is a diagram showing the functional configuration of the information processing apparatus according to the first embodiment. As shown in FIG. 12, the information processing apparatus 30 includes an information acquisition unit 31, a first analysis unit 32, a structured information derivation unit 33, a second analysis unit 34, a specification unit 35, and a display control unit 36. Then, by executing the information processing program 42, the CPU 41 functions as the information acquisition unit 31, the first analysis unit 32, the structured information derivation unit 33, the second analysis unit 34, the specification unit 35, and the display control unit 36.

[0071] The information acquisition unit 31 acquires the target medical image G0 to be read from the image server 5 according to an instruction from the input device 45 by the reading doctor who is the operator.

[0072] The first analysis unit 32 analyzes the target medical image G0 using the first derivation model 32A constructed by the learning device 7 described above, and derives the first feature amount V1 for objects such as lesions included in the target medical image G0. In the present embodiment, it is assumed that the target medical image G0 includes two objects, and the first feature amounts V1-1 and V1-2 are derived for each of the two objects.

[0073] Here, in the information processing apparatus 30 according to the first embodiment, in the reading WS3, the reading doctor reads the target medical image G0, and a reading report is generated by inputting the findings text including the reading result using the input device 45.

[0074] The structured information derivation unit 33 derives structured information from the input findings text. The derivation of the structured information is performed in the same manner as the structured information derivation unit 23 of the learning device 7.

[0075] The second analysis unit 34 analyzes the structured information derived from the input findings text using the second derivation model 34A constructed by the learning device 7 described above, and derives the second feature amount V2 for the input findings text.

[0076] The specific part 35 derives the distance in the feature space between the first feature quantity V1 derived by the first analysis part 32 and the second feature quantity V2 derived by the second analysis part 34. Then, based on the derived distance, it specifies the first feature quantity V1 corresponding to the second feature quantity V2. FIG. 13 is a diagram for explaining the specification of the first feature quantity. Note that in FIG. 13, the feature space is shown two-dimensionally for the sake of explanation. As shown in FIG. 13, in the feature space, when comparing the distance d3 between the first feature quantity V1-1 and the second feature quantity V2 and the distance d4 between the first feature quantity V1-2 and the second feature quantity V2, d3 < d4. Therefore, the specific part 35 specifies the first feature quantity corresponding to the second feature quantity V2 as the first feature quantity V1-1.

[0077] The display control part 36 displays, in the target medical image G0, the object from which the specified first feature quantity is derived, distinguishing it from other areas. FIG. 14 is a diagram showing the creation screen of the reading report displayed on the reading WS3. As shown in FIG. 14, the creation screen 70 of the reading report has an image display area 71 and a text display area 72. The target medical image G0 is displayed in the image display area 71. In FIG. 14, the target medical image G0 is one tomographic image constituting a three-dimensional image of the chest. The finding text input by the reading doctor is displayed in the text display area 72. In FIG. 14, the finding text "There is a 10-mm solid nodule in the right lung S6." is displayed. Note that the right lung S6 is synonymous with the right lower lobe S6 of the right lung.

[0078] The target medical image G0 shown in FIG. 14 includes a lesion 73 in the right lung and a lesion 74 in the left lung. When comparing the first feature quantity V1-1 derived for the lesion 73 in the right lung and the first feature quantity V1-2 derived for the lesion 74 in the left lung, the distance from the second feature quantity V2 derived for the finding text "There is a 10-mm solid nodule in the right lung S6." is smaller for the first feature quantity V1-1. Therefore, the display control part 36 displays, in the target medical image G0, the lesion 73 in the right lung, distinguishing it from other areas. In FIG. 14, the lesion 73 in the right lung is surrounded by a rectangular mark 75 to distinguish the lesion 73 from other areas, but it is not limited to this. Marks of any shape such as arrows can be used.

[0079] Next, the processing performed in the first embodiment will be described. FIG. 15 is a flowchart of the learning process according to the first embodiment. It is assumed that the images and the radiology reports used for learning are acquired from the image server 5 and the report server 6 by the information acquisition unit 21 and stored in the storage 13. Also, it is assumed that the end condition of learning is to perform learning a predetermined number of times.

[0080] First, the first derivation unit 22 derives a first feature amount V1 regarding the object included in the image by the first neural network 61 (step ST1). Also, the structured information derivation unit 23 derives structured information from the sentence including the description regarding the object (step ST2). Subsequently, the second derivation unit 24 derives a second feature amount V2 regarding the sentence including the description regarding the object by the second neural network 62 from the structured information (step ST3). Note that the processes of steps ST2 and ST3 may be performed first, or the processes of step ST1 and steps ST2 and ST3 may be performed in parallel.

[0081] Next, the learning unit 25 learns the first neural network and the second neural network so that the distance between the derived first feature amount V1 and the second feature amount V2 becomes small according to the correspondence relationship between the image and the sentence (step ST4). Further, the learning unit 25 determines whether or not learning has been performed a predetermined number of times (predetermined number of times of learning: step ST5). If step ST5 is negated, the process returns to step ST1, and the processes of steps ST1 to ST5 are repeated. If step ST5 is affirmed, the process ends.

[0082] Next, the information processing according to the first embodiment will be described. FIG. 16 is a flowchart of the information processing according to the first embodiment. It is assumed that the target medical image G0 to be processed is acquired by the information acquisition unit 31 and stored in the storage 43. First, the first analysis unit 32 analyzes the target medical image G0 using the first derivation model 32A to derive the first feature amount V1 regarding an object such as a lesion included in the target medical image G0 (step ST11).

[0083] Next, the information acquisition unit 31 acquires the findings text input by the radiologist using the input device 45 (step ST12), and the structured information derivation unit 33 derives structured information from the input findings text (step ST13). Next, the second analysis unit 34 analyzes the derived structured information using the second derivation model 34A to derive the second feature amount V2 regarding the input findings text (step ST14).

[0084] Subsequently, the specifying unit 35 derives the distance in the feature space between the first feature amount V1 derived by the first analysis unit 32 and the second feature amount V2 derived by the second analysis unit 34, and specifies the first feature amount V1 corresponding to the second feature amount V2 based on the derived distance (step ST15). Then, the display control unit 36 displays the object from which the specified first feature amount V1 is derived in the target medical image G0 by distinguishing it from other regions (step ST16), and the process ends.

[0085] As described above, in the learning device according to the first embodiment, structured information regarding a sentence is derived by structuring a sentence including a description regarding an object included in an image, and the second feature amount V2 regarding the sentence is derived from the structured information. Then, when the object included in the image and the object described in the sentence correspond to each other, the first neural network 61 and the second neural network 62 are learned so that the distance between the derived first feature amount V1 and the second feature amount V2 becomes small in the feature space to which the first feature amount V1 and the second feature amount V2 belong, thereby constructing the first derivation model 32A and the second derivation model 34A.

[0086] Therefore, even if there are variations in the expression of the sentences for training the second neural network 62, substantially the same structured information will be derived as long as the content is the same. In particular, if the structured information is normalized, the same structured information will be derived. As a result, since the second neural network 62 is trained using substantially the same unique expression, the second derivation model 34A can be constructed to derive the second feature amount without being affected by the variation in the expression. Therefore, without preparing a large number of sentences for training the second neural network 62, the first derivation model 32A and the second derivation model 34A that can accurately associate an image with a sentence can be constructed.

[0087] Further, by applying the first derivation model 32A and the second derivation model 34A constructed by learning to the information processing apparatus 30 according to the first embodiment, even if there are variations in the expression of the input sentence, the image including the corresponding object and the sentence including the description of the object are associated with each other, and the first feature amount V1 and the second feature amount V2 are derived so that the medical image including the non-corresponding object and the sentence including the description of the object are not associated with each other. Therefore, by using the derived first feature amount V1 and second feature amount V2, the association between the image and the sentence can be accurately performed.

[0088] Further, since the association between the image and the sentence can be accurately performed, when creating a reading report for a medical image, the object described in the input finding sentence can be accurately specified in the medical image.

[0089] In the learning device according to the first embodiment described above, a second derivation unit may be constructed by further using the relationships included in the structured information. Hereinafter, this will be described as a second embodiment of the learning device. Note that the configuration of the learning device according to the second embodiment is the same as that of the learning device 7 shown in FIG. 4, except that the second derivation unit 24 uses a second neural network having a Graph Convolutional Network (hereinafter referred to as GCN) instead of the second neural network 62 to derive the second feature amount. Therefore, detailed description of the device will be omitted here.

[0090] FIG. 17 is a diagram schematically showing a second neural network learned by the learning device according to the second embodiment. As shown in FIG. 17, the second neural network 80 in the second embodiment includes an embedding layer 80A and a GCN 80B. In the second embodiment, the second derivation unit 24 inputs the input structured information before normalization to the embedding layer 80A. The embedding layer 80A outputs a feature vector 81 for the structured information. The GCN 80B derives the second feature amount V2 based on the feature vector 81 and the relationships derived by the structured information derivation unit 23.

[0091] FIG. 18 is a diagram for explaining the derivation of the second feature amount by GCN. FIG. 18 shows the structured information before normalization in a graph structure based on the relationships derived by the structured information derivation unit 23. That is, the node of the specific expression “solid nodule (finding+)” representing the finding is related to the node of the specific expression “10 mm (size)” representing the size, the node of the specific expression “right lung S6 (position)” representing the position, and the node of the specific expression “obscure (finding+)” representing the finding, but is not related to the node of the specific expression “margin (position)” representing the position. The node of the specific expression “margin (position)” representing the position is shown to be related to the node of the specific expression “obscure (finding+)” representing the finding. In FIG. 18, “solid nodule (finding+)”, “10 mm (size)”, “right lung S6 (position)”, “obscure (finding+)”, and “margin (position)” are shown as “solid nodule”, “10 mm”, “right lung S6”, “obscure”, and “margin”.

[0092] In GCN80B, at each node, convolution is performed between the feature vector of its own node and the feature vectors of adjacent nodes, and the feature vector of each node is updated. Then, convolution using the updated feature vector is repeatedly performed, and the feature vector for the solid nodule representing the feature of a typical lesion in the structured information is output as the second feature amount V2.

[0093] In the second embodiment, the learning unit 25 learns the first neural network 61 and the second neural network 80 using the first feature amount V1 and the second feature amount V2 in the same manner as in the first embodiment. Thereby, in the second embodiment, the second feature amount V2 can be derived in consideration of the relationships of the specific expressions derived from the text.

[0094] Next, a second embodiment of the information processing apparatus will be described. FIG. 19 is a functional configuration diagram of the information processing apparatus according to the second embodiment. Note that, in FIG. 19, the same components as those in FIG. 11 are given the same reference numerals, and detailed descriptions thereof are omitted. As shown in FIG. 19, the information processing apparatus 30A according to the second embodiment is different from the information processing apparatus according to the first embodiment in that it includes a search unit 37 instead of the specifying unit 35.

[0095] In the information processing apparatus 30A according to the second embodiment, the information acquisition unit 31 acquires a number of medical images stored in the image server 5. Then, the first analysis unit 32 derives a first feature amount V1 for each of the medical images. The information acquisition unit 31 transmits the first feature amount V1 to the image server 5. In the image server 5, the medical images are stored in the image DB 5A in association with the first feature amount V1. The medical images registered in the image DB 5A in association with the first feature amount V1 are referred to as reference images in the following description.

[0096] Also, in the information processing apparatus 30A according to the second embodiment, in the reading WS3, a reading doctor reads the target medical image G0, and a reading report is generated by inputting a findings sentence including the reading result using the input device 45. The structured information derivation unit 33 derives structured information from the input findings sentence. The derivation of the structured information is performed in the same manner as the structured information derivation unit 23 of the learning apparatus 7. The second analysis unit 34 analyzes the derived structured information using the second derivation model 34A constructed by the above-described learning apparatus 7, thereby deriving a second feature amount V2 for the input findings sentence.

[0097] The search unit 37 refers to the image DB 5A and searches for a reference image associated with a first feature amount V1 whose distance from the second feature amount V2 derived by the second analysis unit 34 is close in the feature space. FIG. 20 is a diagram for explaining the search performed in the information processing apparatus 30A according to the second embodiment. Note that, in FIG. 20 as well, the feature space is shown two-dimensionally for the sake of explanation. Also, for the sake of explanation, five first feature amounts V1-11 to V1-15 are plotted in the feature space.

[0098] The search unit 37 identifies the first feature amount whose distance from the second feature amount V2 in the feature space is within a predetermined threshold value. In FIG. 20, a circle 85 with a radius d5 centered on the second feature amount V2 is shown. The search unit 37 identifies the first feature amount included within the circle 85 in the feature space. In FIG. 20, three first feature amounts V1-11 to V1-13 are identified.

[0099] The search unit 37 searches for the reference images associated with the identified first feature amounts V1-11 to V1-13 in the image DB 5A, and acquires the searched reference images from the image server 5.

[0100] The display control unit 36 displays the acquired reference images on the display 44. FIG. 21 is a diagram showing a creation screen of a radiology report in the information processing apparatus 30A according to the second embodiment. As shown in FIG. 21, the creation screen 90 has an image display area 91, a text display area 92, and a result display area 93. The target medical image G0 is displayed in the image display area 91. In FIG. 21, the target medical image G0 is one tomographic image constituting a three-dimensional image of the chest. The finding text input by the radiologist is displayed in the text display area 92. In FIG. 21, the finding text of "There is a 10-mm solid nodule in the right lung S6." is displayed.

[0101] The reference images searched by the search unit 37 are displayed in the result display area 93. In FIG. 21, three reference images R1 to R3 are displayed in the result display area 93.

[0102] Next, the information processing according to the second embodiment will be described. FIG. 22 is a flowchart of the information processing according to the second embodiment. It is assumed that the first feature amount for the reference image is derived by the first analysis unit 32, and a large number of them are registered in the image DB 5A in association with the reference image. Also, it is assumed that the target medical image G0 is displayed on the display 44 by the display control unit 36. In the second embodiment, the information acquisition unit 31 acquires the findings text input by the radiologist using the input device 45 (step ST21), and the structured information derivation unit 33 derives the structured information from the input findings text (step ST22). Next, the second analysis unit 34 analyzes the derived structured information using the second derivation model 34A to derive the second feature amount V2 for the input findings text (step ST23).

[0103] Subsequently, the search unit 37 refers to the image DB 5A and searches for a reference image associated with the first feature amount V1 having a short distance from the second feature amount V2 (step ST24). Then, the display control unit 36 displays the searched reference image on the display 44 (step ST25), and the process ends.

[0104] The reference images R1 to R3 searched in the second embodiment are medical images whose features are similar to the findings text input by the radiologist. Since the findings text relates to the target medical image G0, the reference images R1 to R3 are similar in case to the target medical image G0. Therefore, according to the second embodiment, the target medical image G0 can be read by referring to the reference images with similar cases. Also, the reading report for the reference image can be acquired from the report server 6 and utilized for creating the reading report for the target medical image G0.

[0105] In the first embodiment of the above information processing apparatus, when the display control unit 36 displays a finding sentence, it may notify a unique expression that contributed to the association with the first feature amount of the object included in the image. In this case, the second derivation unit 24 constructed from the second neural network 62 having the network structure combining the RNN and the attention mechanism shown in FIG. 10 described above is used, and according to the magnitude of the weighting in the attention mechanism, the unique expression that contributed to the association with the first feature amount may be specified. Further, the degree of contribution may be derived according to the magnitude of the weight coefficient.

[0106] FIG. 23 is a diagram showing another example of a creation screen of a radiographic report displayed on the radiographic WS3. In the creation screen 70A of the radiographic report shown in FIG. 23, "right lung S6", "10 mm", and "solid nodule" included in the finding sentence "There is a 10 mm solid nodule in the right lung S6" displayed in the text display area 72 are unique expressions that contributed to the association with the first feature amount of the lesion 73 included in the target medical image G0, and these are highlighted. In FIG. 23, the difference in the degree of contribution of each unique expression is shown by the difference in the interval and the number of lines of the hatched lines. In FIG. 23, the unique expressions included in the finding sentence are in the order of "right lung S6", "solid nodule", and "10 mm" in descending order of the degree of contribution. Thus, by notifying the unique expression that contributed to the association with the first feature amount, important keywords in the finding sentence can be easily recognized.

[0107] Of course, it is also possible to perform the notification of the unique expression that contributed to the association with the first feature amount on the display screen 90 shown in FIG. 21 in the second embodiment of the information processing apparatus.

[0108] Further, in the above embodiment, a derivation model for deriving the feature amounts of medical images and finding sentences about medical images is constructed, but the present disclosure is not limited thereto. For example, of course, the technique of the present disclosure can also be applied when constructing a derivation model for deriving the feature amounts of photographic images and sentences such as comments on photographic images.

[0109] Also, in the above embodiment, for example, as the hardware structure of the processing unit (Processing Unit) that executes various processes such as the information acquisition unit 21, the first derivation unit 22, the structured information derivation unit 23, the second derivation unit 24, and the learning unit 25 in the learning device 7, and the information acquisition unit 31, the first analysis unit 32, the structured information derivation unit 33, the second analysis unit 34, the specifying unit 35, the display control unit 36, and the search unit 37 in the information processing devices 30 and 30A, the following various processors (Processor) can be used. As described above, in addition to the CPU, which is a general-purpose processor that executes software (program) and functions as various processing units, the above various processors include a programmable logic device (Programmable Logic Device: PLD), which is a processor whose circuit configuration can be changed after manufacture, such as an FPGA (Field Programmable Gate Array), and a dedicated electric circuit, which is a processor having a circuit configuration designed specifically to execute specific processes, such as an ASIC (Application Specific Integrated Circuit).

[0110] One processing unit may be composed of one of these various processors, or may be composed of a combination of two or more processors of the same type or different types (for example, a combination of multiple FPGAs or a combination of a CPU and an FPGA). Also, a plurality of processing units may be composed of one processor. As an example of composing a plurality of processing units with one processor, first, as represented by computers such as clients and servers, there is a form in which one processor is composed of a combination of one or more CPUs and software, and this processor functions as a plurality of processing units. Second, as represented by a system on chip (System On Chip: SoC), there is a form in which a processor that realizes the functions of the entire system including a plurality of processing units with one IC (Integrated Circuit) chip is used. Thus, the various processing units are configured using one or more of the above various processors as the hardware structure.

[0111] Furthermore, as the hardware structure of these various processors, more specifically, an electrical circuit (Circuitry) combined with circuit elements such as semiconductor elements can be used.

[0112] Hereinafter, the appended claims of the present disclosure will be described. (Appended Claim 1) Comprising at least one processor, wherein the processor Derives a first feature amount for an object included in an image by a first neural network, Derives structured information about the sentence by structuring a sentence including a description of an object included in the image, Derives a second feature amount for the sentence from the structured information by a second neural network, When the object included in the image corresponds to the object described in the sentence, in the feature space to which the first feature amount and the second feature amount belong, the distance between the first feature amount and the second feature amount derived is smaller than when the object included in the image does not correspond to the object described in the sentence. By learning the first neural network and the second neural network, a first derivation model for deriving a feature amount for an object included in an image and a second derivation model for deriving a feature amount for a sentence including a description of the object are constructed. (Appended Claim 2) The learning device according to Appended Claim 1, wherein when the object included in the image does not correspond to the object described in the sentence, the processor learns the first neural network and the second neural network so that the distance between the first feature amount and the second feature amount derived is larger than when the object included in the image corresponds to the object described in the sentence in the feature space. (Appended Claim 3) The learning device according to claim 1 or 2, wherein the processor extracts one or more specific expressions related to the object from the text, and determines the factuality of the specific expressions, thereby deriving the specific expressions and the determination result of the factuality as the structured information. (Claim 4) The specific expression represents at least one of a position, a finding, and a size of the object, The learning device according to claim 3, wherein the determination result of the factuality represents any one of positive, negative, and doubtful regarding the finding. (Claim 5) The learning device according to claim 3 or 4, wherein the processor further derives the relationship between the specific expressions as the structured information when a plurality of the specific expressions are extracted. (Claim 6) The learning device according to claim 5, wherein the relationship represents whether each of the plurality of specific expressions is related. (Claim 7) The learning device according to any one of claims 3 to 6, wherein the processor derives normalized structured information by normalizing the specific expressions and the factuality. (Claim 8) The image is a medical image, the object included in the image is a lesion included in the medical image, The learning device according to any one of claims 1 to 7, wherein the text is a finding text in which findings about the lesion are described. (Claim 9) Comprising at least one processor, The processor is Deriving a first feature amount for one or more objects included in the target image by a first derivation model constructed by the learning device according to any one of claims 1 to 8, Deriving structured information about the target text by structuring one or more target texts including descriptions about the object, Deriving a second feature amount for the target text from the structured information about the target text by a second derivation model constructed by the learning device according to any one of claims 1 to 8, Based on the distance in the feature space between the derived first feature quantity and the second feature quantity, identify the first feature quantity corresponding to the second feature quantity, An information processing apparatus that displays, by distinguishing from other regions in the target image, an object from which the identified first feature quantity was derived. (Appended Claim 10) Comprising at least one processor, The processor: Receives an input of a target sentence including a description of an object, Derives structured information about the target sentence by structuring the target sentence, Derives a second feature quantity of the input target sentence from the structured information about the target sentence by a second derivation model constructed by the learning apparatus according to any one of Claims 1 to 8, By referring to a database in which first feature quantities of one or more objects included in each of a plurality of reference images, derived by a first derivation model constructed by the learning apparatus according to any one of Claims 1 to 8, are associated with the respective reference images, identify at least one of the first feature quantities corresponding to the second feature quantity based on the distance in the feature space between the first feature quantities of the plurality of reference images and the derived second feature quantity, An information processing apparatus that identifies a reference image associated with the identified first feature quantity. (Appended Claim 11) The processor is the information processing apparatus according to Claim 9 or 10 that notifies an inherent expression that contributed to the association with the first feature quantity. (Appended Claim 12) Derive a first feature quantity of an object included in an image by a first neural network, Derive structured information about the sentence by structuring a sentence including a description of an object included in the image, Derive a second feature quantity of the sentence from the structured information by a second neural network, When the object included in the image corresponds to the object described in the text, in the feature space to which the first feature amount and the second feature amount belong, the distance between the first feature amount and the second feature amount derived is smaller than the case where the object included in the image does not correspond to the object described in the text. By learning the first neural network and the second neural network, a first derivation model for deriving a feature amount for an object included in an image and a second derivation model for deriving a feature amount for a text including a description about the object are constructed. (Appended Claim 13) Derive a first feature amount for one or more objects included in the target image by the first derivation model constructed by the learning device according to any one of Claims 1 to 8. Derive structured information about the target text by structuring one or more target texts including a description about the object. Derive a second feature amount for the target text from the structured information about the target text by the second derivation model constructed by the learning device according to any one of Claims 1 to 8. Specify the first feature amount corresponding to the second feature amount based on the distance in the feature space between the derived first feature amount and the second feature amount. An information processing method for displaying the object from which the specified first feature amount is derived separately from other regions in the target image. (Appended Claim 14) Receive an input of a target text including a description about the object. Derive structured information about the target text by structuring the target text. Derive a second feature amount for the input target text from the structured information about the target text by the second derivation model constructed by the learning device according to any one of Claims 1 to 8. By referring to a database in which first feature amounts for one or more objects included in each of a plurality of reference images, which are derived by a first derivation model constructed by the learning device according to any one of Supplementary Notes 1 to 8, are associated with the respective reference images, at least one of the first feature amounts corresponding to the second feature amount is specified based on a distance in a feature space between the first feature amounts for the plurality of reference images and the derived second feature amount. An information processing method for specifying a reference image associated with the specified first feature amount. (Supplementary Note 15) A procedure for deriving a first feature amount for an object included in an image by a first neural network, and A procedure for deriving structured information about the sentence by structuring a sentence including a description about an object included in the image, and A procedure for deriving a second feature amount for the sentence from the structured information by a second neural network, and When the object included in the image and the object described in the sentence correspond to each other, in a feature space to which the first feature amount and the second feature amount belong, the distance between the first feature amount and the second feature amount derived is made smaller than when the object included in the image and the object described in the sentence do not correspond to each other. By learning the first neural network and the second neural network, a learning program for causing a computer to execute a procedure for constructing a first derivation model for deriving a feature amount for an object included in an image and a second derivation model for deriving a feature amount for a sentence including a description about the object. (Supplementary Note 16) A procedure for deriving a first feature amount for one or more objects included in a target image by a first derivation model constructed by the learning device according to any one of Supplementary Notes 1 to 8, and A procedure for deriving structured information about the target sentence by structuring one or more target sentences including a description about the object, and A procedure for deriving a second feature quantity for the target sentence from the structured information for the target sentence by means of a second derivation model constructed by the learning device according to any one of Supplementary Notes 1 to 8, A procedure for specifying the first feature quantity corresponding to the second feature quantity based on the distance in the feature space between the derived first feature quantity and the second feature quantity, An information processing program for causing a computer to execute a procedure for displaying, in the target image, the object from which the specified first feature quantity was derived, distinguished from other regions. (Supplementary Note 17) A procedure for receiving an input of a target sentence including a description about an object, A procedure for deriving structured information about the target sentence by structuring the target sentence, A procedure for deriving a second feature quantity for the input target sentence from the structured information for the target sentence by means of a second derivation model constructed by the learning device according to any one of Supplementary Notes 1 to 8, By referring to a database in which first feature quantities for one or more objects included in each of a plurality of reference images, derived by a first derivation model constructed by the learning device according to any one of Supplementary Notes 1 to 8, are associated with each of the reference images, based on the distance in the feature space between the first feature quantities for the plurality of reference images and the derived second feature quantity, a procedure for specifying at least one of the first feature quantities corresponding to the second feature quantity, An information processing program for causing a computer to execute a procedure for specifying a reference image associated with the specified first feature quantity.

Explanation of Signs

[0113] 1 Medical information system 2 Imaging device 3 Reading WS 4 Medical treatment WS 5 Image server 5A Image DB 6 Report server 6A Report DB 7 Learning device 10 Network 11,41 CPU 12 Study Programs 13,43 Storage 14,44 display 15,45 Input Devices 16,46 memory 17,47 Network I / F 18,48 bus 21 Information Acquisition Department 22 First derivation part 23 Structured information derivation part 24 Second Derivation 25 Learning Department 30, 30A Information processing equipment 31 Information Acquisition Department 32 1st Analysis Department 32A First Derived Model 33 Structured information derivation part 34 2nd analysis section 34A Second Derived Model 35 Specific part 36 Display control unit 37 Search Section 42 Information Processing Program 51 Medical Imaging 52 Image interpretation report 53,54 Observations 53A,54A Structured information 55 Tomographic images 55A Object 59 List 61 First Neural Network 62 Second Neural Network 62A Buried Layer 62B Addition mechanism 62C Transformer 65,66,68 feature vector 67 Network Structure 67A RNN layer 67B Attention Mechanism 70, 70A, 90 Image interpretation report creation screen 71,91 Image display area 72,92 Article display area 73,74 Lesion 75 Mark 80 Second neural network 80A Embedding layer 80B GCN layer 81 Feature vector 85 Circle 93 Result display area d1, d2, d3, d4 Distance d5 Radius G0 Target medical image R1, R2, R3 Reference image uw Vector V1, V1-1, V1-2, V1-11~V1-15 First feature quantity V2, V2-1, V2-2 Second feature quantity

Claims

1. Comprising at least one processor, wherein the processor is a learning device that derives a first feature amount for an object included in an image by means of a first neural network, derives structured information about the sentence by structuring a sentence including a description of the object included in the image, derives a second feature amount for the sentence from the structured information by means of a second neural network, when the object included in the image corresponds to the object described in the sentence, in the feature space to which the first feature amount and the second feature amount belong, the distance between the first feature amount derived and the second feature amount is made smaller than when the object included in the image does not correspond to the object described in the sentence, and by learning the first neural network and the second neural network, a first derivation model for deriving a feature amount for an object included in an image and a second derivation model for deriving a feature amount for a sentence including a description of the object are constructed, derives a first feature amount for one or more objects included in a target image by means of the first derivation model constructed by the learning device, derives structured information about the target sentence by structuring one or more target sentences including a description of the object, derives a second feature amount for the target sentence from the structured information about the target sentence by means of the second derivation model constructed by the learning device, identifies the first feature amount corresponding to the second feature amount based on the distance in the feature space between the derived first feature amount and the second feature amount, and an information processing device that displays the object from which the identified first feature amount is derived separately from other regions in the target image.

2. Comprising at least one processor, wherein the processor is a learning device that receives an input of a target sentence including a description of an object, derives structured information about the target sentence by structuring the target sentence, derives a first feature amount for an object included in an image by means of a first neural network, derives structured information about the sentence by structuring a sentence including a description of the object included in the image, derives a second feature amount for the sentence from the structured information by means of a second neural network, ​ When the object included in the image corresponds to the object described in the text, in the feature space to which the first feature amount and the second feature amount belong, the distance between the first feature amount and the second feature amount derived is smaller than when the object included in the image does not correspond to the object described in the text. By learning the first neural network and the second neural network, a learning device that constructs a first derivation model for deriving a feature amount for an object included in an image and a second derivation model for deriving a feature amount for a text including a description of the object, Derive a second feature amount for the input target text from the structured information for the target text by the second derivation model constructed by By referring to a database in which the first feature amounts for one or more objects included in each of a plurality of reference images, derived by the first derivation model constructed by the learning device, are associated with each of the reference images, at least one of the first feature amounts corresponding to the second feature amount is specified based on the distance in the feature space between the first feature amount for the plurality of reference images and the derived second feature amount, An information processing device that specifies a reference image associated with the specified first feature amount.

3. The information processing device according to claim 1 or 2, wherein the processor notifies a unique expression that contributed to the association with the first feature amount.

4. A learning device, Derive a first feature amount for an object included in an image by a first neural network, Derive structured information for the text by structuring the text including a description of the object included in the image, Derive a second feature amount for the text from the structured information by a second neural network, When the object included in the image corresponds to the object described in the text, in the feature space to which the first feature amount and the second feature amount belong, the distance between the first feature amount and the second feature amount derived is smaller than when the object included in the image does not correspond to the object described in the text. By training the first neural network and the second neural network, a first derivation model for deriving a feature amount of an object included in an image and a second derivation model for deriving a feature amount of a text including a description of the object are constructed. A learning device, Derive a first feature amount for one or more objects included in the target image by the first derivation model constructed by, Derive structured information about the target text by structuring one or more target texts including descriptions of the object, Derive a second feature amount for the target text from the structured information about the target text by the second derivation model constructed by the learning device, Specify the first feature amount corresponding to the second feature amount based on the distance in the feature space between the derived first feature amount and the second feature amount, An information processing method for displaying, in the target image, the object from which the specified first feature amount is derived, distinguished from other regions.

5. Receive an input of a target text including a description of an object, Derive structured information about the target text by structuring the target text, A learning device, Derive a first feature amount of an object included in an image by a first neural network, Derive structured information about the text by structuring the text including a description of the object included in the image, Derive a second feature amount of the text from the structured information by a second neural network, When the object included in the image corresponds to the object described in the text, in the feature space to which the first feature amount and the second feature amount belong, the distance between the first feature amount and the second feature amount derived is smaller than when the object included in the image does not correspond to the object described in the text. By learning the first neural network and the second neural network, a learning device that constructs a first derivation model for deriving a feature amount of an object included in an image and a second derivation model for deriving a feature amount of a text including a description of the object. Derive a second feature amount of the input target text from the structured information of the target text by the second derivation model constructed by By referring to a database in which the first feature amounts of one or more objects included in each of a plurality of reference images derived by the first derivation model constructed by the learning device are associated with each of the reference images, based on the distance in the feature space between the first feature amounts of the plurality of reference images and the derived second feature amount, identify at least one of the first feature amounts corresponding to the second feature amount. An information processing method for identifying a reference image associated with the identified first feature amount.

6. A learning device, Derive a first feature amount of an object included in an image by a first neural network, Derive structured information about the text by structuring the text including a description of the object included in the image, Derive a second feature amount of the text from the structured information by a second neural network, When the object included in the image corresponds to the object described in the text, in the feature space to which the first feature amount and the second feature amount belong, the distance between the first feature amount and the second feature amount derived is smaller than when the object included in the image does not correspond to the object described in the text. By learning the first neural network and the second neural network, a learning device that constructs a first derivation model for deriving a feature amount of an object included in an image and a second derivation model for deriving a feature amount of a text including a description of the object. The procedure of deriving first feature amounts for one or more objects included in a target image by means of a first derivation model constructed thereby, and the procedure of deriving structured information about the target text by structuring one or more target texts including descriptions about the object, and the procedure of deriving second feature amounts for the target text from the structured information about the target text by means of a second derivation model constructed by the learning device, and the procedure of specifying the first feature amount corresponding to the second feature amount based on the distance in the feature space between the derived first feature amount and the second feature amount, and An information processing program for causing a computer to execute the procedure of displaying, in the target image, the object from which the specified first feature amount was derived separately from other areas.

7. The procedure of receiving an input of a target text including a description about an object, and the procedure of deriving structured information about the target text by structuring the target text, and A learning device, deriving first feature amounts for objects included in an image by means of a first neural network, deriving structured information about the text by structuring the text including a description about an object included in the image, deriving second feature amounts for the text from the structured information by means of a second neural network, When the object included in the image corresponds to the object described in the text, in the feature space to which the first feature amount and the second feature amount belong, the distance between the first feature amount and the second feature amount to be derived is made smaller than when the object included in the image does not correspond to the object described in the text. By learning the first neural network and the second neural network, a learning device that constructs a first derivation model for deriving feature amounts for objects included in an image and a second derivation model for deriving feature amounts for a text including a description about an object, the procedure of deriving second feature amounts for the input target text from the structured information about the target text by means of the second derivation model constructed thereby, By referring to a database in which first feature amounts for one or more objects included in each of a plurality of reference images, derived by a first derivation model constructed by the learning device, are associated with the respective reference images, based on distances in a feature space between the first feature amounts for the plurality of reference images and the derived second feature amounts, a procedure for specifying at least one of the first feature amounts corresponding to the second feature amount; An information processing program that causes a computer to execute a procedure for specifying a reference image associated with the specified first feature amount.

Citation Information

Patent Citations

  • Learning data generation support apparatus and learning data generation support method and learning data generation support program

    JP2019008349A

  • Method and device for presenting high risk patient having high possibility of causing skeletal-related event

    JP2021002334A

  • Efficient cross-modal retrieval via deep binary hashing and quantization

    JP2021099803A

  • Medical document creation device, method, and program, learning device, method, and program, and learned model

    WO2020241857A1

  • Information processing method, program, and information processing device

    JP2020013594A