Information processing device, method, and program

The information processing device accurately derives sentence features by dividing sentences into units, determining attributes and relationships, and using machine learning to weight these units, addressing the need for extensive training data in existing methods.

JP7838936B2Active Publication Date: 2026-04-01FUJIFILM CORP
View PDF 11 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-06
Publication Date
2026-04-01

AI Technical Summary

Technical Problem

Existing methods require a large amount of training data to accurately derive sentence features, making it difficult to construct analytical models for tasks such as sentence classification.

Method used

An information processing device that divides sentences into units, determines attributes, factual nature, and relationships of each unit, and uses a derivation model through machine learning to weight and derive sentence features without requiring extensive training data.

Benefits of technology

Enables accurate derivation of sentence features without the need for large amounts of training data, facilitating tasks like identifying disease names and searching for corresponding images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007838936000001
    Figure 0007838936000001
  • Figure 0007838936000002
    Figure 0007838936000002
  • Figure 0007838936000003
    Figure 0007838936000003
Patent Text Reader

Abstract

To provide an information processing device, method and program, which allow for accurately deriving feature quantities of sentences without having to prepare a large amount of teacher data.SOLUTION: A processor of an information processing device provided herein is configured to divide a sentence into predetermined units, determine at least one of attributes, factuality and relationships of each unit, determine the weight of each unit according to a determination result, derive a feature quantity of each unit using a derivation model built by machine learning, and weighting the feature quantity of each unit according to the weight to derive a feature quantity of the sentence.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus, method, and program.

Background Art

[0002] Various methods for analyzing sentences have been proposed. For example, Patent Document 1 proposes a method of structuring and presenting terms belonging to categories such as part names, disease names, and sizes by analyzing a plurality of terms extracted from a radiology report. In addition, Patent Document 2 proposes a method of creating a summary of a medical record using importance information indicating the importance of each element generated based on the elements constituting the medical record text obtained by performing text analysis on the medical record text and the summary text corresponding to the medical record text.

[0003] In addition, a method of deriving feature amounts of a sentence by analyzing the sentence and performing various inferences based on the feature amounts has been proposed. For example, a method of deriving feature amounts of a sentence by analyzing the sentence and performing a task of classifying the sentence based on the feature amounts has been proposed (see Non-Patent Document 1). In the method as described in Non-Patent Document 1, an analysis model constructed by machine learning a neural network using teacher data associating the sentence and the classification result is used. )]]

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Non-Patent Documents

[0005]

Non-Patent Document 1

[0006] On the other hand, a large amount of training data is required to derive sentence features that can accurately perform various tasks. However, since the number of sentences is limited, it is difficult to prepare a large amount of training data. For this reason, it is difficult to construct an analytical model that can accurately derive sentence features.

[0007] This disclosure is made in light of the circumstances described above, and aims to enable the accurate deriving of textual features without the need to prepare a large amount of training data. [Means for solving the problem]

[0008] The information processing device according to this disclosure comprises at least one processor, The processor divides the statement into predetermined units, Determine at least one of the attributes, factual nature, and relationships of each unit, Based on the results of the assessment, determine the weight for each unit. Using a derivation model constructed through machine learning, the features of each unit are derived, and then the features of the sentence are derived by weighting the features of each unit based on the weights.

[0009] Furthermore, in the information processing device described herein, the processor may determine the weights for words determined to have predetermined attributes, facts, and relationships to be greater than the weights for words determined to have attributes, facts, and relationships other than those predetermined.

[0010] Furthermore, in the information processing device described herein, predetermined attributes, facts, and relationships may be determined according to a task using the derived sentence features.

[0011] Furthermore, in the information processing apparatus according to this disclosure, the processor may determine the weights using a derived model.

[0012] Furthermore, in the information processing device described herein, the processor may display sentences by emphasizing words determined to have predetermined attributes, facts, and relationships.

[0013] Furthermore, in the information processing device according to this disclosure, the processor may perform the task of identifying the content of a sentence using the derived feature quantities.

[0014] Furthermore, in the information processing device described herein, the processor may perform the task of searching for an image corresponding to a sentence using the derived feature quantities.

[0015] The information processing method disclosed herein divides a sentence into predetermined units, Determine at least one of the attributes, factual nature, and relationships of each unit, Based on the results of the assessment, determine the weight for each unit. Using a derivation model constructed through machine learning, the features of each unit are derived, and then the features of the sentence are derived by weighting the features of each unit based on the weights.

[0016] Note that it may be provided as a program for causing a computer to execute the information processing method according to the present disclosure.

Advantages of the Invention

[0017] According to the present disclosure, even without preparing a large amount of teacher data, feature amounts of sentences can be accurately derived.

Brief Description of the Drawings

[0018] [Figure 1] Diagram showing a schematic configuration of a medical information system to which an information processing apparatus according to a first embodiment of the present disclosure is applied [Figure 2] Diagram showing a schematic configuration of an information processing apparatus according to a first embodiment [Figure 3] Functional configuration diagram of an information processing apparatus according to a first embodiment [Figure 4] Diagram showing a creation screen of a radiology report in a first embodiment [Figure 5] Diagram for explaining determination of attributes and factuality in a first embodiment [Figure 6] Diagram schematically showing a derivation model [Figure 7] Diagram schematically showing another example of a derivation model <00001​​​​​​​​​​​​​​​​​​​​​A diagram illustrating the determination of attributes, facts, and relationships in the third embodiment. [Figure 16] A diagram schematically showing the weighting in the derivation model in the third embodiment. [Modes for carrying out the invention]

[0019] Embodiments of this disclosure will be described below with reference to the drawings. First, the configuration of the medical information system 1 to which the information processing device according to the first embodiment is applied will be described. Figure 1 is a diagram showing the schematic configuration of the medical information system 1. The medical information system 1 shown in Figure 1 is a system for taking photographs of the subject's area to be examined, storing the medical images acquired by the photographs, interpreting the medical images by a radiologist and creating an interpretation report, and allowing the requesting physician in the clinical department to view the interpretation report and observe the medical images of the subject in detail, based on an examination order from a physician in a clinical department using a known ordering system.

[0020] Each device is a computer on which an application program is installed to function as a component of the medical information system 1. The application program is stored in a memory device of a server computer connected to network 10, or in network storage, in a state that is accessible from the outside, and is downloaded and installed on the computer upon request. Alternatively, it may be recorded on a recording medium such as a DVD (Digital Versatile Disc) or CD-ROM (Compact Disc Read Only Memory) and distributed, and then installed on the computer from that recording medium.

[0021] The imaging device 2 is a device (modality) that generates a medical image representing the area to be diagnosed by imaging the area of ​​the subject. Specifically, this includes plain X-ray imaging devices, CT (Computed Tomography) devices, MRI (Magnetic Resonance Imaging) devices, and PET (Positron Emission Tomography) devices. The medical image generated by imaging device 2 is transmitted to the image server 5 and stored in the image database 6.

[0022] The Image Interpretation WS3 is a computer used, for example, by a radiologist in the radiology department for interpreting medical images and creating interpretation reports, and it incorporates the information processing device 20 according to this embodiment. The Image Interpretation WS3 performs tasks such as requesting the image server 5 to view medical images, performing various image processing on medical images received from the image server 5, displaying medical images, and accepting input of findings text related to medical images. The Image Interpretation WS3 also performs analysis processing on medical images and input findings text, assists in creating interpretation reports based on the analysis results, requests the report server 7 to register and view interpretation reports, and displays interpretation reports received from the report server 7. These processes are performed by the Image Interpretation WS3 executing software programs for each process.

[0023] The Clinical WS4 is a computer used by physicians in various departments for detailed examination of images, viewing of image interpretation reports, and creation of electronic medical records. It consists of a processing unit, display devices such as a display, and input devices such as a keyboard and mouse. The Clinical WS4 handles requests to view images from the Image Server 5, displays images received from the Image Server 5, requests to view image interpretation reports from the Report Server 7, and displays image interpretation reports received from the Report Server 7. These processes are carried out by the Clinical WS4 executing software programs for each process.

[0024] Image server 5 is a general-purpose computer with software programs installed that provide the functionality of a database management system (DBMS). Image server 5 also includes storage that constitutes image DB 6. This storage may be a hard disk drive connected to image server 5 via a data bus, or a disk drive connected to a NAS (Network Attached Storage) or SAN (Storage Area Network) connected to network 10. Furthermore, when image server 5 receives a request to register a medical image from imaging device 2, it formats the medical image into a database format and registers it in image DB 6.

[0025] Image DB6 stores image data and associated information of medical images acquired by imaging device 2. The associated information includes, for example, an image ID (identification) to identify individual medical images, a patient ID to identify the subject, an examination ID to identify the examination, a unique ID (UID: unique identification) assigned to each medical image, the date and time of the examination in which the medical image was generated, the type of imaging device used in the examination to acquire the medical image, patient information such as patient name, age, and gender, the examination site (imaging site), imaging information (imaging protocol, imaging sequence, imaging method, imaging conditions, use of contrast agent, etc.), and information such as the series number or acquisition number if multiple medical images were acquired in a single examination.

[0026] Furthermore, when the image server 5 receives a viewing request from the image interpretation WS3 and the medical treatment WS4 via the network 10, it searches for medical images registered in the image database 6 and sends the retrieved medical images to the requesting image interpretation WS3 and medical treatment WS4.

[0027] The report server 7 incorporates software programs that provide the functionality of a database management system to a general-purpose computer. When the report server 7 receives a request to register an image interpretation report from the image interpretation WS3, it formats the image interpretation report into a database format and registers it in the report DB8.

[0028] The report DB8 registers image interpretation reports that include at least the findings statement created in the image interpretation WS3. The image interpretation report may include, for example, the medical image being interpreted, an image ID to identify the medical image, a radiologist ID to identify the radiologist who performed the interpretation, the name of the lesion, location information of the lesion, information for accessing the medical image containing a specific region, and characteristic information.

[0029] Furthermore, when the report server 7 receives a request to view an image interpretation report from the image interpretation WS3 and medical treatment WS4 via the network 10, it searches the report DB8 for the image interpretation report and sends the retrieved report to the requesting image interpretation WS3 and medical treatment WS4.

[0030] Furthermore, medical images are not limited to CT images; any medical images, such as MRI images and simple two-dimensional images obtained using a simple X-ray machine, can be used.

[0031] Network 10 is a wired or wireless local area network that connects various devices within the hospital. If the image interpretation WS3 is installed in another hospital or clinic, Network 10 may be configured to connect the local area networks of each hospital via the Internet or a dedicated line.

[0032] Next, an information processing device according to the first embodiment will be described. Figure 2 illustrates the hardware configuration of the information processing device according to the first embodiment. As shown in Figure 2, the information processing device 20 includes a CPU (Central Processing Unit) 11, non-volatile storage 13, and memory 16 as a temporary storage area. The information processing device 20 also includes a display 14 such as a liquid crystal display, input devices 15 such as a keyboard and mouse, and a network I / F (Interface) 17 connected to a network 10. The CPU 11, storage 13, display 14, input devices 15, memory 16, and network I / F 17 are connected to a bus 18. Note that the CPU 11 is an example of a processor in this disclosure.

[0033] Storage 13 is implemented by HDD (Hard Disk Drive), SSD (Solid State Drive), flash memory, etc. Information processing programs are stored in storage 13 as a storage medium. The CPU 11 reads the information processing program 12 from storage 13, expands it into memory 16, and executes the expanded information processing program 12.

[0034] Next, the functional configuration of the information processing device according to the first embodiment will be described. Figure 3 is a diagram showing the functional configuration of the information processing device according to the first embodiment. As shown in Figure 3, the information processing device 20 includes an information acquisition unit 21, a division unit 22, a determination unit 23, an analysis unit 24, and a display control unit 25. The CPU 11 executes the information processing program 12, and the CPU 11 functions as the information acquisition unit 21, the division unit 22, the determination unit 23, the analysis unit 24, and the display control unit 25. The information processing device according to this embodiment performs the task of identifying the disease name expressed in the findings statement created by the radiologist. The task of identifying the disease name is an example of a task that identifies the content of a statement.

[0035] The information acquisition unit 21 acquires the target medical image G0 for creating a radiology report from the image server 5 based on instructions from the radiologist operator via the input device 15. The target medical image G0 is displayed on the display 14 by the display control unit 25. Specifically, the target medical image G0 is displayed on the display 14 on the radiology report creation screen.

[0036] Figure 4 shows the screen for creating a medical imaging report. As shown in Figure 4, the creation screen 50 has an image display area 51, a text display area 52, and a specific result display area 53. The image display area 51 displays the target medical image G0 acquired by the information acquisition unit 21. In Figure 4, the target medical image G0 is a single tomographic image that constitutes a three-dimensional image of the chest. The text display area 52 displays the findings entered by the physician. As shown in Figure 4, the findings are: "A 13 mm partially solid nodule is observed in the right lung S8. Spicula are observed at the periphery, and bronchial radiolucency is observed internally." The specific result display area 53 will be described later.

[0037] The division unit 22 divides the findings statement into predetermined units. Specifically, it divides the findings statement into words by performing morphological analysis on the findings statement. As a result, the findings statement displayed in the text display area 52 is divided into the words: "A 13mm partially solid nodule is observed in the right lung / S8. A spicula is observed at the margin, and a bronchial radiolucency is observed inside." The division unit 22 may also divide the findings statement into phrases.

[0038] The determination unit 23 determines at least one of the attributes, factuality, and relationships of each divided word. In this embodiment, the determination unit 23 determines the attributes and factuality of each divided word. Figure 5 is a diagram illustrating the determination of attributes and factuality. Regarding attributes, the determination unit 23 determines, for example, whether a word has the attributes of location, size, characteristics, lesion, change, or disease name. For example, regarding the above finding statement, the determination unit 23 determines that "right lung," "S8," "margin," and "internal" have the attribute of location, "13 mm" has the attribute of size, "partially solid," "spicule," and "bronchial radiolucency" have the attribute of characteristics, and "nodule" has the attribute of lesion. Note that the types of attributes are not limited to location, size, characteristics, lesion, change, and disease name.

[0039] Regarding factual accuracy, the determination unit 23 determines whether the attributes of the characteristics, lesion, and disease name indicate a negative, positive, or suspected result. In the findings statement above, the words determined to have the attributes of characteristics, lesion, and disease name are "partially solid type," "nodule," "spicule," and "bronchial radiolucency." Since all of these words end in "observed" in the context, they are all positive. Therefore, the determination unit 23 determines that the factual accuracy of the words "partially solid type," "nodule," "spicule," and "bronchial radiolucency" is all positive. In Figure 5, a positive result is indicated by adding a "+" sign after the attribute. Alternatively, a "-" sign could be added for a negative result, and a "±" sign for a suspected result.

[0040] The analysis unit 24 derives the feature quantities of each word, determines the weights for each word according to the judgment result of the judgment unit 23, and derives the feature quantities of the observation text by performing a weighted calculation on the feature quantities of each word based on the determined weights. To this end, the analysis unit 24 derives the feature quantities of the observation text using a derivation model constructed by machine learning a neural network.

[0041] Figure 6 is a schematic diagram of the derivation model. As shown in Figure 6, the derivation model 30 has an embedding layer 31, a recurrent neural network (RNN) layer 32, a weighting calculation mechanism 33, and a multi-layer perceptron (MLP) 34. The analysis unit 24 inputs each word 40 derived by the division unit 22 by dividing the observation sentence into the embedding layer 31. The embedding layer 31 outputs a feature vector 41 for each word 40. In Figure 6, the feature vectors are shown as black circles. The feature vector 41 for each word 40 is an n-dimensional vector. The RNN layer 32 outputs a feature vector 42 for each word 40, taking into account the context of the feature vectors 41 output by the embedding layer 31. The weighting calculation mechanism 33 derives the feature vectors of the observation sentence input to the analysis unit 24 as feature quantities V0 by performing a weighting calculation on the feature vectors 42 for each word 40. Feature vector V0 is also an n-dimensional vector. MLP34 outputs the disease name represented by the findings statement input from feature vector V0 as the identified result 44.

[0042] The weighting calculation mechanism 33 determines the weights to be used when performing weighting calculations on the feature vectors 42 of each word 40. In this embodiment, the weighting calculation mechanism 33 determines that the weights for words determined to have predetermined attributes and facts are greater than the weights for words determined to have attributes and facts other than those predetermined. The predetermined attributes can be, for example, characteristics and disease names. The predetermined facts can be positive.

[0043] Therefore, the weighting calculation mechanism 33 determines the weights such that the weights for "partially rich," "nodule," "spicule," and "bronchial radiolucency" are greater than the weights for the other words. In Figure 6, the arrows in the weighting calculation mechanism 33 for the feature vectors 42 of "partially rich," "nodule," "spicule," and "bronchial radiolucency" output by the RNN layer 32 are made thicker than the other arrows, indicating that the weights for "partially rich," "nodule," "spicule," and "bronchial radiolucency" are greater than the weights for the other words. In addition, in Figure 6, the weights for "partially rich," "nodule," "spicule," and "bronchial radiolucency" among the words 40 input to the embedding layer 31 are also shown to be greater than the weights for the other words by making them bold.

[0044] In this embodiment, the sum of the weights is set to 1. The weighting calculation mechanism 33 determines the weights such that, for example, 80% of the weights are assigned to predetermined attributes and facts. In this embodiment, the observation sentence is divided into 20 words by the division unit 22, of which four words are determined to be predetermined attributes and facts: "partially enriched," "nodule," "spicule," and "bronchial radiance." Therefore, the weighting calculation mechanism 33 determines the weight for each of these four words to be 0.8 / 4 = 0.2. In addition, the weight for words other than "partially enriched," "nodule," "spicule," and "bronchial radiance" is determined to be 0.2 / 16 = 0.0125.

[0045] Then, the weighting calculation mechanism 33 derives the feature quantity V0 by weighting and adding all the feature vectors 42 according to the determined weights.

[0046] In this embodiment, the analysis unit 24 performs the task of identifying the disease name in the findings statement using the feature quantity V0. For this reason, the weights of attributes and factual information that are less relevant to the task may be determined to be smaller. For example, the attributes necessary for identifying the disease name are characteristics and lesions. For this reason, the weights of words for which characteristics and lesions are positive attributes may be made larger than the weights of other words.

[0047] On the other hand, as shown in Figure 7, the weighting calculation mechanism 33 may be provided with a context vector 36 that has been trained to give higher weights to predetermined attributes and facts. The context vector 36 is trained so that attributes and facts that contribute more to the derivation of the feature quantity V0, i.e., predetermined attributes and facts, are assigned larger weights. The context vector 36 then outputs a vector that gives larger weights to predetermined attributes and facts. The weighting calculation mechanism 33 determines the weight for each feature vector 42 by taking the dot product of the vector derived by the context vector 36 and each feature vector 42.

[0048] In Figure 7, arrows are added only to the feature vectors 42 of "partially enriched," "nodule," "spicule," and "bronchial radiolucency," which have larger weights due to the context vector 36, to indicate that the weights from the context vector 36 are applied. However, the weights from the context vector 36 are also applied to the feature vectors 42 of other words.

[0049] Furthermore, the context vector 36 may be constructed by learning to reduce the weight of attributes and facts that are less relevant to the task.

[0050] MLP34 is a fully connected neural network that, when given feature vectors V0 as input, is built by machine learning the network to identify the disease name expressed in the findings text. The disease name identified by MLP34 also includes the possibility that the condition is benign.

[0051] MLP34 outputs scores for multiple types of lung disease, identifies the disease with the highest score as the disease in the findings, and outputs identification result 44. For example, if MLP34 is trained to identify three disease names: "benign," "adenocarcinoma," and "squamous cell carcinoma," MLP34 will output scores for "benign," "adenocarcinoma," and "squamous cell carcinoma." Then, MLP34 identifies the disease name with the highest score as the disease name in the findings. For example, if the scores for "benign," "adenocarcinoma," and "squamous cell carcinoma" are 0.1, 0.8, and 0.1 respectively, MLP34 will identify "adenocarcinoma" as the disease name and output identification result 44.

[0052] The display control unit 25 displays the target medical image G0 and the findings text, as well as the disease name identification result, on the image interpretation report creation screen. Figure 8 shows the image interpretation report creation screen with the disease name identification result displayed. As shown in Figure 8, the display control unit 25 displays the disease name represented by the findings text identified by the analysis unit 24 in the identification result display area 53. In Figure 8, the disease name is "lung adenocarcinoma".

[0053] Furthermore, the display control unit 25 may highlight words in the findings text displayed in the text display area 52 that have been determined to have predetermined attributes and factual accuracy. In Figure 8, the words "partially enriched," "nodule," "spicule," and "bronchial radiolucency" are highlighted by underlining them. Note that highlighting is not limited to underlining; it may also be done by emphasizing characters or adding markers to words.

[0054] Furthermore, when weighting is performed using context vector 36, the degree of emphasis on words may differ depending on the magnitude of the vector output by context vector 36. For example, suppose the magnitudes of the vectors output by context vector 36 for "partially full," "nodule," "spicule," and "bronchial radiolucency" are "partially full" = "nodule" < "spicule" = "bronchial radiolucency." In this case, as shown in Figure 9, the degree of emphasis on the words "spicule" and "bronchial radiolucency" in the findings text may be greater than the degree of emphasis on "partially full" and "nodule." Note that in Figure 9, the difference in the degree of emphasis is indicated by the difference in the spacing of the hatching lines.

[0055] Next, the processing performed in the first embodiment will be described. Figure 10 is a flowchart showing the processing performed in the first embodiment. The target medical image G0 is assumed to be acquired from the image server 5 by the information acquisition unit 21 and stored in the storage 13.

[0056] First, the display control unit 25 displays the screen for creating the image interpretation report (step ST1) and accepts input of the findings (step ST2). Next, the division unit 22 divides the findings into words (step ST3), and the determination unit 23 determines at least one of the attributes, factuality, and relationships of each divided word (step ST4).

[0057] Then, the weighting calculation mechanism 33 of the analysis unit 24 determines the weight for each word according to the judgment result (step ST5), and derives the feature quantity V0 of the findings text by performing a weighting calculation on the feature quantity for each word based on the determined weight (step ST6). Furthermore, the MLP 34 of the analysis unit 24 identifies the disease name represented by the findings text based on the feature quantity V0 (step ST7). Then, the display control unit displays the identified disease name on the image interpretation report creation screen 50 (step ST8), and the process ends.

[0058] Thus, in this embodiment, the findings are divided into predetermined units, such as words, at least one of the attributes, factuality, and relationships of each unit is determined, a weight for each unit is determined according to the determination result, and the features of each unit are derived by performing a weighted calculation based on the determined weights. For this reason, it is possible to derive features V0 that are effective for tasks such as identifying the disease name represented by the findings without having to construct a derivation model 30 using a large amount of training data.

[0059] Furthermore, by displaying the weights determined by the weighting calculation mechanism 33, it is easy to confirm the units that were used as the basis for deriving the features.

[0060] Next, a second embodiment of the information processing device according to the present disclosure will be described. Figure 11 is a functional configuration diagram of the information processing device according to the second embodiment. In Figure 11, the same reference numerals are used for components identical to those in Figure 3, and detailed explanations are omitted. As shown in Figure 11, the information processing device 20A according to the second embodiment differs from the first embodiment in that it further includes a search unit 26 to perform the task of searching for medical images. In the second embodiment, the analysis unit 24 performs the process up to the derivation of the feature quantity V0. Therefore, the derivation model 30 according to the second embodiment may not include the MLP 34.

[0061] In the second embodiment, it is assumed that the image DB6 stores a large number of medical images, each associated with a feature vector V1. The feature vector V1 of the stored medical images is derived by an unillustrated derivation model that has been machine-learned to derive feature vector V1 from the medical images. The feature vector V1 of the medical image and the feature vector V0 derived by the analysis unit 24 are n-dimensional vectors distributed in the same feature space. The medical images stored in the image DB6 will be referred to as reference images in the following description.

[0062] Furthermore, in the information processing device 20A according to the second embodiment, similar to the first embodiment, the radiologist interprets the target medical image G0 in the image interpretation WS3 and inputs a report containing the interpretation results using the input device 15. The division unit 22 divides the input report into words, and the determination unit 23 derives the attribute and factual determination results for each word. Then, the analysis unit 24 performs a weighting calculation of the feature vector 41 of each word, similar to the first embodiment, to derive the feature vector of the report as a feature quantity V0.

[0063] The search unit 26 refers to the image DB6 and searches for a reference image in the feature space that has a corresponding feature quantity V1 that is close in distance to the feature quantity V0 derived by the analysis unit 24. Figure 12 is a diagram illustrating the search performed in the information processing device 20A according to the second embodiment. In Figure 12, the feature space is shown in two dimensions for illustrative purposes. Also, five feature quantities V1-1 to V1-5 are plotted in the feature space for illustrative purposes.

[0064] The search unit 26 identifies features in the feature space whose distance from feature V0 is within a predetermined threshold. In Figure 12, a circle 60 with radius d1 centered on feature V0 is shown. The search unit 26 identifies features contained within circle 60 in the feature space. In Figure 12, three features V1-1 to V1-3 are identified.

[0065] The search unit 26 searches the image database 6 for reference images associated with the identified feature quantities V1-1 to V1-3, and retrieves the retrieved reference images from the image server 5.

[0066] The display control unit 25 displays the acquired reference image on the image interpretation report creation screen. Figure 13 shows the image interpretation report creation screen in the second embodiment. As shown in Figure 13, the creation screen 70 has an image display area 71, a text display area 72, and a result display area 73. The target medical image G0 is displayed in the image display area 71. The findings entered by the radiologist are displayed in the text display area 72. In Figure 13, the findings are displayed as, "There is a 10 mm solid nodule in the right lung S6."

[0067] The result display area 73 displays the reference images found by the search unit 26. In Figure 13, three reference images R1 to R3 are displayed in the result display area 73.

[0068] Next, the processing performed in the second embodiment will be described. Figure 14 is a flowchart showing the processing performed in the second embodiment. The target medical image G0 is assumed to be acquired from the image server 5 by the information acquisition unit 21 and stored in the storage 13.

[0069] First, the display control unit 25 displays the screen for creating the image interpretation report (step ST11) and accepts input of the findings text (step ST12). Next, the division unit 22 divides the findings text into words (step ST13), and the determination unit 23 determines at least one of the attributes, factuality, and relationships of each divided word (step ST14). Then, the weighting calculation mechanism 33 of the analysis unit 24 determines a weight for each word according to the determination result (step ST15), and derives the feature quantity V0 of the findings text by performing a weighting calculation on the feature quantities for each word based on the determined weights (step ST16).

[0070] Next, the search unit 26 refers to the image DB6 and searches for a reference image associated with a feature V1 that is close in distance to feature V0 (step ST17). Then, the display control unit 25 displays the retrieved reference image on the display 14 (step ST18), and the process ends.

[0071] In the second embodiment, the reference images R1 to R3 retrieved are medical images that have similar characteristics to the findings entered by the radiologist. Since the findings relate to the target medical image G0, the reference images R1 to R3 are similar in case to the target medical image G0. Therefore, according to the second embodiment, the interpretation of the target medical image G0 and the creation of an interpretation report can be performed by referring to reference images with similar cases. In addition, the interpretation report for the reference images can be obtained from the report server 7 and used to create an interpretation report for the target medical image G0.

[0072] In the first and second embodiments described above, weights are determined according to attributes and facts, but the invention is not limited to these. In addition to attributes and facts, relationships may also be used to determine the weights. This will be described below as the third embodiment.

[0073] Figure 15 is a diagram illustrating the determination of attributes, facts, and relationships. In the third embodiment, the findings statement is assumed to be "A 13 mm partially solid nodule is observed in the right lung S8. Spicula are observed at the margin. Scarring is present at the left lung apex." Furthermore, the divided section 22 is assumed to be a division of the findings statement into "A 13 mm partially solid nodule is observed in the right lung / S8 / . Spicula are observed at the margin. Scarring is present at the left lung apex."

[0074] In the third embodiment, the determination unit 23 determines the relationships in addition to the attributes and factual nature of each divided word. With respect to the above findings sentence, the determination unit 23 determines that "right lung," "S8," "margin," and "left lung apex" have the attribute of location, "13 mm" has the attribute of size, "partially solid," "spicule," and "scar pattern" have the attribute of characteristics, and "nodule" has the attribute of lesion.

[0075] Regarding the factual accuracy, the determination unit 23 determines that the factual accuracy of the words "partially solid type," "nodule," "spicule," and "scar pattern" is all positive.

[0076] Regarding relationships, the determination unit 23 derives the relationships between words. For example, among the words included in the findings statement, the word "nodule," which is related to the task of identifying the disease name in the first embodiment, is related to the size attribute word "13 mm," the location attributes "right lung" and "S8," and the characteristic attributes "partially solid" and "spicule," but not to the location attribute "left lung apex" or the characteristic attribute "scarred appearance." Also, the characteristic attribute "spicule" is related to the location attribute "margin," but not to the location attribute "left lung apex" or the characteristic attribute "scarred appearance." Furthermore, the characteristic attribute "scarred appearance" is related to the location attribute "left lung apex."

[0077] The relationships can be derived by referring to a predefined table that shows whether or not there are relationships between a large number of words. Alternatively, the relationships can be derived using a derivation model constructed by machine learning to output whether or not there are relationships between words. Another approach is to identify words related to the task of identifying disease names as keywords, and then identify all words that modify these keywords as related words.

[0078] In the third embodiment, the weighting calculation mechanism 33 of the derived model 30 identifies words related to the task performed by the information processing device, and determines that the weights of the identified words determined to have predetermined attributes and facts are greater than the weights of the words determined to have attributes and facts other than the predetermined attributes and facts. Figure 16 is a diagram illustrating the weighting in the second embodiment. Note that in Figure 16, the same reference numerals are used for components identical to those in Figure 6, and detailed explanations are omitted here.

[0079] Here, the word related to the task is "nodule," and the words related to "nodule" are "13mm," "right lung," "S8," "partially solid," and "spicule." Therefore, the weighting calculation mechanism 33 determines the weights such that, in addition to "nodule," the weights for "partially solid" and "spicule" among "13mm," "right lung," "S8," and "partially solid" are greater than the weights for the other words. In Figure 16, the arrows in the weighting calculation mechanism 33 for the feature vectors 42 of "partially solid," "nodule," and "spicule" output by the RNN layer 32 are thicker than the other arrows, indicating that the weights for "partially solid," "nodule," and "spicule" are greater than the weights for the other words.

[0080] As in the third embodiment, by determining the weight for each unit using relationships in addition to attributes and factual information, a feature quantity V0 that is effective for tasks such as identifying the disease name expressed in the findings can also be derived.

[0081] In the first and second embodiments described above, the attributes and factual nature of each divided word are determined, and in the third embodiment described above, the attributes, factual nature, and relationships of each divided word are determined, but the method is not limited to these. Only one of the attributes, factual nature, and relationships of each divided word may be determined, or any combination of any two of these may be determined.

[0082] Furthermore, in the first embodiment described above, a derivation model for deriving feature quantities of the observation text for medical images is used to identify the disease name represented by the observation text, but the invention is not limited to this. For example, the technology of this disclosure can of course be applied to tasks such as identifying the content of comments using a derivation model for deriving feature quantities of text such as comments on photographic images.

[0083] Furthermore, in the above embodiment, the hardware structure of the Processing Unit, which executes various processes such as the information acquisition unit 21, the division unit 22, the determination unit 23, the analysis unit 24, the display control unit 25, and the search unit 26, can be the various processors shown below. As mentioned above, these various processors include a CPU, which is a general-purpose processor that executes software (programs) and functions as various processing units, as well as a Programmable Logic Device (PLD), which is a processor whose circuit configuration can be changed after manufacturing, such as an FPGA (Field Programmable Gate Array), and a dedicated electrical circuit, which is a processor with a circuit configuration specifically designed to execute a particular process, such as an ASIC (Application Specific Integrated Circuit).

[0084] A single processing unit may be composed of one of these various processors, or it may be composed of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs or a combination of a CPU and an FPGA). Alternatively, multiple processing units may be composed of a single processor. Examples of composing multiple processing units with a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, as is typical of computers such as client and server systems, and this processor functions as multiple processing units. Secondly, a configuration using a processor that realizes the functions of the entire system, including multiple processing units, on a single IC (Integrated Circuit) chip, as is typical of a System on a Chip (SoC). Thus, various processing units are configured as hardware structures using one or more of the above-mentioned various processors.

[0085] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits (Circuitry) that combine circuit elements such as semiconductor elements. [Explanation of symbols]

[0086] 1. Medical Information System 2. Imaging device 3 Image Interpretation Workshop 4 Clinical Department WS 5 Image Server 6 Image Database 7. Report Server 8 Report Database 10 Networks 11 CPU 12. Information Processing Programs 13 Storage 14 displays 15 Input Devices 16 memory 17 Network Interface 18 bus 21 Information Acquisition Department 22 Division 23 Judgment section 24 Analysis Department 25 Display Control Unit 26 Search Section 30 Derivation Model 31. Embedding layer 32 RNN layers 33. Weighting Calculation Mechanism 34 MLP 36 Context vectors 40 words 41,42 Feature vectors 44 Specific results 50, 70 creation screen 51, 71 Image display area 52, 72 Text display area 53 Specific result display area 60 yen 73 Results display area d1 radius G0 Target medical images

Claims

1. Equipped with at least one processor, The aforementioned processor, Divide the sentence into predetermined units, Determine at least one of the attributes, factual nature, and relationships of each of the aforementioned units, Based on the results of the above determination, the weight for each of the above units is determined, An information processing device that derives the feature quantities of each unit using a derivation model constructed by machine learning, and then derives the feature quantities of the sentence by performing a weighting calculation on the feature quantities of each unit based on the weights.

2. The information processing apparatus according to claim 1, wherein the processor determines the weights for words determined to have predetermined attributes, facts, and relationships to be greater than the weights for words determined to have attributes, facts, and relationships other than the predetermined attributes, facts, and relationships.

3. The information processing apparatus according to claim 2, wherein the predetermined attributes, facts, and relationships are determined according to a task using the derived feature quantities of the sentence.

4. The information processing apparatus according to claim 2 or 3, wherein the processor determines the weights by the derivation model.

5. The information processing apparatus according to any one of claims 1 to 4, wherein the processor displays the sentence by highlighting words determined to have predetermined attributes, facts, and relationships.

6. The information processing apparatus according to any one of claims 1 to 5, wherein the processor performs the task of identifying the content of the sentence using the derived feature quantities.

7. The processor uses the derived feature quantities to search for images corresponding to the sentence. An information processing device according to any one of claims 1 to 5, which performs a task.

8. The information processing apparatus according to any one of claims 1 to 7, wherein the processor determines at least one of the attributes and facts of each unit.

9. The information processing apparatus according to claim 8, wherein the processor determines both the attributes and factual status of each unit.

10. The information processing apparatus according to claim 9, wherein the processor further determines the relationship between each of the units.

11. The computer divides the sentence into predetermined units, Determine at least one of the attributes, factual nature, and relationships of each of the aforementioned units, Based on the results of the above determination, the weight for each of the above units is determined, An information processing method for deriving the feature quantities of a sentence by using a derivation model constructed by machine learning to derive the feature quantities of each unit, and by performing a weighting calculation on the feature quantities of each unit based on the weights.

12. The procedure for dividing a sentence into predetermined units, A procedure for determining at least one of the attributes, factual nature, and relationships of each of the aforementioned units, A procedure for determining the weight for each unit according to the result of the above determination, An information processing program that causes a computer to perform the following steps: deriving the feature quantities of each unit using a derivation model constructed by machine learning, and deriving the feature quantities of the sentence by performing a weighting calculation on the feature quantities of each unit based on the weights.

Citation Information

Patent Citations

  • Text classification method, device and equipment

    CN111753525A

  • Multi-source social network construction method based on user identity association

    CN111815468A

  • Fire-fighting plan classification method based on deep learning

    CN112069814A

  • Natural language processing method and device and electronic equipment

    CN112528654A

  • Information processing device and information processing program

    JP2015135640A