Medical diagnostic support system, medical diagnostic support method, and program
Patent Information
- Application Number
- JP2022118380
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-07-26
- Publication Date
- 2026-09-14
- Estimated Expiration
- 2042-07-26
AI Technical Summary
【0007】 本発明によれば、医療従事者の音声情報を用いて、医用画像を精度よく特定することができる。
Smart Images

Figure 0007919939000001 
Figure 0007919939000002 
Figure 0007919939000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a medical diagnosis support system, a medical diagnosis support method, and a program for specifying medical images using voice information of medical staff. [Background Art]
[0002] In a conventional medical diagnosis support system, image character information corresponding to each of a plurality of medical images is compared with predetermined character information, and at least one medical image related to the predetermined character information is specified from the plurality of medical images. (For example, Patent Document 1) [Prior Art Literature] [Patent Literature]
[0003] [Patent Document 1] Japanese Unexamined Patent Application Publication No. 2019-169049 [Summary of the Invention] [Problem to be Solved by the Invention]
[0004] However, in Patent Document 1, there is a possibility that the image character information does not match the predetermined character information, which may make it difficult to specify a medical image. Accordingly, an object of the present invention is to accurately specify a medical image using voice information of a medical staff member by converting information corresponding to the medical image into other information. [Means for Solving the Problem]
[0005] To achieve the object of the present invention, the medical diagnosis support system of the present invention comprises: an image acquiring means for acquiring a plurality of medical images; and the plurality of medical images Each associated with , showing the characteristics of the lesion in each of the multiple medical images. a finding information acquiring means for acquiring finding information; the finding information These are simple terms. a converting means for converting the finding information into term information; and a medical staff member The speech that is directed at the subjectAn analysis means analyzes speech to obtain textual information, and a conversion means converts the term information and then, based on the textual information obtained by the analysis means, selects from among multiple medical images. at least one It comprises a means for identifying medical images.
[0006] Furthermore, the medical diagnostic support method of the present invention includes the steps of acquiring multiple medical images and the multiple medical images Each Linked to , showing the characteristics of the lesion in each of the multiple medical images. Steps to obtain findings information, and the findings information These are simple terms. Steps to convert into terminology information, and healthcare professionals This is done by the subject. The process involves analyzing speech to obtain textual information, and then, based on the terminology and textual information, selecting from multiple medical images. at least one The process includes the step of identifying a medical image. [Effects of the Invention]
[0007] According to the present invention, medical images can be accurately identified using voice information from medical professionals. [Brief explanation of the drawing]
[0008] [Figure 1] This figure shows the configuration of the medical diagnostic support system of the present invention. [Figure 2] A diagram showing a modified example of the medical diagnostic support system of the present invention. [Figure 3] This figure shows specific examples of the image identification unit 126 and the conversion unit 132 in the medical diagnostic support system of the present invention. [Figure 4] This figure shows one display mode of the display selection unit 128 in the medical diagnostic support system of the present invention. [Figure 5] This figure shows one display mode of the display selection unit 128 in the medical diagnostic support system of the present invention. [Figure 6] A flowchart illustrating the operation of the medical diagnostic support system of the present invention. [Modes for carrying out the invention]
[0009] The embodiments for carrying out the present invention will be described below with reference to the drawings.
[0010] Embodiments of the present invention will be described with reference to Figure 1. Figure 1 is a diagram showing the configuration of a medical diagnostic support system according to an embodiment of the present invention.
[0011] The medical diagnostic support system includes a microphone (sound acquisition unit) 110 that acquires sound, and a storage unit 112 that stores the sound data acquired by the microphone (sound acquisition unit) 110.
[0012] The medical diagnostic support system includes a control unit 120 that analyzes the voice data of 100 medical professionals (doctors, nurses, etc.) and identifies medical images from the analyzed voice data of the medical professionals 100 and terminology information based on findings corresponding to multiple medical images captured by the imaging device 130, and a display unit 140 that displays the medical images. The display unit 140 is, for example, an LCD monitor or a CRT monitor. The components (functions) of the medical diagnostic support system and the control unit 120 are realized, for example, by a processor such as a CPU (Central Processing Unit) or GPU (Graphics Processing Unit) executing a program (software) stored in memory.
[0013] The microphone (sound acquisition unit) 110 acquires the voice of the medical professional 100 and the voice of the subject (patient) 102. The microphone (sound acquisition unit) 110 performs AD conversion on the voice and generates audio data.
[0014] When one microphone (sound acquisition unit) 110 is placed in the imaging room, the microphone (sound acquisition unit) 110 will acquire both the voice of the medical professional 100 and the voice of the subject 102. Therefore, in a later stage, it is necessary to separate the audio data into the voice data of the medical professional 100 and the voice data of the subject 102. In this embodiment, a configuration in which one microphone (sound acquisition unit) 110 is placed in the imaging room is described, but the embodiment is not limited to this configuration.
[0015] Note that when two microphones (audio acquisition units) 114 and 116 are placed in an imaging room, there is no need to separate the audio data into the voice of the medical worker 100 and the voice of the subject 102.
[0016] The storage unit 112 stores the audio data acquired by the microphone (audio acquisition unit) 110 while the medical worker 100 and the subject 102 are having a conversation. Specifically, when the volume of the audio data acquired by the microphone (audio acquisition unit) 110 is equal to or higher than a predetermined threshold, the storage unit 112 stores the audio data acquired by the microphone (audio acquisition unit) 110, triggered by the timing when the volume of the audio data becomes equal to or higher than the predetermined threshold. That is, the storage unit 112 can store the audio data acquired by the microphone (audio acquisition unit) 110 while the medical worker 100 and the subject 102 are conversing at a predetermined volume.
[0017] If the state where the volume of the audio data acquired by the microphone (audio acquisition unit) 110 is less than a predetermined threshold continues for a predetermined time (for example, 1 minute), the storage unit 112 determines that the medical worker 100 and the subject 102 are not having a conversation, and does not store the audio data acquired by the microphone (audio acquisition unit) 110.
[0018] The control unit 120 includes: a voice separation unit 122 that separates the audio data of the medical worker 100 from the audio data acquired by the microphone (audio acquisition unit) 110; an analysis unit 124 that analyzes the audio data of the medical worker 100; an image identification unit 126 that identifies a medical image from the audio data of the medical worker 100 and term information obtained by converting finding information corresponding to a plurality of medical images; and a display selection unit 128 that selects a medical image to be displayed from among the identified medical images.
[0019] The control unit 120 acquires audio data from the storage unit 112. The control unit 120 can also acquire audio data from the storage unit 112 in real time in accordance with the speech of the medical professional 100. Once audio data is acquired from the storage unit 112, the audio separation unit 122 separates the audio data from the storage unit 112 into the audio data of the medical professional 100 and the audio data of the subject 102. The audio separation unit 122 then extracts the audio data of the medical professional 100.
[0020] In a hospital examination room, the medical professional 100 is generally fixed, while the patient 102 changes with each examination. Therefore, if the voice separation unit 122 learns the characteristics of the medical professional 100's voice data in advance, it can distinguish between the voice data of the medical professional 100 and the voice data of the patient 102, which is different.
[0021] Specifically, the voice separation unit 122 pre-acquires voice data from the medical professional 100, and pre-trains the voice separation unit 122 with the feature quantities of the medical professional 100's voice data. The feature quantities of voice data are at least one of the following: volume, frequency, tone, intonation, and speaking speed. Volume indicates the loudness of the voice. Frequency and tone indicate the pitch of the voice. Intonation is the intonation of the voice, and speaking speed is the length of speech per unit of time.
[0022] The voice separation unit 122 compares the pre-trained feature quantities of the voice data of the medical professional 100 with the feature quantities of the voice data input to the voice separation unit 122, and separates the voice data of the medical professional 100 from the voice data of the subject 102 based on the likelihood that the voice is likely to be that of the medical professional (the person themselves).
[0023] In this way, the voice separation unit 122 can separate the voice data output from the storage unit 112 into the voice data of the medical professional 100 and the voice data of the subject 102. In other words, the voice separation unit 122 can estimate the portion of the voice data belonging to the medical professional 100 from the voice data containing both the voices of the medical professional 100 and the subject 102. To put it another way, the voice separation unit 122 can estimate the voice data of the medical professional 100 from the voice data containing both the voices of the medical professional 100 and the subject 102.
[0024] Furthermore, the voice separation unit 122 can extract the voice data acquired from the storage unit 112 as text information through voice recognition processing, and separate the voice data of the medical professional 100 and the voice data of the subject 102 from the text information of the voice data.
[0025] The voice separation unit 122 pre-acquires the voice data of the medical professional 100 as text information, and learns the content of the text information based on the voice data of the medical professional 100. The content of the text information based on the voice data is the content of the medical professional 100's statements regarding the medical interview and the content of the subject 102's statements. Examples of the medical professional 100's statements regarding the medical interview include, "What brings you here today?", "Do you have any pain?", "Have you been able to eat?", "There is a blockage in your blood vessels.", "There is swelling.", "You will be admitted to the hospital." Examples of the subject 102's statements include, "I have a headache.", "I feel fine.", "I ate breakfast this morning."
[0026] The voice separation unit 122 can estimate the portion of the audio data acquired from the storage unit 112 that is the voice of the medical professional 100 if the content of the medical professional's speech regarding the medical interview is recognized. The voice separation unit 122 can estimate the portion of the audio data acquired from the storage unit 112 that is the voice of the subject 102 if the content of the subject's speech is recognized.
[0027] The voice data of the medical professional 100 and the voice data of the subject 102, separated by the voice separation unit 122, are output to the analysis unit 124.
[0028] The analysis unit 124 analyzes the voice data of the medical professional 100 that was separated by the voice separation unit 122. The analysis unit 124 analyzes the voice data of the medical professional 100 and extracts specific textual information (keywords) related to the lesion. The analysis unit 124 can also analyze the voice data of the medical professional 100 and extract specific textual information (keywords) that are highly likely to be related to the findings. For example, from the voice data "There is a blockage in the blood vessel," "blockage" is extracted. From the voice data "There is swelling," "swelling" is extracted. From the voice data "A tumor can be seen in the medical image. It is spherical in shape here," "tumor" and "spherical" are extracted.
[0029] The analysis unit 124 analyzes the voice data of the medical professional 100 and outputs the acquired specific character information (keywords) to the image identification unit 126.
[0030] Furthermore, the medical diagnostic support system includes an imaging device 130 that photographs the subject 102 and acquires medical images, a conversion unit 132 that converts the findings information associated with the medical images into terminology information, and an operation unit 134 that performs operations on the control unit 120. The imaging device 130 may also have a conversion unit 132.
[0031] The imaging device 130 is a device that acquires medical image data of a patient, such as an X-ray CT (Computed Tomography) device, an MRI (Magnetic Resonance Imaging) device, or an ultrasound diagnostic device.
[0032] An X-ray CT scanner is equipped with an X-ray source and an X-ray detector. CT image data is generated by rotating the X-ray source and detector around the patient, irradiating the patient with X-rays from the X-ray source, and projecting the data detected by the X-ray detector.
[0033] An MRI device generates a predetermined magnetic field for a subject placed in a static magnetic field, and then performs a Fourier transform on the acquired data to produce MRI image data.
[0034] An ultrasound diagnostic device transmits ultrasound waves to a patient and receives the reflected ultrasound waves from the patient to generate ultrasound image data.
[0035] The medical images (CT image data, MRI image data, ultrasound image data, etc.) generated by the imaging device 130 are three-dimensional data (volume data) and two-dimensional data.
[0036] In this embodiment, a magnetic resonance imaging apparatus, a CT scanner, an ultrasound scanner, etc., are given as examples of imaging devices 130 and explained, but other imaging devices may be used as long as medical images can be acquired.
[0037] The operating unit 132 may be, for example, a mouse or keyboard. The operating unit 132 may also be a touch panel monitor that integrates a display panel for displaying medical images with an operating panel.
[0038] The imaging device 130 may have a function to analyze the captured medical image and extract a specific lesion area. One method for extracting a specific lesion area in a medical image is to use computer-aided diagnosis (CAD) technology. Alternatively, a region growing method may be used to extract a specific lesion area in a medical image. The region growing method is an image extraction method in which a reference point for the area to be extracted is set via the operation unit 134, and the area of pixels whose difference from the pixel value of the reference point falls within the set range is extracted. The operator sets a reference point in the area of interest in the medical image using the operation unit 132 and sets the set range. The region growing method extracts the area of pixels whose difference from the pixel value of the reference point falls within the set range, and the area of pixels that falls within the set range is extracted.
[0039] The imaging device 130 can associate findings information with a medical image if a specific lesion area is present in the medical image. Finding information is information that describes the characteristics of the lesion in the medical image and is entered by a medical professional 100. Finding information is shown in the form of, for example, "thrombotic infarction in the left middle cerebral artery" or "mild hernia signs in the upper right brain." It may also be shown based on the morphology of the lesion in the medical image, such as "spherical," "lobulated," "polygonal," "wedge-shaped," "flat," or "irregular." Furthermore, by associating finding information with a specific lesion area, it is also possible to link finding information to a specific lesion area in a medical image. For example, "thrombotic infarction in the left middle cerebral artery" can be associated with an area suspected of having a thrombotic infarction in a medical image. Similarly, "mild hernia signs in the upper right brain" can be associated with an area suspected of having a hernia in a medical image.
[0040] Furthermore, findings information can also be obtained using a neural network trained on training data. For example, the imaging device 130 stores a pre-trained model that has been trained to output (infer) findings information in medical images. The pre-trained model is generated using a neural network, for example, but among neural network technologies, deep learning technologies such as CNN and RNN (Recurrent Neural Network), or models derived from CNN and RNN, such as Vision Transformer (ViT), may also be used. In addition, other machine learning technologies such as support vector machines, logistic regression, and random forests may be used, or rule-based methods may be used.
[0041] Training data is pre-generated using multiple medical images and findings. The training data is determined according to the inference task and classification target that the neural network will perform. When training a neural network to perform a classification task, training data is generated by pairing medical image data with ground truth labels, which are findings that indicate the characteristics of lesions captured in the medical images.
[0042] The conversion unit 132 pre-stores a database of terms for converting findings information associated with medical images into simpler terms. Simple conversion is a conversion that generalizes (simplifies) the findings information.
[0043] When the conversion unit 132 acquires finding information associated with a medical image, it outputs term information that matches a word pre-stored in the database for each word that makes up the finding information. The database in the conversion unit 132 stores multiple words and term information as a set. The database in the conversion unit 132 may also store a table in which words and term information are set one-to-one. For example, the term information for "infarction" is stored in correspondence with the term "blockage" and "blockage". Similarly, the term information for "hernia" is stored in correspondence with the term "swelling".
[0044] Here, if the findings information of medical image 1 contains the word "infarction," the conversion unit 132 can output term information for "blockage" or "blockage." Similarly, if the findings information of medical image 2 contains the word "hernia," the conversion unit 132 can output term information for "swelling." In this way, the conversion unit 132 can output term information that is a simplified version of various words in the findings information.
[0045] The conversion unit 132 can also output the word in the findings information and the converted term information as a set. Specifically, the conversion unit 132 outputs the word in the findings information before conversion and the converted term information. For example, if the findings information contains the word "infarction," then "infarction," "blockage," and "blockage" will be output. If the findings information contains the word "hernia," then "hernia" and "swelling" will be output.
[0046] In this way, the conversion unit 132 can output various words in the findings information and terminology information obtained by converting those various words in the findings information into simpler terms to the image identification unit 126.
[0047] The image identification unit 126 identifies the medical image to be displayed on the display unit 140 from specific character information (keywords) obtained by analyzing the voice data of the medical professional 100 and term information converted from finding information based on multiple medical images. The image identification unit 126 identifies the medical image to which term information corresponding to specific character information (keywords) related to the lesion is linked.
[0048] The display unit 140 displays the medical image identified by the image identification unit 126. The display unit 140 may also display the medical image together with findings information or terminology information. In addition, the display unit 140 may also display the medical image together with specific textual information (keywords) related to the lesion.
[0049] Therefore, the medical diagnostic support system can display a medical image corresponding to a predetermined keyword on the display unit 140 when a predetermined keyword appears in the speech that the medical professional 100 is giving to the patient 102. In other words, the medical diagnostic support system stores a link between medical images and terminology information, which is a simplified version of the findings information corresponding to the medical image. Then, when the stored terminology information appears in the explanatory speech that the medical professional 100 is giving to the patient 102, the system can recommend a medical image corresponding to the terminology information.
[0050] Figure 2 shows a modified example of the medical diagnostic support system according to an embodiment of the present invention.
[0051] The medical diagnostic support system includes a medical image server 200 that stores multiple medical images taken with various imaging devices. The medical image server 200 is connected to a conversion unit 132 and a control unit 120.
[0052] The conversion unit 132 includes an image acquisition unit 150 that acquires medical images from the medical image server 200, a findings information acquisition unit 152 that acquires findings information associated with the medical images, and a findings information conversion unit 154 that converts various words in the findings information into terminology information.
[0053] The image acquisition unit 150 acquires medical images corresponding to the subject 102 from the medical image server 200. The image acquisition unit 150 can also analyze the voice data of the subject 102 separated by the voice separation unit 122, identify the subject 102's ID from the analyzed voice data, and acquire medical images corresponding to the subject 102's ID.
[0054] Medical images acquired by the image acquisition unit 150 are associated with findings information. The findings information acquisition unit 152 acquires the findings information associated with the medical images. The findings information acquisition unit 152 can acquire various words that make up the findings information.
[0055] The findings information conversion unit 154 converts findings information into terminology information. For each word that makes up the findings information, it outputs terminology information that matches a word pre-stored in the findings information conversion unit 154. The medical image server 200 then stores the medical image and the terminology information converted from the findings information in association.
[0056] The other components of the medical diagnostic support system are the same as those in Figure 1, so their explanation will be omitted.
[0057] Figure 3 shows a specific example of the image identification unit 126 and the conversion unit 132 according to an embodiment of the present invention.
[0058] Here, it is assumed that medical image 1 and medical image 2 of subject 102 have been taken by imaging device 130. Imaging device 130 or image server 200 stores medical image 1 and medical image 2. Finding information is associated with medical image 1 and medical image 2, respectively. The finding information for medical image 1 is "Thrombotic infarction in the left middle cerebral artery." The finding information for medical image 2 is "Mild signs of herniation in the upper right brain."
[0059] The conversion unit 132 extracts multiple words from the finding information of medical image 1, "Thrombotic infarction in the left middle cerebral artery." These words are "left middle cerebral artery," "thrombotic," and "infarction." The conversion unit 132 outputs term information that matches pre-registered words. The pre-registered words include the word "infarction." The conversion unit 132 also stores corresponding term information for "blockage" and "blockage" for the word "infarction."
[0060] The conversion unit 132 converts the word "infarction" into the term information for "blockage" or "blockage".
[0061] Furthermore, the conversion unit 132 extracts multiple words from the findings information of medical image 2, "Signs of mild hernia in the upper right brain." These words are "upper right brain," "mild," "hernia," and "signs." The conversion unit 132 outputs term information that matches pre-registered words. The pre-registered words include the word "hernia." The conversion unit 132 has term information for "swelling" associated with the word "hernia." The conversion unit 132 converts the word "hernia" into term information for "swelling."
[0062] The analysis unit 124 analyzes the voice data of the medical professional 100 that was separated by the voice separation unit 122. The analysis unit 124 analyzes the voice data of the medical professional 100 and extracts textual information. For example, from the voice data, "The examination revealed a small blockage in the blood vessels of the brain. You will be hospitalized for a while," the textual information "blockage," which is related to the lesion, is extracted. Note that "blockage" is specific textual information that is highly likely to be related to the findings.
[0063] The image identification unit 126 identifies the medical image to be displayed on the display unit 140 based on the text information "blockage" obtained by analyzing the voice data of the medical professional 100, and the term information "blockage," "blockage," and "swelling" converted from finding information based on multiple medical images.
[0064] The image identification unit 126 identifies medical image 1 and outputs it because the text information "blockage" obtained by analyzing the voice data of the medical professional 100 matches the term information "blockage" converted from the findings information based on medical image 1. The display unit displays medical image 1.
[0065] The display selection unit 128 selects the medical image to be displayed on the display unit 140 from among the multiple medical images identified by the image identification unit 126. However, if only one medical image is identified by the image identification unit 126, it is not necessary to select a medical image using the display selection unit 128.
[0066] Specifically, the image identification unit 126 identifies multiple medical images to be displayed on the display unit 140 based on the text information obtained by analyzing the voice data of the medical professional 100 and the terminology information converted from the findings information based on multiple medical images. The display selection unit 128 then selects the medical image to be displayed on the display unit 140. The display selection unit 128 may use the display unit 140 as a display interface for image selection, or it may maintain a different display interface. Furthermore, the image identification unit 126 can identify a medical image even if the text information obtained by analyzing the voice data of the medical professional 100 does not match the terminology information converted from the findings information based on the medical images. Specifically, the image identification unit 126 identifies a medical image from among multiple medical images based on the similarity between the text information obtained by analyzing the voice data of the medical professional 100 and the terminology information converted from the findings information based on the medical images. The similarity between the text information and the terminology information is pre-stored in a database.
[0067] For example, the image identification unit 126 identifies medical image A because the text information "blockage" obtained by analyzing the voice data of the medical professional 100 is similar to the term information "thrombus" converted from the findings information based on medical image A (high similarity). The display unit 140 displays medical image A. On the other hand, the image identification unit 126 does not identify medical image B because the text information "blockage" obtained by analyzing the voice data of the medical professional 100 is not similar to the term information "swelling" converted from the findings information based on medical image B (low similarity).
[0068] Furthermore, the image identification unit 126 can also use a neural network to infer and identify medical images to be displayed on the display unit 140. The image identification unit 126 has a trained model stored in it that learns using terminology information obtained by analyzing the voice data of the medical professional 100 and converting it with findings information based on medical images. When text information obtained by analyzing the voice data of the medical professional 100 is input, the trained model outputs (infers) a medical image corresponding to the voice data of the medical professional 100. The image identification unit 126 can use the trained model to identify medical images to be displayed on the display unit 140.
[0069] Figure 4 shows one display mode of the display selection unit 128 of the present invention.
[0070] In this case, the image identification unit 126 identifies medical image 3 and medical image 4 from the text information obtained by analyzing the voice data of the medical professional 100 and the term information converted from the findings information based on multiple medical images. If medical image 3 and medical image 4 are identified, they are displayed on the display interface of the display selection unit 128.
[0071] Medical images 3 and 4 are assigned an image ID, terminology (keywords) converted from the findings information associated with medical images 3 and 4, and the name of the imaging device (modality) used to capture medical images 3 and 4. Therefore, when multiple medical images are presented as candidates, the display selection unit 128 can simultaneously display which terminology (keywords) each medical image is associated with. Thus, the medical professional 100 can check the terminology (keywords) converted from the findings information associated with medical images 3 and 4 and select a medical image using the display selection unit 128. The medical professional 100 can select a medical image using the selection mark (arrow) 160. The display unit 140 displays the medical image selected by the display selection unit 128.
[0072] Furthermore, when displaying medical image 3 and medical image 4, the display selection unit 128 may display partial images 162 and 164, which are excised areas of interest such as lesion regions, based on terminology information (keywords) converted from the findings information associated with medical image 3 and medical image 4. Specifically, the image identification unit 126 identifies the lesion region corresponding to "blockage" in the terminology information (keywords) associated with medical image 3. The image identification unit 126 generates partial image 162 by excising the image to include the lesion region corresponding to "blockage". Partial image 162 is an enlarged image of the area of interest that includes the lesion region corresponding to "blockage". Partial image 162 is positioned adjacent to medical image 3.
[0073] The image identification unit 126 identifies the lesion area corresponding to "swelling" in the terminology information (keywords) associated with the medical image 4. The image identification unit 126 generates a partial image 164 by cropping out the area including the lesion area corresponding to "swelling". The partial image 164 is an enlarged image of the area of interest that includes the lesion area corresponding to "swelling". The partial image 164 is positioned adjacent to the medical image 4.
[0074] Furthermore, the display selection unit 128 may also display non-image information, such as findings information linked to medical images 3 and 4, and electronic medical record data, as supplementary information associated with medical images 3 and 4 identified by the image identification unit 126, including information about the imaging device used to photograph the subject 102. Healthcare professionals can then review the information related to medical images 3 and 4 and determine which medical image to select.
[0075] Figure 5 shows one display mode of the display selection unit 128 of the present invention.
[0076] Here, the image identification unit 126 identifies medical image 5 and medical image 6 from text information obtained by analyzing the voice data of the medical professional 100 and terminology information converted from findings information based on multiple medical images. When medical image 5 and medical image 6 are identified, they are displayed on the display unit 140. Medical image 5 was taken with an MRI device, and medical image 6 was taken with an X-ray CT device. The display selection unit 128 can also simultaneously display medical images from other imaging devices if the same location has been captured with multiple imaging devices. In other words, the display selection unit 128 can simultaneously display medical images taken with multiple imaging devices.
[0077] Specifically, the image identification unit 126 identifies the lesion area corresponding to "blockage" in the terminology information (keywords) associated with the medical image 5. The image identification unit 126 generates a partial image 166 by cutting out the lesion area corresponding to "blockage" to include it. The partial image 166 is an enlarged image of the area of interest that includes the lesion area corresponding to "blockage". The partial image 166 is positioned adjacent to the medical image 5.
[0078] The image identification unit 126 identifies the area in medical image 6 that corresponds to the lesion area in the term information (keyword) associated with medical image 5, which is at the same location. The image identification unit 126 generates a partial image 168 by cutting out the medical image 6 to include the identified area. The partial image 166 is placed adjacent to medical image 6.
[0079] Medical professionals 100 can review medical images taken with different imaging devices and determine which medical image, or both, should be selected.
[0080] Figure 6 is a flowchart showing the operation of a medical diagnostic support system according to an embodiment of the present invention.
[0081] S100: The imaging device 130 photographs the subject 102 and acquires multiple medical images. The imaging device 130 assigns the subject 102's identification information (ID) to each medical image.
[0082] S102: If a specific lesion area is present in the medical image captured by the imaging device 130, the medical professional 100 inputs findings information regarding the specific lesion in the medical image. The findings information regarding the specific lesion is linked to the medical image. The conversion unit 132 (finding information acquisition unit 152) acquires the findings information corresponding to the medical image.
[0083] S104: Here, the medical professional 100 decides whether or not to display a medical image using voice information via the operation unit 134. This decision information is transmitted to the control unit 120. If the medical image is not to be displayed using voice information, the process proceeds to S106; if the medical image is to be displayed using voice information, the process proceeds to S108.
[0084] S106: If audio information is not used to display medical images, the control unit 120 outputs multiple medical images captured by the imaging device 130 and the findings information corresponding to the multiple medical images to the display unit 140. The display unit 140 displays the multiple medical images and the findings information corresponding to the multiple medical images.
[0085] S108: When displaying medical images using audio information, first, the findings information is converted into terminology information. The conversion unit 132 outputs terminology information obtained by converting various words in the findings information corresponding to multiple medical images into simpler terms. The terminology information converted by the conversion unit 132 is linked to the medical image corresponding to the findings information before conversion. In other words, by identifying the terminology information, the medical image can be identified.
[0086] S110: The analysis unit 124 analyzes the voice data of the medical professional 100 that was separated by the voice separation unit 122. The analysis unit 124 analyzes the voice data of the medical professional 100 and obtains specific textual information (keywords).
[0087] S112: The image identification unit 126 identifies the medical image to be displayed on the display unit 140 based on specific character information (keywords) obtained by analyzing the voice data of the medical professional 100 and term information converted from finding information corresponding to multiple medical images.
[0088] S114: The display unit 140 displays the medical image identified and output by the image identification unit 126. If there are multiple medical images identified and output by the image identification unit 126, the display selection unit 128 displays multiple medical images. When multiple medical images are identified by the image identification unit 126, the display selection unit 128 selects from the displayed medical images to display on the display unit 140.
[0089] As described above, the medical diagnostic support system of the present invention comprises: an image acquisition means for acquiring multiple medical images; a findings information acquisition means for acquiring findings information associated with the multiple medical images; a conversion means for converting the findings information into terminology information; an analysis means for analyzing the voice information of medical professionals to acquire text information; and an identification means for identifying a medical image from among the multiple medical images based on the terminology information converted by the conversion means and the text information acquired by the analysis means. Therefore, medical images can be identified with high accuracy using the voice information of medical professionals.
[0090] The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions. This program and the computer-readable storage medium storing the program are included in the present invention.
[0091] The embodiments of the present invention described above are merely examples of how the invention can be implemented, and the technical scope of the invention should not be interpreted as being limited by them. In other words, the present invention can be implemented in various forms without departing from its technical concept or its main features. [Explanation of Symbols]
[0092] 100 healthcare workers 102 Subjects 110 Microphone (Voice Acquisition Unit) 112 Storage section 120 Control Unit 122 Audio separation unit 124 Analysis Department 126 Image Identification Section 130 Imaging device 132 Conversion section 134 Operation section 140 Display section
Claims
1. Image acquisition means for acquiring multiple medical images, A means for acquiring findings information that is linked to each of multiple medical images and acquires findings information indicating the characteristics of lesions in each of the multiple medical images, A conversion means for converting the aforementioned findings information into terminology information in simple terms, An analysis means for analyzing speech data of utterances made by medical professionals to test subjects and obtaining textual information, A medical diagnostic support system comprising: an identification means for identifying at least one medical image from among the plurality of medical images based on the term information converted by the conversion means and the character information acquired by the analysis means.
2. The medical diagnostic support system according to claim 1, characterized in that the analysis means analyzes the voice data of the medical professional and extracts specific character information related to the lesion.
3. The medical diagnostic support system according to claim 1, characterized in that the conversion means pre-stores terminology information in a database for converting the findings information into simpler terms.
4. The medical diagnostic support system according to claim 3, characterized in that the conversion means outputs term information that matches a word pre-stored in the database for each word constituting the findings information.
5. The medical diagnostic support system according to claim 1, characterized in that the conversion means outputs a set of words constituting the findings information and the converted term information.
6. The medical diagnostic support system according to claim 1, characterized in that the identification means identifies a medical image to which term information corresponding to specific character information related to a lesion is linked.
7. The medical diagnostic support system according to claim 1, further comprising a display means for displaying the medical image identified by the identification means together with the findings information, the terminology information, or specific character information related to the lesion.
8. The medical diagnostic support system according to claim 1, characterized in that, when multiple medical images are identified by the identification means, it includes a selection means for selecting a medical image to be displayed on the display means from among the multiple medical images.
9. The medical diagnostic support system according to claim 8, characterized in that when a plurality of medical images are identified by the identification means, the display means displays term information associated with each medical image.
10. The medical diagnostic support system according to claim 7, characterized in that the display means displays a partial image generated by cutting out a lesion region corresponding to the term information in a medical image identified by the identification means.
11. The medical diagnostic support system according to claim 7, characterized in that the display means displays information of the imaging device used to photograph the subject, added to the findings information or electronic medical record data associated with the medical image.
12. The plurality of medical images include medical images taken by each of the plurality of imaging devices, The medical diagnostic support system according to claim 7, characterized in that when the identification means identifies a plurality of medical images taken of the same part of the subject by each of the plurality of imaging devices using the term information converted by the conversion means and the character information acquired by the analysis means, the display means displays the plurality of medical images taken by the plurality of imaging devices together.
13. The medical diagnostic support system according to claim 1, characterized in that the identifying means identifies a medical image from among the plurality of medical images based on the similarity between the text information obtained by analyzing the voice data of the medical professional and the term information converted from the findings information linked to the medical image.
14. The medical diagnostic support system according to claim 1, further comprising a voice separation unit for separating the voice data of the medical professional from the voice data of the subject, wherein the analysis means analyzes the separated voice data of the medical professional.
15. The medical diagnostic support system according to claim 14, characterized in that the voice separation unit compares the feature quantities of the voice data of a medical professional that have been learned in advance with the feature quantities of the voice data acquired by the voice acquisition means, and separates the voice data of the medical professional from the voice data of the subject based on the likelihood that the voice sounds like that of a medical professional.
16. The medical diagnostic support system according to claim 1, characterized in that the analysis means analyzes the voice data acquired in real time in accordance with the speech made by the medical professional to the subject.
17. The medical diagnostic support system according to claim 1, characterized in that the means for acquiring the findings information acquires the findings information using a trained model that has been trained using a plurality of medical images and findings information contained in each of the plurality of medical images.
18. The medical diagnostic support system according to claim 1, characterized in that the identifying means identifies at least one medical image corresponding to the text information acquired by the analysis means, using a trained model that has been trained using text information acquired by analyzing voice data of medical professionals and terminology information converted from findings information linked to medical images.
19. Steps to acquire multiple medical images, The steps include: obtaining findings information that is linked to each of multiple medical images and indicates the characteristics of the lesion in each of the multiple medical images; The steps include converting the aforementioned findings information into terminology information using simpler terms, The process involves analyzing the speech spoken by medical professionals to the subject to obtain textual information, and A medical diagnostic support method characterized by comprising the step of identifying at least one medical image from among the plurality of medical images based on the term information and the character information.
20. A program for causing a computer to execute the medical diagnostic support method described in claim 19.
Citation Information
Patent Citations
Information providing program, information providing device and information providing method
JP2012194928A
Information processor
JP2017054267A
Medical image specification device, method, and program
JP2019169049A
Medical diagnosis assistance system, medical diagnosis assistance device, medical diagnosis assistance method, and program
JP2020181422A
Systems and methods for using ai to identify regions of interest in medical images
US20220164586A1