Swallowing evaluation system and swallowing evaluation method

The system accurately evaluates swallowing by combining ultrasonic images and swallowing sounds through machine learning, addressing the challenge of timing variations and improving detection of swallowing abnormalities.

JP7848134B2Active Publication Date: 2026-04-20FUJIFILM CORP
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
FUJIFILM CORP
Filing Date
2021-11-25
Publication Date
2026-04-20

AI Technical Summary

Technical Problem

Existing swallowing evaluation systems face challenges in accurately assessing swallowing ability due to variations in swallowing sound timing among subjects, making it difficult to distinguish between the piriform fossa and food residues in ultrasonic images.

Method used

A swallowing evaluation system utilizing an ultrasonic probe, sound acquisition unit, and evaluation unit that employs machine learning to combine ultrasonic images and swallowing sounds, analyzing image and sound features through multiple neural networks to evaluate swallowing accuracy.

Benefits of technology

Enables highly accurate evaluation of swallowing by integrating ultrasonic images and swallowing sounds, improving the detection of swallowing abnormalities and residue presence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007848134000001
    Figure 0007848134000001
  • Figure 0007848134000002
    Figure 0007848134000002
  • Figure 0007848134000003
    Figure 0007848134000003
Patent Text Reader

Abstract

This swallowing evaluation system (1) comprises an ultrasonic probe (2), an image acquisition unit that acquires an ultrasonic image within the pharynx of a subject by transmitting and receiving ultrasonic beams using the ultrasonic probe (2), a sound acquisition unit that acquires the sound of swallowing by the subject, and an evaluation unit (19) that evaluates the swallowing of the subject by utilizing machine learning that combines the ultrasonic image acquired by the image acquisition unit and the swallowing sound acquired by the sound acquisition unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a swallowing evaluation system and a swallowing evaluation method for evaluating the swallowing of a subject using ultrasonic images.

Background Art

[0002] Conventionally, an examination has been performed to determine whether a subject can swallow normally based on an ultrasonic image of the pharyngeal region of the subject. Usually, when swallowing is not performed normally, food often remains in the vallecula epiglottica or the piriform fossa of the subject. It is known that it is difficult to distinguish between the piriform fossa and food residues and the surrounding tissues in ultrasonic images, and an examiner such as a doctor requires a certain degree of skill to confirm the piriform fossa and food residues in the ultrasonic image. In order to easily perform such an examination, for example, a swallowing ability measurement system as disclosed in Patent Document 1 has been developed.

[0003] The swallowing ability measurement system of Patent Document 1 acquires an ultrasonic image of the inside of the neck and the swallowing sound of the subject, synchronizes the acquisition timing of the ultrasonic image and the temporal change in the frequency of the swallowing sound with each other, and acquires at least one of the moving speeds of the wall of the tubular organ in the pharyngeal region of the subject and the food passing through the tubular organ, thereby evaluating the swallowing ability of the subject.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, because the timing of swallowing sounds varies among subjects, simply synchronizing the timing of ultrasound image acquisition with the time change of swallowing sound frequency, as in the technology disclosed in Patent Document 1, sometimes made it difficult to accurately evaluate the subject's swallowing.

[0006] The present invention aims to provide a swallowing evaluation system and a swallowing evaluation method that can evaluate a subject's swallowing with high accuracy. [Means for solving the problem]

[0007] The swallowing evaluation system according to the present invention is characterized by comprising an ultrasonic probe, an image acquisition unit that acquires an ultrasonic image of the pharynx of a subject by transmitting and receiving an ultrasonic beam using the ultrasonic probe, a sound acquisition unit that acquires swallowing sounds by the subject, and an evaluation unit that evaluates the swallowing of a subject by using machine learning that combines the ultrasonic image acquired by the image acquisition unit and the swallowing sounds acquired by the sound acquisition unit.

[0008] The evaluation unit preferably performs swallowing evaluation by inputting ultrasound images and swallowing sounds into a neural network. In this process, the evaluation unit can evaluate swallowing by inputting ultrasound images into a first neural network to calculate image features, inputting swallowing sounds into a second neural network to calculate sound features, and then inputting the image features and sound features into a third neural network.

[0009] Furthermore, the evaluation unit can also evaluate swallowing by inputting swallowing sounds into a second neural network to calculate sound features, and if the sound features exceed a predetermined feature threshold, inputting the ultrasound image and sound features into a fourth neural network. Furthermore, the evaluation unit can also evaluate swallowing by inputting ultrasound images into a first neural network to calculate image features, and then inputting swallowing sounds and image features into a fifth neural network if the image features exceed a predetermined feature threshold. Furthermore, the evaluation unit can also evaluate swallowing by inputting both ultrasound images and swallowing sounds into the same neural network.

[0010] Furthermore, the evaluation unit can input time-series data of ultrasound images and swallowing sounds into a neural network. The sound acquisition unit can acquire swallowing sounds during swallowing, and the image acquisition unit can acquire ultrasound images after swallowing. Furthermore, the image acquisition unit can also acquire ultrasound images during swallowing. Furthermore, the sound acquisition unit can acquire at least one of the breath sounds, either the breath sound before swallowing or the breath sound after swallowing.

[0011] The sound acquisition unit may have a microphone built into the ultrasonic probe. Furthermore, the sound acquisition unit may also have a microphone that is independent of the ultrasound probe and is in contact with the pharyngeal region of the subject.

[0012] Furthermore, the evaluation unit can output the presence or absence of swallowing residue in the subject's pharynx as an evaluation result. Furthermore, the evaluation unit can also output the presence or absence of swallowing difficulties in the subject as an evaluation result. Furthermore, the evaluation unit can also output the appropriate consistency of dysphagia-friendly food for the subject as an evaluation result.

[0013] Furthermore, it is preferable that the swallowing evaluation system includes a monitor that displays information representing the ultrasound image acquired by the image acquisition unit and the swallowing sound acquired by the sound acquisition unit.

[0014] The swallowing evaluation method according to the present invention is characterized by acquiring an ultrasonic image in the pharynx of a subject by transmitting and receiving an ultrasonic beam using an ultrasonic probe, acquiring a swallowing sound made by the subject, and evaluating the swallowing of the subject by using machine learning with the acquired ultrasonic image and the swallowing sound as inputs.

Effects of the Invention

[0015] According to the present invention, since the swallowing evaluation system includes an ultrasonic probe, an image acquisition unit that acquires an ultrasonic image in the pharynx of a subject by transmitting and receiving an ultrasonic beam using the ultrasonic probe, a sound acquisition unit that acquires a swallowing sound made by the subject, and an evaluation unit that evaluates the swallowing of the subject by using machine learning that combines the ultrasonic image acquired by the image acquisition unit and the swallowing sound acquired by the sound acquisition unit, it is possible to accurately evaluate the swallowing of the subject.

Brief Description of the Drawings

[0016] [Figure 1] It is a block diagram showing the configuration of a swallowing evaluation system according to Embodiment 1 of the present invention. [Figure 2] It is a block diagram showing the configuration of a transmission / reception circuit in Embodiment 1 of the present invention. [Figure 3] It is a block diagram showing the configuration of an image generation unit in Embodiment 1 of the present invention. [Figure 4] It is a block diagram showing the configuration of an evaluation unit in Embodiment 1 of the present invention. [Figure 5] It is a flowchart showing the operation of a swallowing evaluation system according to Embodiment 1 of the present invention. [Figure 6] It is a diagram showing an example of displaying information representing an ultrasonic image and a swallowing sound on a monitor. [Figure 7] It is a flowchart showing the detailed operation of swallowing evaluation in Embodiment 1 of the present invention. [Figure 8] It is a block diagram showing the configuration of an evaluation unit in Embodiment 3 of the present invention. [Figure 9]It is a flowchart showing the detailed operation of swallowing evaluation in Embodiment 3 of the present invention. [Figure 10] It is a block diagram showing the configuration of the evaluation unit in Embodiment 4 of the present invention. [Figure 11] It is a flowchart showing the detailed operation of swallowing evaluation in Embodiment 4 of the present invention.

Embodiments for Carrying Out the Invention

[0017] Hereinafter, embodiments of this invention will be described based on the accompanying drawings. The description of the constituent elements described below is made based on typical embodiments of the present invention, but the present invention is not limited to such embodiments. In this specification, a numerical range represented by "~" means a range including the numerical values described before and after "~" as the lower limit value and the upper limit value. In this specification, "identical" and "the same" include the error ranges generally acceptable in the technical field.

[0018] Embodiment 1 FIG. 1 shows the configuration of a swallowing evaluation system 1 according to Embodiment 1 of the present invention. The swallowing evaluation system 1 includes an ultrasonic probe 2, a device main body 3, and a microphone 4. The ultrasonic probe 2 and the device main body 3 are connected to each other, and the device main body 3 and the microphone 4 are connected to each other.

[0019] The ultrasonic probe 2 has a transducer array 11, and a transmission / reception circuit 12 is connected to the transducer array 11.

[0020] The main body of the device 3 has an image generation unit 13, which is connected to the transmitting / receiving circuit 12 of the ultrasonic probe 2. The transmitting / receiving circuit 12 and the image generation unit 13 constitute the image acquisition unit. A display control unit 14 and a monitor 15 are sequentially connected to the image generation unit 13. An image memory 16 is also connected to the image generation unit 13. The main body of the device 3 also has a sound processing unit 17, which is connected to a microphone 4. The microphone 4 and the sound processing unit 17 constitute the sound acquisition unit. A sound memory 18 is also connected to the sound processing unit 17. The sound memory 18 is connected to the display control unit 14. An evaluation unit 19 is connected to the image memory 16 and the sound memory 18, and the display control unit 14 is connected to the evaluation unit 19.

[0021] Furthermore, the control unit 20 is connected to the transmitting / receiving circuit 12, image generation unit 13, display control unit 14, image memory 16, sound processing unit 17, sound memory 18, and evaluation unit 19. In addition, the input device 21 is connected to the control unit 20. Furthermore, the processor 22 is composed of an image generation unit 13, a display control unit 14, a sound processing unit 17, an evaluation unit 19, and a control unit 20.

[0022] The transducer array 11 has a plurality of transducers arranged in one or two dimensions. Each of these transducers transmits ultrasound according to a drive signal supplied from the transmitting / receiving circuit 12, and also receives ultrasound echoes from the subject and outputs a signal based on the ultrasound echoes. Each transducer is constructed by forming electrodes at both ends of a piezoelectric body made of, for example, a piezoelectric ceramic represented by PZT (Lead Zirconate Titanate), a polymer piezoelectric element represented by PVDF (Poly Vinylidene Di Fluoride), or a piezoelectric single crystal represented by PMN-PT (Lead Magnesium Niobate-Lead Titanate).

[0023] The transmitting / receiving circuit 12 transmits ultrasonic waves from the transducer array 11 and generates a sound line signal based on the received signal acquired by the transducer array 11, under the control of the control unit 20. As shown in Figure 2, the transmitting / receiving circuit 12 includes a pulser 31 connected to the transducer array 11, and an amplifier 32, an AD (Analog Digital) converter 33, and a beamformer 34 connected sequentially in series from the transducer array 11.

[0024] The pulser 31 includes, for example, multiple pulse generators and supplies drive signals to multiple transducers of the transducer array 11, adjusting the delay amount, so that the ultrasonic waves transmitted from the transducers form an ultrasonic beam, based on a transmission delay pattern selected according to a control signal from the control unit 20. In this way, when a pulsed or continuous wave voltage is applied to the electrodes of the transducers of the transducer array 11, the piezoelectric material expands and contracts, generating pulsed or continuous wave ultrasonic waves from each transducer, and an ultrasonic beam is formed from the combined wave of these ultrasonic waves.

[0025] The transmitted ultrasonic beam is reflected from a target, such as a part of the subject, and propagates toward the transducer array 11 of the ultrasonic probe 2. The ultrasonic echo propagating toward the transducer array 11 is received by each of the transducers that make up the transducer array 11. At this time, each transducer that makes up the transducer array 11 expands and contracts upon receiving the propagating ultrasonic echo, generating a received signal, which is an electrical signal, and outputs these received signals to the amplification unit 32.

[0026] The amplification unit 32 amplifies the signals input from each transducer constituting the transducer array 11 and transmits the amplified signals to the AD conversion unit 33. The AD conversion unit 33 converts the signals transmitted from the amplification unit 32 into digital received data and transmits this received data to the beamformer 34. The beamformer 34 performs so-called receive focus processing by adding each received data converted by the AD conversion unit 33 with a corresponding delay, according to the sound velocity or sound velocity distribution set based on the reception delay pattern selected according to the control signal from the control unit 20. Through this receive focus processing, each received data converted by the AD conversion unit 33 is phase-aligned and added together, and a sound ray signal with a focused ultrasonic echo is obtained. This sound ray signal is sent to the image generation unit 13.

[0027] As shown in Figure 3, the image generation unit 13 has a configuration in which a signal processing unit 35, a DSC (Digital Scan Converter) 36, and an image processing unit 37 are connected in series in sequence. The signal processing unit 35 applies distance-dependent attenuation correction to the sound line signal transmitted from the transmitting / receiving circuit 12 according to the depth of the ultrasonic reflection position, and then performs envelope detection processing to generate a B-mode image signal, which is tomographic image information about the tissue within the subject.

[0028] The DSC36 converts the B-mode image signal generated by the signal processing unit 35 into an image signal that follows the scanning method of a normal television signal (raster conversion). The image processing unit 37 performs various necessary image processing, such as gradation processing, on the B-mode image signal input from the DSC 36, and then sends the B-mode image signal to the display control unit 14 and the image memory 16 in accordance with a command from the control unit 20. The B-mode image signal processed by the image processing unit 37 is simply called an ultrasound image.

[0029] The image memory 16 is a memory for storing and retrieving ultrasound images generated by the image generation unit 13 under the control of the control unit 20. The ultrasound images stored in the image memory 16 are read under the control of the control unit 20 and sent to the evaluation unit 19.

[0030] For the image memory 16, recording media such as flash memory, HDD (Hard Disc Drive), SSD (Solid State Drive), FD (Flexible Disc), MO disk (Magneto-Optical disc), MT (Magnetic Tape), RAM (Random Access Memory), CD (Compact Disc), DVD (Digital Versatile Disc), SD card (Secure Digital card), or USB memory (Universal Serial Bus memory) can be used.

[0031] Microphone 4 is located independently of the ultrasound probe 2 and near the subject's throat, and is used to acquire the subject's swallowing sounds as analog data. The swallowing sounds acquired by microphone 4 are sent to the sound processing unit 17. Microphone 4 can be, for example, brought into contact with the subject's pharynx by the user's hand, or it can be attached to the subject's pharynx. Microphone 4 also has a mounting device (not shown) with a frame shape or the like for attaching to the subject's neck, and by attaching this mounting device to the subject's neck, microphone 4 can be positioned near the subject's pharynx.

[0032] The sound processing unit 17 converts the analog data of the swallowing sound acquired by the microphone 4 into digital data and sends the obtained digital data to the sound memory 18. The sound processing unit 17 also generates information representing the swallowing sound, such as a waveform graph showing the time change of the amplitude of the swallowing sound, based on the digital data of the swallowing sound, and sends this information to the sound memory 18.

[0033] The sound memory 18 is a memory for storing and retrieving digital data of swallowing sounds sent from the sound processing unit 17, as well as information representing swallowing sounds, such as waveform graphs, under the control of the control unit 20. The swallowing sound data stored in the sound memory 18 is read under the control of the control unit 20 and sent to the evaluation unit 19. The information representing swallowing sounds stored in the sound memory 18 is also read under the control of the control unit 20 and sent to the display control unit 14, where it is displayed on the monitor 15.

[0034] The evaluation unit 19 evaluates the swallowing of a subject by using machine learning (multimodal learning) that combines ultrasound images transmitted from the image memory 16 and swallowing sound data of the subject transmitted from the sound memory 18.

[0035] As shown in Figure 4, the evaluation unit 19 includes an image analysis unit 38, a sound analysis unit 39, and an evaluation result output unit 40. The evaluation result output unit 40 is connected to the image analysis unit 38 and the sound analysis unit 39, and the display control unit 14 is connected to the evaluation result output unit 40.

[0036] The image analysis unit 38 receives an ultrasound image of the subject's pharynx from the image memory 16 and inputs this ultrasound image into a pre-trained first neural network to calculate image features. Image features are indicators calculated based on the ultrasound image that represent the degree to which there is an abnormality in the subject's swallowing, such as the probability that food remains in the subject's pharynx or the probability that the subject is experiencing dysphagia. A larger image feature value indicates a higher probability that the subject is experiencing a swallowing abnormality, while a smaller image feature value indicates a lower probability that the subject is experiencing a swallowing abnormality.

[0037] In cases where a subject is experiencing swallowing difficulties, food is usually found to be present in the vallecula or piriform fossa within the subject's pharynx. Therefore, if a food-like structure is detected in the vallecula or piriform fossa in the ultrasound image, it can be concluded that there is a high probability that the subject is experiencing swallowing difficulties.

[0038] The first neural network used by the image analysis unit 38 learns features such as structures depicted in ultrasound images based on ultrasound images of the pharynx of multiple subjects, including cases where food remains in the piriform fossa and cases where food does not remain in the piriform fossa. It then outputs image features by comparing the learned features with the features in the input ultrasound images.

[0039] The sound analysis unit 39 receives swallowing sound data from the sound memory 18 and inputs this swallowing sound data into a trained second neural network to calculate sound features. Sound features are indicators calculated based on the swallowing sound data of the subject, representing the degree to which there is an abnormality in the subject's swallowing, such as the probability that food remains in the subject's pharynx or the probability that the subject is experiencing dysphagia. A larger sound feature indicates a higher probability of an abnormality in the subject's swallowing, while a smaller sound feature indicates a lower probability of an abnormality in the subject's swallowing.

[0040] In cases where there is an abnormality in the subject's swallowing, the swallowing sounds often contain abnormal noise and sounds with abnormal frequencies compared to when the subject swallows normally. Therefore, if the swallowing sounds contain abnormal noise, or if frequency analysis of the swallowing sounds reveals peak amplitude values ​​in abnormal frequency bands compared to normal swallowing sounds, it can be concluded that there is a high probability that there is an abnormality in the subject's swallowing.

[0041] The second neural network used by the sound analysis unit 39 outputs sound features by comparing, for example, the typical temporal changes and frequency distribution of the amplitude of a normal swallowing sound, which has been learned based on swallowing sound data from multiple subjects, with the temporal changes and frequency distribution of the amplitude of the input swallowing sound.

[0042] The evaluation result output unit 40 inputs the image features calculated by the image analysis unit 38 and the sound features calculated by the sound analysis unit 39 into a trained third neural network and outputs evaluation results such as whether or not the subject has dysphagia. The third neural network has learned the relationship between combinations of image feature values ​​and sound feature values ​​and evaluation results regarding the subject's swallowing based on evaluations of multiple subjects, and when image features and sound features are input, it outputs a swallowing evaluation result based on the learned results.

[0043] The control unit 20 controls each part of the swallowing evaluation system 1 according to a pre-recorded program or the like. The display control unit 14, under the control of the control unit 20, performs predetermined processing on the ultrasound image, the time change in the amplitude of the subject's swallowing sound, and the swallowing evaluation results, and displays them on the monitor 15.

[0044] The monitor 15 displays various information under the control of the display control unit 14. The monitor 15 includes, for example, display devices such as an LCD (Liquid Crystal Display) or an organic EL display (Organic Electroluminescence Display).

[0045] The input device 21 is for the user to perform input operations. The input device 21 consists of, for example, a keyboard, mouse, trackball, touchpad, and touch panel, or other devices for the user to perform input operations.

[0046] The processor 22, which includes an image generation unit 13, a display control unit 14, a sound processing unit 17, an evaluation unit 19, and a control unit 20, is composed of a CPU (Central Processing Unit) and a control program for causing the CPU to perform various processes. However, it may also be composed of an FPGA (Field Programmable Gate Array), a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), a GPU (Graphics Processing Unit), or other ICs (Integrated Circuits), or a combination thereof.

[0047] Furthermore, the image generation unit 13, display control unit 14, sound processing unit 17, evaluation unit 19, and control unit 20 of the processor 22 can be partially or entirely integrated into a single CPU or the like.

[0048] Next, the operation of the swallowing evaluation system 1 according to Embodiment 1 of the present invention will be explained using the flowchart shown in Figure 5.

[0049] First, the user places the ultrasound probe 2 in contact with the subject's throat and positions the microphone 4 on the subject's throat. In this state, in step S1, the control unit 20 receives an instruction to start acquiring the subject's swallowing sounds for evaluation. The control unit 20 determines that the instruction to start acquiring swallowing sounds has been received, for example, when the user inputs an instruction to start acquiring swallowing sounds via the input device 21. Furthermore, when step S1 is completed, the subject begins swallowing, for example, by prompting the subject to swallow food. Here, the food swallowed by the subject may be, for example, a jelly food commonly available for subjects with dysphagia.

[0050] Next, the microphone 4 captures the swallowing sounds of the subject during swallowing. The analog data of the swallowing sounds captured by the microphone 4 is converted into digital data by the sound processing unit 17 and stored in the sound memory 18. The sound processing unit 17 also generates information representing the swallowing sounds, such as a waveform graph showing the time change of the amplitude of the swallowing sounds, and this information is displayed on the monitor 15, for example, as shown in Figure 6. In the example in Figure 6, a waveform graph W showing the time change of the amplitude of the swallowing sounds is displayed as information representing the swallowing sounds.

[0051] In step S3, the control unit 20 determines whether or not to terminate the acquisition of swallowing sounds. For example, the control unit 20 determines to terminate the acquisition of swallowing sounds if the user inputs an instruction to terminate the acquisition of swallowing sounds via the input device 21, and determines to continue acquiring swallowing sounds if the user does not input an instruction to terminate the acquisition of swallowing sounds via the input device 21.

[0052] If it is determined in step S3 to continue acquiring swallowing sounds, the process returns to step S2, and swallowing sounds are acquired anew. In this way, the processes of steps S2 and S3 are repeated until it is determined in step S3 to terminate the acquisition of swallowing sounds. As a result, the swallowing sounds of the subject acquired through the repetition of steps S2 and S3 are stored in the sound memory 18.

[0053] If, for example, the user determines in step S3 that the subject has finished swallowing and inputs an instruction to stop acquiring swallowing sounds via the input device 21, and it is determined that the acquisition of swallowing sounds has ended, the process proceeds to step S4.

[0054] In step S4, an ultrasound image U of the subject's pharynx after swallowing is obtained. By obtaining an ultrasound image U of the subject after swallowing in this way, if there is any abnormality in the subject's swallowing, an ultrasound image U that visualizes food remaining in the piriform fossa is often obtained.

[0055] When an ultrasound image U is acquired, the transmitting / receiving circuit 12, under the control of the control unit 20, performs reception focus processing using a preset sound velocity value to generate an acoustic ray signal. The acoustic ray signal generated in this way by the transmitting / receiving circuit 12 is sent to the image generation unit 13. The image generation unit 13 generates an ultrasound image using the acoustic ray signal sent from the transmitting / receiving circuit 12. For example, as shown in Figure 6, the ultrasound image U generated in this way is sent to the display control unit 14 and displayed on the monitor 15.

[0056] In step S5, the control unit 20 determines whether or not to save the ultrasound image U acquired in step S4 to the image memory 16. For example, if the user inputs an instruction to save the ultrasound image U via the input device 21, the control unit 20 determines to save the ultrasound image U acquired in step S4 and saves the ultrasound image U to the image memory 16. If the user does not input an instruction to save the ultrasound image U via the input device 21, the control unit 20 determines not to save the ultrasound image U acquired in step S4.

[0057] For example, as shown in Figure 6, the control unit 20 can display a freeze save button B on the monitor 15, and when the freeze save button B is selected by the user via the input device 21, it can freeze the ultrasound image U displayed on the monitor 15 and save the ultrasound image U to the image memory 16. If the freeze save button B is not selected, it can determine that the ultrasound image U should not be saved.

[0058] If it is determined in step S5 not to save the ultrasound image U, the process returns to step S4, and a new ultrasound image U is acquired. Therefore, steps S4 and S5 are repeated until it is determined in step S5 to save the ultrasound image U. If it is determined in step S5 to save the ultrasound image U, and the ultrasound image U is saved in the image memory 16, the process proceeds to step S6.

[0059] In step S6, the evaluation unit 19 evaluates the subject's swallowing by using machine learning that combines the swallowing sound data of the subject during swallowing stored in the sound memory 18 through the repetition of steps S2 and S3, and the ultrasound image U stored in the image memory 16 in step S5. The process in step S6 will be explained in detail using the flowchart shown in Figure 7.

[0060] When the processing of step S6 begins, the processing of step S8 is performed first. In step S8, the image analysis unit 38 of the evaluation unit 19 inputs the ultrasound image U stored in the image memory 16 in step S5 into the first neural network to calculate image features that represent the degree to which the subject has an abnormal swallowing problem.

[0061] In this process, the first neural network compares features learned from ultrasound images of multiple subjects' pharynges, including cases where food remains in the piriform fossa and cases where food does not, with features in the input ultrasound image U. The network then outputs the probability that the remaining food is depicted in the ultrasound image U as an image feature. For example, the clearer a structure resembling the remaining food is depicted in the ultrasound image U, the larger the image feature value obtained.

[0062] Next, in step S9, the sound analysis unit 39 of the evaluation unit 19 inputs the swallowing sound data obtained by repeating steps S2 and S3 into a second neural network to calculate sound features that represent the degree to which there is an abnormality in the subject's swallowing.

[0063] In this process, the second neural network outputs sound features by comparing the typical temporal changes and frequency distribution of the amplitude of a normal swallowing sound, learned based on data of swallowing sounds from multiple subjects, with the temporal changes and frequency distribution of the input swallowing sound amplitude. For example, a larger sound feature value is obtained when the amplitude of the noise sound is detected at an abnormal timing compared to the typical temporal changes of the amplitude of a normal swallowing sound, or when the peak value of the amplitude is detected in an abnormal frequency band compared to the typical frequency distribution of a normal swallowing sound.

[0064] In the subsequent step S10, the evaluation result output unit 40 of the evaluation unit 19 outputs the evaluation results of the subject's swallowing by inputting the image features calculated in step S8 and the sound features calculated in step S9 into a third neural network. The evaluation results of swallowing include, for example, whether or not there is residual food in the piriform fossa, or whether or not there is suspicion of dysphagia in the subject. The evaluation result output unit 40 can also output the evaluation results of swallowing in multiple stages, such as "no residue," "a little residue," "a lot of residue," and "a lot of residue," for example, the amount of food remaining in the piriform fossa. Furthermore, the evaluation result output unit 40 can also output the appropriate hardness of a dysphagia-friendly food for the subject being evaluated as part of the swallowing evaluation results. For example, the hardness of dysphagia-friendly food described in "Swallowing Adjustment Society Classification 2013, Japanese Journal of Swallowing Rehabilitation 17(3):255-267, 2013" can be used.

[0065] In this way, in step S10, the final evaluation result of the subject's swallowing is output based on both image features representing the degree of abnormality in the subject's swallowing, calculated based on the analysis of the ultrasound image U, and sound features representing the degree of abnormality in the subject's swallowing, calculated based on the analysis of the subject's swallowing sounds. Therefore, for example, it is possible to obtain an evaluation result with higher accuracy than if the subject's swallowing were evaluated using only one of the image features calculated based on the analysis of the ultrasound image U or the sound features calculated based on the analysis of the subject's swallowing sounds.

[0066] In this way, the processes in steps S8 to S10 are performed, and the process in step S6 is carried out. Once the process in step S6 is completed, the process proceeds to step S7. In step S7, information representing the swallowing evaluation results of the subject, which were output in step S6, is displayed on monitor 15. This completes the operation of the swallowing evaluation system 1 according to Embodiment 1.

[0067] As described above, according to the swallowing evaluation system 1 of Embodiment 1 of the present invention, the swallowing of the subject is evaluated by machine learning that combines both the ultrasound image U of the subject's pharynx and the subject's swallowing sounds, thus enabling highly accurate evaluation of the subject's swallowing.

[0068] As shown in Figure 1, the transmitting and receiving circuit 12 is provided in the ultrasonic probe 2, but it may also be provided in the main body of the device 3 instead. Furthermore, although the image generation unit 13 is provided in the main body of the device 3, it may also be provided in the ultrasonic probe 2 instead of being provided in the main body of the device 3. Furthermore, as shown in Figure 3, the image generation unit 13 includes a signal processing unit 35, a DSC 36, and an image processing unit 37, but the signal processing unit 35 can also be included in the ultrasonic probe 2.

[0069] Furthermore, the method of connecting the ultrasound probe 2 and the main unit 3 of the device is not particularly limited; it may be a wired connection or a wireless connection. Similarly, the method of connecting the main unit 3 and the microphone 4 is not particularly limited; it may be a wired connection or a wireless connection. Furthermore, the main unit 3 of the device may be a so-called handheld type that can be easily carried by the user, or it may be a so-called stationary type.

[0070] Furthermore, although it is explained that microphone 4 is independent of ultrasound probe 2, it may also be attached to ultrasound probe 2, such as being built into it. In this case, there is no need to place microphone 4 independently near the subject's pharynx, and the subject's swallowing sounds can also be acquired by bringing ultrasound probe 2 into contact with the subject's pharynx to capture the ultrasound image U.

[0071] Furthermore, while it has been explained that the sound processing unit 17 and microphone 4 of the main unit 3 constitute a sound acquisition unit for acquiring the swallowing sounds of the subject, the swallowing evaluation system 1 may also have a sound acquisition unit that has a microphone 4 and a sound processing unit 17 and is independent of the main unit 3. In this case, the swallowing sound data of the subject acquired by the sound acquisition unit is input to the main unit 3.

[0072] Furthermore, in steps S2 to S5, after acquiring and saving swallowing sounds during swallowing, ultrasound images U after swallowing are acquired and saved. However, ultrasound images U during swallowing may also be acquired and saved. In this case, when image features are calculated in step S8, the ultrasound images U during swallowing can be analyzed in addition to the ultrasound images U after swallowing, thereby improving the accuracy of the image features.

[0073] Alternatively, multiple continuous frames of ultrasound images U can be acquired over a certain period of time from the start of swallowing to the end of swallowing, and these ultrasound images U can be stored in the image memory 16. In this case, image features can be calculated by analyzing the time-series data of the ultrasound images U acquired within a certain period of time in step S8. This allows for the analysis of the movement of structures within the pharynx based on multiple frames of ultrasound images U, for example, enabling the calculation of image features with higher accuracy than, for example, calculating image features based on a single frame of ultrasound image U.

[0074] Furthermore, if there is any abnormality in the subject's swallowing, it is possible that the respiratory sounds will contain abnormal noise and different respiratory sounds will be acquired compared to when swallowing is performed normally. For this reason, for example, the sound acquisition unit having a microphone 4 and a sound processing unit 17 can acquire data on swallowing sounds during swallowing, as well as data on respiratory sounds before swallowing and after swallowing, and this respiratory sound data can be stored in the sound memory 18. In this case, in step S9, the subject's swallowing sound data and respiratory sound data are analyzed to calculate sound features. This makes it possible to calculate sound features with higher accuracy than, for example, calculating sound features based on the subject's swallowing sounds.

[0075] Furthermore, in the flowchart shown in Figure 7, step S9 is performed after step S8, but step S8 may be performed after step S9, or steps S8 and S9 may be performed simultaneously.

[0076] Furthermore, the method of notifying the user of the evaluation results is not limited to displaying the evaluation results on the monitor 15. For example, the swallowing evaluation system 1 may be equipped with a lamp (not shown), and the user may be notified of the presence or absence of swallowing impairment in the subject by the lamp's light color and flashing pattern. In this case, the lamp may be placed on the ultrasound probe 2, on the main body of the device 3, or independently of the ultrasound probe 2 and the main body of the device 3.

[0077] Embodiment 2 In Embodiment 1, the evaluation unit 19 includes an image analysis unit 38, a sound analysis unit 39, and an evaluation result output unit 40, and evaluates the swallowing of a subject using three neural networks: a first neural network, a second neural network, and a third neural network. However, the method for evaluating the swallowing of a subject is not limited to this.

[0078] For example, the evaluation unit 19 can evaluate the subject's swallowing using only the same neural network that has learned the relationship between the subject's ultrasound image U, the subject's swallowing sounds, and combinations thereof, and the evaluation results.

[0079] Thus, even when the swallowing evaluation system 1 outputs swallowing evaluation results using only one neural network, it can evaluate the subject's swallowing with high accuracy by performing machine learning that combines both the ultrasound image U of the subject's pharynx and the subject's swallowing sounds, just as it does when using three neural networks as in Embodiment 1.

[0080] Embodiment 3 For example, it is possible to use two neural networks to output an evaluation of the subject's swallowing ability. Figure 8 shows the configuration of the evaluation unit 19A in Embodiment 3. The evaluation unit 19A has a sound analysis unit 39 and an evaluation result output unit 40A, with the evaluation result output unit 40A connected to the sound analysis unit 39. The display control unit 14 is also connected to the evaluation result output unit 40A.

[0081] The sound analysis unit 39, similar to the sound analysis unit 39 in Embodiment 1, receives data of the subject's swallowing sounds and inputs this swallowing sound data into a second neural network to calculate sound features. The calculated sound features are output to the evaluation result output unit 40A.

[0082] The evaluation result output unit 40A has a defined sound feature threshold for sound features and determines whether the sound features calculated by the sound analysis unit 39 exceed the sound feature threshold. Furthermore, if the sound features calculated by the sound analysis unit 39 exceed the sound feature threshold, the evaluation result output unit 40A outputs the evaluation result of the subject's swallowing by inputting those sound features and the ultrasound image U generated by the image generation unit 13 into the fourth neural network.

[0083] The fourth neural network used by the evaluation result output unit 40A calculates image features by analyzing the input ultrasound image U, similar to the first neural network used by the image analysis unit 38 in Embodiment 1, and outputs the evaluation result of the subject's swallowing based on the calculated image features and sound features transmitted from the sound analysis unit 39, similar to the third neural network used by the evaluation result output unit 40 in Embodiment 1.

[0084] Next, the swallowing evaluation process in Embodiment 3 will be explained using the flowchart shown in Figure 9. First, in step S11, the sound analysis unit 39 calculates sound features by inputting the swallowing sound data of the subject into the second neural network.

[0085] In step S12, the evaluation result output unit 40A determines whether the sound feature quantities calculated in step S11 exceed the sound feature threshold. If it is determined that the sound feature quantities calculated in step S11 exceed the sound feature threshold, the evaluation result output unit 40A determines that there is a high possibility that some abnormality is occurring in the subject's swallowing and proceeds to the process in step S13.

[0086] In step S13, the evaluation result output unit 40A inputs the sound features calculated in step S11 and the ultrasound image U of the subject's pharynx into the fourth neural network, and outputs a swallowing evaluation result that takes into account the analysis results of the ultrasound image U.

[0087] Furthermore, if in step S12 it is determined that the sound feature quantity is below the sound feature quantity threshold, the evaluation result output unit 40A determines that there is little possibility that the subject has an abnormality in swallowing and proceeds to the process in step S14.

[0088] In step S14, the evaluation result output unit 40A calculates the swallowing evaluation result based on the swallowing sound analysis result in step S11. For example, without using a fourth neural network, the evaluation result output unit 40A determines from the result in step S12 that the sound feature quantity is below the sound feature quantity threshold that there is a low possibility that there is an abnormality in the subject's swallowing, and outputs an evaluation result that there is no abnormality in the subject's swallowing. In this case, the evaluation result output unit 40A can output the evaluation result of the subject's swallowing using only the second neural network. In this way, the swallowing evaluation operation in Embodiment 3 is completed.

[0089] Thus, in Embodiment 3, two neural networks, the second neural network and the fourth neural network, are used only when the sound feature quantity exceeds the sound feature quantity threshold, and only the second neural network is used when the sound feature quantity is below the sound feature quantity threshold. Therefore, compared to Embodiment 1, the computational load on the processor 22 when using the neural networks can be reduced. Consequently, in Embodiment 3, the processing time and power consumption required to evaluate the swallowing of a subject, which are caused by the computational load on the processor 22, can be reduced.

[0090] In this case, if the device body 3 is a so-called handheld or portable type, it is preferable that power is supplied to each part of the device body 3 by a portable battery (not shown). Also, if the device body 3 is a handheld or portable type, it is smaller in size compared to the case where the device body 3 is a so-called stationary type, making it difficult to install a large and high-performance processor 22. For this reason, the embodiment of the third embodiment is particularly useful when the device body 3 is a handheld or portable type.

[0091] Furthermore, it is generally known that in ultrasound images U of the subject's pharynx, food remaining in the piriform fossa and pharynx is difficult to distinguish from surrounding tissue. On the other hand, swallowing sounds can be acquired with relatively high sensitivity, so it is possible to make a more accurate judgment by judging whether or not there is an abnormality in the subject's swallowing than by judging by the swallowing sounds rather than by the ultrasound images U. In Embodiment 3, sound features are first calculated, and if the calculated sound features exceed the sound feature threshold, the swallowing sounds are evaluated by taking into account the analysis results of the ultrasound images U, thereby enabling a highly accurate evaluation of the subject's swallowing.

[0092] Embodiment 4 Figure 10 shows the configuration of the evaluation unit 19B in Embodiment 4. The evaluation unit 19B has an image analysis unit 38 and an evaluation result output unit 40B, with the evaluation result output unit 40B connected to the image analysis unit 38. In addition, the display control unit 14 is connected to the evaluation result output unit 40B.

[0093] The image analysis unit 38, similar to the image analysis unit 38 in Embodiment 1, receives an ultrasound image U of the subject's pharynx and inputs the ultrasound image U into the first neural network to calculate image features. The calculated image features are output to the evaluation result output unit 40B.

[0094] The evaluation result output unit 40B has a predetermined image feature threshold for image features and determines whether the image features calculated by the image analysis unit 38 exceed the image feature threshold. Furthermore, if the image features calculated by the image analysis unit 38 exceed the sound feature threshold, the evaluation result output unit 40B inputs the image features and the swallowing sound data of the subject acquired by the sound acquisition unit, which has a microphone 4 and a sound processing unit 17, into a fifth neural network to output the evaluation result of the subject's swallowing.

[0095] The fifth neural network used by the evaluation result output unit 40B calculates sound features by analyzing the input swallowing sound data, similar to the second neural network used by the sound analysis unit 39 in Embodiment 1, and outputs the evaluation result of the subject's swallowing based on the calculated sound features and image features sent from the image analysis unit 38, similar to the third neural network used by the evaluation result output unit 40 in Embodiment 1.

[0096] Next, the swallowing evaluation process in Embodiment 4 will be explained using the flowchart shown in Figure 11. First, in step S21, the image analysis unit 38 calculates image features by inputting the ultrasound image U of the subject's pharynx into the first neural network.

[0097] In step S22, the evaluation result output unit 40B determines whether the image feature quantities calculated in step S21 exceed the image feature threshold. If, in step S22, it is determined that the image feature quantities calculated in step S21 exceed the image feature threshold, the evaluation result output unit 40B determines that there is a high probability that some abnormality is occurring in the subject's swallowing and proceeds to the process in step S23.

[0098] In step S23, the evaluation result output unit 40B inputs the image features calculated in step S21 and the swallowing sound data of the subject into the fifth neural network, and outputs a swallowing evaluation result that takes into account the swallowing sound analysis results.

[0099] Furthermore, if in step S22 the image feature quantity is determined to be below the image feature quantity threshold, the evaluation result output unit 40B determines that there is little possibility that the subject has an abnormality in swallowing and proceeds to the process in step S24.

[0100] In step S24, the evaluation result output unit 40B calculates the swallowing evaluation result based on the swallowing sound analysis result in step S21. For example, without using a fifth neural network, the evaluation result output unit 40B determines from the result in step S22 that the image feature quantity is below the image feature quantity threshold that there is a low possibility that there is an abnormality in the subject's swallowing, and outputs an evaluation result that there is no abnormality in the subject's swallowing. In this case, the evaluation result output unit 40B can output the evaluation result of the subject's swallowing using only the first neural network. In this way, the swallowing evaluation operation in Embodiment 4 is completed.

[0101] Thus, in Embodiment 4, two neural networks, the first neural network and the fifth neural network, are used only when the image features exceed the image feature threshold, and only the first neural network is used when the image features are below the image feature threshold. Therefore, compared to Embodiment 1, the computational load on the processor 22 when using the neural networks can be reduced. As a result, in Embodiment 4, similar to Embodiment 3, the processing time and power consumption required to evaluate the swallowing of the subject, which are caused by the computational load on the processor 22, can be reduced.

[0102] Furthermore, as described above, the embodiment of Embodiment 4 is particularly useful when the device body 3 is handheld or portable, similar to the embodiment of Embodiment 3. [Explanation of symbols]

[0103] 1 Swallowing evaluation system, 2 Ultrasound probe, 3 Main unit, 4 Microphone, 11 Transducer array, 12 Transmit / receive circuit, 13 Image generation unit, 14 Display control unit, 15 Monitor, 16 Image memory, 17 Sound processing unit, 18 Sound memory, 19, 19A, 19B Evaluation unit, 20 Control unit, 21 Input device, 22 Processor, 31 Pulsar, 32 Amplifier, 33 AD conversion unit, 34 Beamformer, 35 Signal processing unit, 36 DSC, 37 Image processing unit, 38 Image analysis unit, 39 Sound analysis unit, 40, 40A, 40B Evaluation result output unit, B Freeze save button, U Ultrasound image, W Waveform graph.

Claims

1. Ultrasound probe and An image acquisition unit that acquires an ultrasound image of the inside of the subject's pharynx by transmitting and receiving an ultrasound beam using the ultrasound probe, A sound acquisition unit that acquires swallowing sounds from the subject, An evaluation unit that evaluates the swallowing of the subject by using machine learning that combines the ultrasound image acquired by the image acquisition unit and the swallowing sound acquired by the sound acquisition unit. Equipped with, The evaluation unit is a swallowing evaluation system that outputs, as evaluation results, the presence or absence of swallowing residue in the pharyngeal region of the subject and the appropriate hardness of a swallowing food for the subject, or the appropriate hardness of a swallowing food for the subject.

2. The swallowing evaluation system according to claim 1, wherein the evaluation unit evaluates swallowing by inputting the ultrasound image and the swallowing sound into a neural network.

3. The swallowing evaluation system according to claim 2, wherein the evaluation unit inputs the ultrasound image into a first neural network to calculate image features, inputs the swallowing sound into a second neural network to calculate sound features, and inputs the image features and sound features into a third neural network to evaluate swallowing.

4. The swallowing evaluation system according to claim 2, wherein the evaluation unit inputs the swallowing sound into a second neural network to calculate sound features, and if the sound features exceed a predetermined sound feature threshold, the system evaluates swallowing by inputting the ultrasound image and the sound features into a fourth neural network.

5. The swallowing evaluation system according to claim 2, wherein the evaluation unit inputs the ultrasound image to a first neural network to calculate image features, and when the image features exceed a predetermined image feature threshold, the swallowing sound and the image features are input to a fifth neural network to evaluate swallowing.

6. The swallowing evaluation system according to claim 2, wherein the evaluation unit evaluates swallowing by inputting both the ultrasound image and the swallowing sound into the same neural network.

7. The swallowing evaluation system according to any one of claims 2 to 6, wherein the evaluation unit inputs the time-series data of the ultrasound image and the swallowing sound into the neural network.

8. The sound acquisition unit acquires swallowing sounds during swallowing. The swallowing evaluation system according to any one of claims 1 to 7, wherein the image acquisition unit acquires an ultrasound image after swallowing.

9. The swallowing evaluation system according to claim 8, wherein the image acquisition unit also acquires ultrasound images during swallowing.

10. The swallowing evaluation system according to claim 8 or 9, wherein the sound acquisition unit also acquires at least one of the breath sounds, namely the breath sounds before swallowing and the breath sounds after swallowing.

11. The swallowing evaluation system according to any one of claims 1 to 10, wherein the sound acquisition unit has a microphone built into the ultrasonic probe.

12. The swallowing evaluation system according to any one of claims 1 to 10, wherein the sound acquisition unit has a microphone that is independent of the ultrasonic probe and is in contact with the pharyngeal region of the subject.

13. The swallowing evaluation system according to any one of claims 1 to 12, wherein the evaluation unit further outputs whether or not the subject has swallowing difficulties as an evaluation result.

14. A swallowing evaluation system according to any one of claims 1 to 13, comprising a monitor that displays the ultrasound image acquired by the image acquisition unit and information representing the swallowing sound acquired by the sound acquisition unit.

15. By using an ultrasound probe to transmit and receive ultrasound beams, an ultrasound image of the pharynx of the subject is obtained. The swallowing sounds of the subject are obtained, The swallowing of the subject is evaluated by using machine learning with the acquired ultrasound image and swallowing sound as input. The evaluation results output the presence or absence of swallowing residue in the pharynx of the subject and the appropriate consistency of swallowing food for the subject, or the appropriate consistency of swallowing food for the subject. Swallowing evaluation methods.

Citation Information

Patent Citations

  • Food kit for food test for ingestion / Swallow difficulty judgement

    JP2003119158A

  • Living body movement identification system and living body movement identification method

    JP2018000871A

  • Swallowing ability measuring system, swallowing ability measuring method, and sensor holder

    JP2020089613A

  • Apparatus, system and method for evaluation of esophageal function

    US20050124888A1