Artificial Intelligence-Assisted Calculation Method and System for Pharyngeal Imaging Parameters
Patent Information
- Application Number
- TW111143975
- Authority / Receiving Office
- TW · TW
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-11-17
- Publication Date
- 2024-09-11
- Estimated Expiration
- 2042-11-16
Smart Images

Figure TWG2TB001786892_001 
Figure TWG2TB001786892_002 
Figure TWG2TB001786892_003
Abstract
Description
[Technical Field]
[0001] This invention relates to an AI-assisted computation in the field of smart healthcare; and more particularly to a method and system for AI-assisted interpretation of vocal cord images and calculation parameters. [Previous Technology]
[0002] The normalized glottal gap area is an important medical parameter used in clinical research on voice medicine to assess the vocal cord phonation status. This medical parameter is obtained by first obtaining an image of the larynx using larynngeal stroboscopy or flexible endoscopy, then circling the glottal gap in the laryngeal image and calculating the area of the glottal gap, and after measuring the length of the vocal cords, manually substituting the values into the formula "(glottal gap area / vocal cord length squared) x 100" to obtain the so-called normalized glottal gap area.
[0003] The concept of standardized glottic space area first appeared in Omori's 1996 publication and has been widely used in various voice medicine literature since then. However, the calculation of standardized glottic space area requires first downloading laryngeal images to a computer, manually selecting the glottic space in the laryngeal images, and then using image processing software such as ImageJ to perform an integral calculation on the area of the glottic space using the area measurement function to obtain the glottic space area in the laryngeal images. Because this medical parameter is selected manually when selecting the glottic space in the laryngeal images, and the subsequent operation of image processing software to calculate the area of the selected range is also time-consuming, the clinical application value of this method is not high. [Summary of the Invention]
[0004] In view of this, the purpose of the present invention is to provide an artificial intelligence object detection technology that, after extracting the glottic image from medical images or videos by identifying vocal cord features, uses artificial intelligence image recognition and segmentation technology to mark the boundary and range of the glottis and thereby obtain the numerical value of medical features, providing physicians with the assistance of quantitatively interpreting the state of the vocal cords.
[0005] To achieve the above objective, the present invention provides an artificial intelligence-assisted method for calculating pharyngeal image parameters. The method includes the following steps: training a model by using multiple laryngeal images, each with a manually selected glottic image range, to train a deep learning object detection software to extract the glottic image from the received laryngeal images; and using multiple glottic images, each with a manually marked anterior glottic gap, to train a deep learning image recognition and segmentation software to identify the anterior glottic gap from the received glottic images. The anterior glottic gap includes the structural features of the left vocal cord process, the right vocal cord process, and the anterior commissure.
[0006] Receive laryngeal images, including a single laryngeal image of the vocal cords in a vocalizing state, multiple laryngeal images captured frame-by-frame, or a laryngeal video. Identify glottal images: using deep learning object detection software, extract one or more glottal images from the aforementioned laryngeal images or the laryngeal video. Identify the anterior glottal gap in the glottal images: using deep learning image recognition and segmentation software, identify the anterior glottal gap in each of the aforementioned glottal images, and output an anterior glottal gap mask corresponding to each glottal image. Obtain medical parameters: perform image processing such as outlining and repair on each anterior glottal gap mask to clearly depict the anterior glottal gap in each anterior glottal gap mask, and obtain medical parameters of the vocal cord anatomy from the clearly depicted anterior glottal gap in each anterior glottal gap mask.
[0007] To achieve the above objectives, the present invention provides an artificial intelligence-assisted system for calculating pharyngeal image parameters, comprising an input unit, a processing unit, and an output unit. The input unit receives a medical image, including a single laryngeal image with the vocal cords in a vocalized state, multiple frame-by-frame laryngeal images, or laryngeal videos. The processing unit is signal-connected to the input unit and executes a deep learning algorithm to calculate the medical image received by the input unit. The deep learning algorithm includes deep learning object detection software and deep learning image recognition and segmentation software.
[0008] The processing unit uses deep learning object detection software to extract one or more glottic images from one or more laryngeal images or laryngeal films; it uses deep learning image recognition and segmentation software to identify the anterior glottic gap in each of the aforementioned glottic images, and outputs an anterior glottic gap mask for each glottic image; the processing unit performs image processing such as outlining and repair on each anterior glottic gap mask to clearly depict the anterior glottic gap in each anterior glottic gap mask, obtains one or more medical parameters of vocal cord anatomy from the clearly depicted anterior glottic gap in each anterior glottic gap mask, and adds an anatomical marker to the medical image corresponding to the position of each medical parameter. The output unit is signal-connected to the processing unit, and the processing unit receives the medical image with the one or more markers and the one or more medical parameters, and outputs the medical image and the one or more medical parameters as a medical parameter and image report.
[0009] The advantage of this invention lies in utilizing artificial intelligence to identify anatomical features of the vocal cords in medical images, such as laryngeal images and glottic images, and performing rapid or real-time calculations to obtain standardized information on medical parameters such as the area of the anterior glottic gap and the amplitude of vocal fold vibration, for use in assisting in the assessment of vocal cord condition. Simultaneously, it avoids the problem that in traditional methods of obtaining image information such as incomplete glottal closure and mucosal wave from laryngeal images obtained through frame-by-frame laryngeal stroboscopic photography, the location and extent of anatomical structures are primarily determined manually, resulting in subjective judgments that are difficult to standardize. [Simplified Explanation of the Diagram]
[0029] Figure 1 is a flowchart of the steps of a preferred embodiment of the present invention.
[0030] Figures 2A to 2C are schematic diagrams of the larynx when the vocal cords are relaxed.
[0031] Figure 3 is a schematic diagram of the larynx during vocalization.
[0032] Figures 4A and 4B are schematic diagrams of the training model of the preferred embodiment of the present invention.
[0033] Figures 5A to 5D are identification images and schematic diagrams of the above-described preferred embodiments of the present invention.
[0034] Figure 6 is a schematic diagram of the above-described preferred embodiment of the present invention.
[0035] Figure 7 is a block diagram of a preferred embodiment of the system of the present invention.
Implementation Method
[0010] To more clearly illustrate the present invention, preferred embodiments are described in detail below with reference to the accompanying drawings. Please refer to Figure 1, which is a flowchart of the steps of a method for artificial intelligence-assisted calculation of pharyngeal image parameters according to a preferred embodiment of the present invention. The steps of the method include:
[0011] (S01) Training Model: Please refer to Figures 2A and 2B, which show images of the larynx when the vocal cords are relaxed. Important anatomical structures of the vocal cords include the left vocal cord process 10, the right vocal cord process 11, the anterior commissure 12, the posterior commissure 13, the vocal cord length L, the glottal angle θ, and the glottis 16. The glottis 16 is the part between the anterior commissure 12 and the posterior commissure 13 that has a gap to allow respiratory gas to pass through. This gap is called the glottal gap 17. The vocal cord length L is the straight-line length from the left vocal cord process 10 or the right vocal cord process 11 to the anterior commissure 12; the glottal angle θ is the angle between the line connecting the left vocal cord process 10 to the anterior commissure 12 and the line connecting the right vocal cord process 11 to the anterior commissure 12. The degree of opening and closing of the glottis 16 can be determined from the glottal angle θ.
[0012] Referring to Figure 2C, the area of the glottic gap 17 is divided into the anterior glottic gap 171 adjacent to the anterior commissure 12 and the posterior glottic gap 172 located on the other side by the line connecting the left vocal cord process 10 and the right vocal cord process 11. Existing literature and technology use the glottic gap 17 to assess the condition of the vocal cords, all of which use a standardized glottic gap area for assessment. However, the area of the glottic gap 17 is easily affected by the larger posterior glottic gap 172, which does not significantly help with phonation. Therefore, this invention uses the standardized anterior glottic gap area calculated from the anterior glottic gap 171 as a medical parameter, which can provide physicians with more accurate numerical values for quantitative assessment of the condition of the vocal cords.
[0013] Please refer to Figure 3, which shows the glottis image when the vocal cords are clamped together, causing the left vocal cord protuberance 10 and the right vocal cord protuberance 11 to come together. This image is obtained by capturing the portion corresponding to the glottis 16 from the laryngeal image. The laryngeal images or videos obtained by the above-mentioned laryngeal stroboscopy or flexible endoscopy are taken in the vocal state. At this time, the airflow passes through the narrow glottis, causing the mucous membrane of the vocal cords to produce wave-like mucous membrane undulations and produce sound. By taking laryngeal stroboscopy or video frame by frame, laryngeal images ordered in chronological order are obtained. Please refer to Figures 2A, 2C, and 3. The vocal function of the subject is evaluated by the standardized anterior glottic gap area 171 and vocal cord amplitude L1 of different vocal cord shapes that are in the vocal state and undergo mucosal undulation. The vocal cord amplitude L1 is the longest distance between the line connecting the left vocal cord process 10 and the right vocal cord process 11 with the anterior commissure 12 when they are closed and the side edge of the anterior glottic gap 171 along the direction perpendicular to the line connecting them. The continuous state of the vocal cords undergoing mucosal undulation can be determined by the vocal cord amplitude L1 at different times.
[0014] When performing the training model S01 step, as shown in Figure 4A, a deep learning object detection software A is trained using multiple laryngeal images 20, each of which has a manually selected glottic image range. In this preferred embodiment, the deep learning object detection software A uses YOLO, and the deep learning object detection software A extracts the glottic image 21 from the received laryngeal image 20. A deep learning image recognition and segmentation software B is trained using multiple glottal images 21, each manually labeled with anterior glottal gap 171. In this preferred embodiment, the deep learning image recognition and segmentation software B employs U-Net. The deep learning image recognition and segmentation software B identifies the anterior glottal gap 171 from the received glottal images 21. The anterior glottal gap 171 is surrounded by the structural features of the left vocal cord process 10, the right vocal cord process 11, and the anterior commissure 12. When training the deep learning image recognition and segmentation software B, the structural features surrounding the anterior glottal gap 171 are also labeled, so that the deep learning image recognition and segmentation software B can also identify the aforementioned structural features contained in the anterior glottal gap 171.
[0015] (S02) Receive laryngeal images: Receive medical images of the laryngeal vocalization state, including a laryngeal image with the vocal cords in a vocalization state, multiple laryngeal images or laryngeal films taken frame by frame, such as receiving multiple laryngeal images taken frame by frame in this preferred embodiment.
[0016] (S03) Identifying the glottal image: Using the object detection software A trained as shown in Figure 4A, extract a glottal image 21 from each of the aforementioned multiple laryngeal images 20, as shown in Figures 5A and 5B, thereby obtaining multiple glottal images 21 corresponding to the multiple laryngeal images 20. In other preferred embodiments, one or more laryngeal images 20 may be received, and the object detection software A extracted a glottal image 21 from each of the aforementioned laryngeal images 20; or a laryngeal video may be received, and the object detection software A extracted multiple frames of laryngeal images 20 from the aforementioned laryngeal video in chronological order, and a glottal image 21 was extracted from each of the multiple frames of laryngeal images 20, thereby obtaining multiple glottal images 21 corresponding to the multiple laryngeal images 20.
[0017] (S04) Identifying the anterior glottal gap in the glottal images: The deep learning image recognition and segmentation software B, trained as shown in FIG4B, identifies the anterior glottal gap 171 in the aforementioned multiple glottal images 21 respectively. As shown in FIG5C, an anterior glottal gap mask 21A is output for each glottal image 21, thereby obtaining multiple anterior glottal gap masks 21A for the multiple glottal images 21. In other preferred embodiments, the deep learning image recognition and segmentation software B can also identify the anterior glottal gap 171 in a single or multiple glottal images 21 respectively, and output an anterior glottal gap mask 21A for each glottal image 21, thereby obtaining a single or multiple anterior glottal gap masks 21A for the single or multiple glottal images 21. In other preferred embodiments, in the steps of training the model S01, receiving the laryngeal image S02, identifying the glottal image S03, and identifying the anterior glottal gap of the glottal image S04, the object detection software A and the image recognition and segmentation software B used in the present invention are not limited to YOLO and U-Net, but can also be other deep learning artificial intelligence models with the same functions.
[0018] (S05) Obtaining medical parameters: Image processing of each anterior glottal gap mask 21A is performed by outlining and repairing, as shown in Figure 5D. The anterior glottal gap 171 in each anterior glottal gap mask 21A is clearly outlined. From the clearly outlined anterior glottal gap 171 in each anterior glottal gap mask 21A, medical parameters of the vocal cord anatomy are obtained. The medical parameters include, but are not limited to, the standardized anterior glottal gap area and vocal cord amplitude L1 of the anterior glottal gap 171. The values of various medical parameters are provided to physicians to make quantitative interpretations of the vocal cord condition in practical applications.
[0019] The method for obtaining the normalized glottal gap area is to calculate the glottal gap area of each glottal image 21 by using the clearly depicted glottal gap 171 of each glottal gap mask 21A, and to obtain the vocal cord length L from the same clearly depicted glottal gap 171. The vocal cord length L is the straight line length from the left vocal cord protrusion 10 or the right vocal cord protrusion 11 to the anterior commissure 12. The normalized glottal gap area (NAGGA) is calculated by substituting the glottal gap area and the vocal cord length L into the formula "(glottal gap area / vocal cord length squared) x 100".
[0020] As mentioned above, the method for obtaining the vocal cord amplitude L1 is to calculate the longest distance between the side edge of the anterior glottis 171 and the line connecting the left vocal cord protrusion 10 and the right vocal cord protrusion 11 when they are brought together and the anterior commissure 12, as indicated by the clearly depicted anterior glottis gap 171 in the anterior glottis gap mask 21A, along the direction perpendicular to the line.
[0021] (S06) Drawing Report Charts: After obtaining medical parameters S05, as shown in Figure 6, a chart 30 is drawn with time or frame number as the horizontal axis and the standardized anterior glottic gap area as the vertical axis. This chart 30 serves as an auxiliary judgment of the vocal cord condition. This step can further add the anatomical structure markers corresponding to the above-mentioned medical parameters to the medical images of the glottic image 21 or the laryngeal image 20. For example, the anatomical structure marker of the vocal cord amplitude L1 can be added to the glottic image 21. The chart 30, multiple glottic images 21 with anatomical structure markers, and the values of the medical parameters corresponding to the multiple glottic images 21 are output as a medical parameter and image report.
[0022] The medical parameters and images in the above-mentioned medical parameter and image report may include more than one type of medical parameter. Since the medical parameter data has been obtained when the medical parameter acquisition step S05 is performed, this step is a step that can be selectively performed or not performed, in order to further produce an intuitive and easy-to-read report output from the medical parameters and medical images obtained in the medical parameter acquisition step S05.
[0023] In the step of identifying the anterior glottal gap S04 of the glottal image, since this preferred embodiment uses U-Net's deep learning image recognition and segmentation software B, as shown in FIG5C, in the process of using U-Net to identify the glottal gap 171, it is necessary to first compress multiple glottal images into a small-sized version of the glottal image 21, for example, a small size of 128 pixels x 128 pixels. Then, the U-Net deep learning image recognition and segmentation software B identifies the anterior glottal gap 171 in the aforementioned small-sized version of the glottal image 21 respectively, and outputs a 128-pixel x 128-pixel small-sized version of the anterior glottal gap mask 21A corresponding to each small-sized version of the glottal image 21. The anterior glottal gap mask 21A of each small-sized version is restored to the original size version, that is, the original size of the anterior glottal gap mask 21A of 384 pixels x 540 pixels, and outputs it to provide the image processing for subsequent outlining and repair.
[0024] The present invention also provides an artificial intelligence-assisted calculation system for pharyngeal image parameters. Please refer to Figures 2A to 4B and Figure 7. In this preferred embodiment, the system is a computer system and includes an input unit 40, a processing unit 41 and an output unit 42 for executing the above-described artificial intelligence-assisted calculation method for pharyngeal image parameters. The input unit 40 may be an input interface, a card reader or a network card, for receiving a medical image. The medical image includes a laryngeal image 20 with the vocal cords in a vocal state, a plurality of laryngeal images 20 taken frame by frame, or a laryngeal video.
[0025] The processing unit 41 includes a computing, control, and memory module. The processing unit 41 is signal-connected to the input unit 40 to receive the medical image. The processing unit 41 executes a deep learning algorithm to process the medical image received by the input unit. The deep learning algorithm includes a deep learning object detection software A and a deep learning image recognition and segmentation software B. In this preferred embodiment, the deep learning object detection software A adopts YOLO and is trained on multiple laryngeal images 20, each of which has a manually selected range of glottic images 21. The deep learning image recognition and segmentation software B adopts U-Net and is trained on multiple glottic images 21, each of which has a manually marked anterior glottic gap 171.
[0026] The processing unit 41 performs the above-described step of identifying glottal images S03. As shown in FIG5A, the deep learning object detection software A extracts one or more glottal images 21 from one or more laryngeal images 20 or laryngeal films. The processing unit 41 performs the above-described step of identifying the anterior glottal gap S04 of the glottal images. As shown in FIG5C, the deep learning image recognition and segmentation software B identifies the anterior glottal gap 171 in one or more glottal images 21 respectively, and outputs an anterior glottal gap mask 21A for each glottal image 21. As shown in FIG5D, the processing unit 41 performs image processing on each anterior glottal gap mask 21A by outlining and repairing the image, so that the anterior glottal gap 171 in each anterior glottal gap mask 21A is clearly depicted. The processing unit 41 performs the above-described step of acquiring medical parameters S05, and acquires medical parameters of one or more vocal cord anatomical structures from the clearly depicted anterior glottic gaps 171 in each anterior glottic gap mask 21A. When performing the above-described step of drawing report charts S06, an anatomical structure mark is added to the position of each medical parameter in the medical image.
[0027] The output unit 42 may be an output interface, a printer, or a display. The output unit 42 is signal-connected to the processing unit 41, which receives the medical images with one or more markers and one or more medical parameters, including the standardized anterior glottic gap area or vocal cord amplitude L1 of the anterior glottic gap 171. The output unit 42 performs the above-described step of drawing the report chart S06, outputting the medical images, such as the glottic image 21 with anatomical structural markers, and the values of the medical parameters corresponding to multiple glottic images 21, as a medical parameter and image report, as shown in FIG6. The medical parameter and image report may further include the chart 30, which is a rectangular coordinate graph with the frame number as the horizontal axis and the standardized anterior glottic gap area as the vertical axis.
[0028] The above description is only a preferred embodiment of the present invention. Any equivalent changes made by applying the present invention specification and the claims should be included within the patent scope of the present invention.
Claims
1. A method for AI-assisted calculation of pharyngeal image parameters, comprising the following steps: Training the model: Using multiple laryngeal images, each with a manually selected glottal area, a deep learning object detection software is trained to extract the glottal image from the received laryngeal images; Using multiple glottal images, each with a manually labeled anterior glottal gap, a deep learning image recognition and segmentation software is trained to identify the anterior glottal gap from the received glottal images, the anterior glottal gap including the structural features of the left vocal cord process, right vocal cord process, and anterior commissure; Receiving laryngeal images: Receiving a single laryngeal image with the vocal cords in a vocalizing state, multiple laryngeal images captured frame by frame, or laryngeal video; Identifying the glottal image: Using the deep learning object detection software, the glottal image is identified from one or more of the aforementioned laryngeal images or the laryngeal... One or more glottal images are extracted from the video; the anterior glottal gap of the glottal image is identified: the anterior glottal gap in the aforementioned one or more glottal images is identified by the deep learning image recognition and segmentation software, and an anterior glottal gap mask is output for each glottal image; medical parameters are obtained: the anterior glottal gap mask is outlined and repaired by image processing, the anterior glottal gap in each anterior glottal gap mask is clearly depicted, the medical parameters of the vocal cord anatomy are obtained from the clearly depicted anterior glottal gap in each anterior glottal gap mask, and the vocal cord length is obtained from the same clearly depicted anterior glottal gap, the vocal cord length is the straight line length from the left or right vocal cord process to the anterior commissure.
2. The method for artificial intelligence-assisted calculation of pharyngeal image parameters as described in claim 1, wherein in the step of obtaining medical parameters, the obtained medical parameter is the standardized anterior glottic gap area, the anterior glottic gap area of each clearly depicted anterior glottic gap is calculated using each glottic gap, and the standardized anterior glottic gap area is calculated using the anterior glottic gap area and the vocal cord length.
3. The method for artificial intelligence-assisted calculation of pharyngeal image parameters as described in claim 1, wherein in the step of acquiring medical parameters, the acquired medical parameter is the vocal cord amplitude, which is the longest distance between the line connecting the left and right vocal cord processes and the anterior commissure when they are brought together and the lateral edge of the anterior glottic gap along a direction perpendicular to the line connecting them.
4. The method for AI-assisted calculation of pharyngeal image parameters as described in claim 2, wherein after the step of obtaining medical parameters, the step of drawing a report chart is to draw a chart with the shooting time or number of frames as the horizontal axis and the standardized anterior glottic space area as the vertical axis.
5. The method for AI-assisted calculation of pharyngeal image parameters as described in any one of claims 1 to 4, wherein in the step of identifying the anterior glottic gap of the glottic image, the glottic image is first compressed into a smaller version of the glottic image, and then the deep learning image recognition and segmentation software is used to identify the anterior glottic gap in the smaller version of the glottic image, and a smaller version of the anterior glottic gap mask is output for each smaller version of the glottic image, and the anterior glottic gap mask of each smaller version is restored to the original size version of the anterior glottic gap mask for subsequent image processing such as outlining and repair.
6. An artificial intelligence-assisted system for calculating pharyngeal image parameters, comprising an input unit, a processing unit, and an output unit, wherein: The input unit receives a medical image, including a laryngeal image with the vocal cords in a vocalizing state, multiple frame-by-frame laryngeal images, or laryngeal videos. The processing unit is signal-connected to the input unit and executes a deep learning algorithm to process the medical image received by the input unit. The deep learning algorithm includes a deep learning object detection software and a deep learning image recognition and segmentation software. The processing unit uses the deep learning object detection software to extract one or more glottal images from the aforementioned laryngeal images or laryngeal videos. The deep learning image recognition and segmentation software identifies the anterior glottic gap in each of the aforementioned glottal images. The processing unit outputs an anterior glottic gap mask corresponding to each glottic image; the processing unit performs image processing such as outlining and repair on each anterior glottic gap mask, clearly depicting the anterior glottic gap in each anterior glottic gap mask, obtaining one or more medical parameters of vocal cord anatomy from the clearly depicted anterior glottic gap in each anterior glottic gap mask, and adding an anatomical structure mark to the position of each medical parameter in the medical image; the output unit is signal-connected to the processing unit, and the processing unit receives the medical image with the one or more marks and the one or more medical parameters, and outputs the medical image and the one or more medical parameters as a medical parameter and image report.
7. The artificial intelligence-assisted calculation system for pharyngeal image parameters as described in claim 6, wherein the deep learning object detection software is trained on multiple laryngeal images, each with a manually selected glottic image range; and the deep learning image recognition and segmentation software is trained on multiple glottic images, each with a manually marked anterior glottic gap.
8. The AI-assisted calculation system for pharyngeal imaging parameters as described in claim 6, wherein the medical parameters are the standardized anterior glottic space area or vocal cord amplitude.
9. The AI-assisted calculation system for pharyngeal image parameters as described in claim 6, wherein the medical images are multiple frame-by-frame images of the larynx, the medical parameters are standardized anterior glottic space area, and the medical parameters and image report include a graph, which is a rectangular coordinate graph with the frame number as the horizontal axis and the standardized anterior glottic space area as the vertical axis.
Citation Information
Patent Citations
Bronchoscope image feature recognition system and method based on deep learning
CN112614566A
Video laryngoscope system and method for quantitatively assessment trachea
WO2022082558A1