Imaging method for visualizing a larynx on the basis of automated image analysis
Patent Information
- Application Number
- EP2023821948
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-07
- Filing Date
- 2023-12-07
- Publication Date
- 2025-08-27
AI Technical Summary
Conventional laryngeal endoscopy methods are prone to errors and lack accuracy due to complex hardware arrangements and noisy audio signals, particularly in environments where patients' verbal instructions contaminate recordings, leading to inadequate visualization of vocal fold oscillations.
An imaging method that records laryngoendoscopic images to detect the fundamental frequency of vocal fold oscillations through automated image analysis, eliminating the need for high-speed cameras and audio signals, using compressed sensing and neural networks to reconstruct oscillation behavior from sparse image data, even below the Shannon-Nyquist criterion.
This approach provides a robust, accurate, and simplified method for visualizing laryngeal structures, enabling precise determination of vocal fold oscillations and glottis opening states with high accuracy, reducing the need for specialized hardware and improving performance in noisy environments.
Smart Images

Figure IMGF000018_0001 
Figure IMGF000018_0002 
Figure IMGF000019_0001
Abstract
Description
[0001] Imaging technique for visualizing a larynx based on automated image analysis
[0002] The present invention relates to an imaging method for visualizing a larynx and / or structures associated with the larynx and a device therefor.
[0003] Laryngeal endoscopy is the most important diagnostic tool for detecting organic and functional disorders in the larynx. Laryngeal videostroboscopy is an audio-mediated imaging technique that visualizes the vibrational behavior of the vocal folds. The audio signal is used to determine the fundamental frequency (Fo)—the vibration frequency of the vocal folds. Fo is used to trigger the stroboscopic illumination unit, which provides a still image or artificial slow motion (minimally modulated FO) of the vocal fold movement.
[0004] However, this process involves a chain of complex, error-prone algorithms, noisy audio signals, and multiple hardware components that must be coordinated.
[0005] Disadvantages of conventional laryngeal endoscopy procedures and devices include the complex hardware configuration, the sensitive audio signal, and its relatively error-prone analysis. Current techniques lack accuracy and robustness in noisy environments. Due to the recording setup and the examination environment, in which physicians constantly give verbal instructions to patients, the recorded audio signals are often contaminated by the examiner.
[0006] Against this background, the technical task is to mitigate or even completely eliminate the disadvantages of the prior art. In particular, the present invention is based on the task of providing a robust, simple, and accurate method for laryngeal video stroboscopy that does not rely on high-speed cameras.
[0007] This object is achieved by the method and device having the features of the independent claims. Advantageous developments of the invention are the subject of the dependent claims.
[0008] A first aspect of the invention relates to an imaging method for visualizing a larynx and / or structures associated with the larynx, comprising the step: a) recording a number of images of a patient's larynx, in particular of the vocal cords and the glottis of the larynx, while the patient intones a sound; wherein, on the basis of the number of images, a fundamental frequency Fo of the oscillation of the vocal cords is detected by automated image analysis.
[0009] In contrast to conventional imaging methods, particularly in laryngeal endoscopy, the fundamental frequency Fo (also known as the fundamental frequency) is not recorded based on a recorded audio signal, but rather based on image data, particularly laryngoendoscopic images. The image recording rate is preferably below the Shannon-Nyquist criterion.
[0010] This does not preclude the recording of audio signals within the scope of a method according to the invention or with a device according to the invention. However, according to the invention, the recording of audio signals is generally not required for the detection of the fundamental frequency Fo of the vocal cord oscillation, since this can be detected exclusively on the basis of image data.
[0011] In other words, a focus of the present invention is on determining the fundamental frequency FO, even if the image recording rate does not meet the Shannon-Nyquist criterion, for example, because at least a conventional video stroboscopy camera and no high-speed camera is used.
[0012] A video stroboscopy camera preferably used in the context of the present invention preferably has an image recording rate of 60 to 300 Hz, in particular between 60 Hz and 250 Hz, preferably between 100 Hz and 200 Hz.
[0013] A high-speed camera, which is usually used in the context of the prior art, preferably has an image recording rate of 1,000 to 16,000 Hz. In principle, such a camera can also be used in the context of the present invention, but it is an advantage of the present invention that high-speed cameras can be dispensed with.
[0014] Within the scope of the present invention, preferably neither an audio signal nor a high-speed camera is used. In other words, the fundamental frequency F0 of the vocal fold oscillation is preferably determined by automated image analysis rather than by analyzing an audio signal.
[0015] This differs from known state-of-the-art methods, which rely on data sources that satisfy the Shannon-Nyquist criterion to calculate Fo. For example, an audio signal and / or a high-speed camera signal are usually used.
[0016] The present invention offers the advantage of determining Fo even under conditions that do not satisfy Shannon-Nyquist, such as when using a video stroboscopy camera. However, the present invention can also be used when Shannon-Nyquist is satisfied. Conversely, however, the known prior art methods cannot determine FO when Shannon-Nyquist is not present.
[0017] Within the scope of the present invention, a fundamental frequency Fo of the oscillation of the vocal cords is preferably detected based on the number of images by automated image analysis using a compressed acquisition. A control unit of a device according to the invention is therefore preferably designed to detect a fundamental frequency Fo of the oscillation of the vocal cords using a compressed acquisition based on the number of images by automated image analysis.
[0018] The number of images recorded in step c) is preferably recorded at random times and / or does not follow a predetermined time interval during a recording period. For example, during a 600 ms recording period of the vocal folds and glottis, 60 to 80 images are recorded at random times.
[0019] According to one embodiment, the fundamental frequency Fo can be determined based on the images recorded at random times, preferably in real time, using a compressed acquisition algorithm. Preferably, the image recording is influenced by the modulation of the stroboscopy unit almost simultaneously, in particular immediately, and without further delay.
[0020] According to one embodiment, a method according to the invention comprises at least one of the following steps: b) storing the number of images in a database and automatically analyzing the images with a means, preferably comprising a neural network, for predicting an oscillation state of the vocal folds and / or for predicting a degree of opening of the glottis; and / or c) creating a function by means of the means for predicting a
[0021] Oscillation state of the vocal cords and / or to predict a
[0022] Degree of opening of the glottis, where the function represents the opening area of the glottis as a function of time.
[0023] In other words, in steps b) and c), the vibration behavior or the oscillation state of the vocal cords or the degree of opening of the glottis is extracted from images of the vocal cords and glottis.
[0024] The degree of glottal opening correlates with the oscillation state of the vocal folds, as the movement of the vocal folds causes the size of the glottal opening area (also called the glottal area) to increase and decrease. For this reason, the size of the glottal area can be used to describe the oscillation behavior of the vocal folds.
[0025] The sinusoidal behavior of the glottal area during vocal cord vibration (glottis fully open, glottis closing, glottis fully closed, glottis opening, etc.) allows the opening and closing behavior of the vocal cords to be represented as a wave function. Based on the recorded images, a function is created according to step c) that indicates the glottal opening area as a function of time.
[0026] According to one embodiment of the invention, the function representing the opening area of the glottis as a function of time is used to classify one or more images showing the vocal folds and glottis in a specific position into the vibration cycle.
[0027] This allows, for example, to determine the relative degree of glottis opening for each image (e.g. on a scale where a value of 0 corresponds to a closed glottis and a value of 1 to a maximally open glottis).
[0028] To automatically determine the relative degree of glottis opening for an image, a means of automatic image analysis can be used. This means can, for example, comprise a convolutional neural network (CNN) configured or trained to automatically determine the relative degree of glottis opening in an image of a glottis.
[0029] To train the neural network, for example, a dataset containing image data from high-speed cameras capable of imaging the vibration behavior of the vocal folds with high temporal resolution can be used. For example, with such a training dataset, the temporal resolution or number of data points can be many times higher than the temporal resolution or number of data points required to determine Fo according to the Shannon-Nyquist criterion. Based on this dataset, it is possible to precisely determine the vibration cycle.
[0030] This dataset can now be used, for example, to fully supervise the neural network's training, so that it outputs the relative degree of glottis opening for each input image. Very high accuracies can be achieved in determining the degree of glottis opening on images unknown to the neural network. For example, accuracies of an average absolute error of 0.07 or 0.09 were achieved when determining the degree of glottis opening on images unknown to the neural network.
[0031] Alternatively, or in addition to a neural network, the means for automatic image analysis can also be configured to detect the relative degree of glottis opening in an image using segmentation. A neural network is therefore not mandatory.
[0032] According to one embodiment, a method according to the invention comprises the step: d) determining a fundamental frequency Fo of the oscillation of the vocal folds by means of a means for compressed sensing (also known as "compressed sensing" or "sparse sensing") on the basis of the number of images recorded in step a).
[0033] For example, during a laryngeal endoscopy, images of the vocal folds and glottis are recorded at random times. These images are automatically analyzed for the relative degree of glottis opening using, for example, the neural network trained as described above. Based on these "snapshots" of the relative degree of glottis opening (each image reflects a relative degree of glottis opening), the compressed acquisition method can reconstruct a sinusoidal function (also known as a glottal area waveform) that reflects the oscillation state of the vocal folds and thus the opening and closing behavior of the glottis.
[0034] Due to the use of the compressed acquisition algorithm, preferably a very small number of images are sufficient to reconstruct the opening and closing behavior of the glottis with sufficient accuracy. In other words, compressed acquisition preferably makes it possible to accurately capture the opening and closing behavior of the glottis even based on a very small number of snapshots. This is possible because the glottal area waveform is sparse in the frequency domain. A prerequisite for the successful use of compressed acquisition is that the underlying signal is sparse in one domain.
[0035] Compressed acquisition thus offers an advantageous algorithmic method for determining the fundamental frequency Fo, even if the available data points or snapshots do not satisfy the Shannon-Nyquist criterion. In other words, this algorithmic method allows us to reconstruct original signals with sparse data under certain assumptions. That is, with appropriate prior knowledge of the data, we can calculate what the sampled signal—in this case, the sequence of opening and closing of the glottal area—might have looked like.
[0036] The means for compressed acquisition is designed, for example, to perform a convex optimization that uses individual data points regarding the oscillation status of the vocal folds, e.g. estimated by the degree of opening of the vocal folds in randomly acquired images, to reconstruct the temporal course of the oscillation behavior of the vocal folds or the relative degree of opening of the glottis.
[0037] In other words, a continuous sine function of the oscillation state of the vocal cords over time can be constructed, preferably by means of individual data points, each of which reflects an oscillation state of the vocal cords.
[0038] The fundamental frequency Fo can then be determined from the continuous sine function of the oscillation state of the vocal cords over time, for example, using suitable methods such as a fast Fourier transform (FFT). This makes it possible to accurately record the oscillation behavior of the vocal cords and determine the fundamental frequency Fo of the oscillation behavior of the vocal cords, even with relatively few images of the vocal cords.
[0039] According to one embodiment, the means for compressed detection may be configured to use a number of data points to determine the fundamental frequency Fo, wherein each data point preferably corresponds to a degree of glottis opening identified in an image, to create a function representing an opening area of the glottis as a function of time, wherein the number of data points is less than a number of data points corresponding to the Shannon-Nyquist criterion of the function.
[0040] Once the fundamental frequency Fo has been determined, it is preferably used to control a light source with the frequency Fo, which is used, for example, in stroboscopic laryngeal endoscopy.
[0041] A method according to the invention can therefore comprise the step: e) stroboscopically controlling a light source, preferably illuminating the vocal cords and the glottis, at the frequency Fo or a multiple thereof, preferably while the patient continues to intonate the sound and / or while images of the patient's larynx, in particular of the vocal cords and the glottis of the larynx, are recorded.
[0042] An example application is the still image of vocal cord vibration. If the light source is controlled in this way with the frequency Fo, it preferably emits a flash of light whenever the glottis is in a certain, consistent state of opening or the vocal cords are in a certain, consistent state of oscillation. For example, the light source always emits a flash of light when the glottis is completely closed. This way, the imaging of the vocal cords is not disturbed by their vibration.
[0043] According to step e), more generally, a stroboscopic control of a light source preferably illuminating the vocal cords and the glottis takes place with a control frequency which was created on the basis of the frequency Fo.
[0044] The control frequency can, for example, be equal to Fo and / or a multiple of Fo (e.g., 2 x Fo; 2.5 x Fo; or even 0.5 x Fo or 0.7 x Fo). Thus, the multiple of Fo can, in principle, also be smaller than Fo.
[0045] The control frequency may also correspond to a frequency which has Fo as a subcomponent of an addition (e.g. Fo + 1 Hz) or subtraction (e.g. Fo - 1 Hz).
[0046] A further aspect of the present invention relates to a device for visualizing a larynx and / or structures associated with the larynx. All features, advantages, and effects disclosed above in the context of a method according to the invention are also applicable to a device according to the invention, even if they are not explicitly listed to avoid redundancies, and vice versa. In other words, a control unit of a device according to the invention can be configured to execute one or more arbitrary method steps of a method according to the invention.
[0047] Accordingly, a device for visualising a larynx and / or structures associated with the larynx is provided, comprising: a means for recording a number of images of a patient's larynx, in particular of the vocal cords and the glottis of the larynx, and - a control unit which is programmed to detect a fundamental frequency Fo of the oscillation of the vocal cords on the basis of the number of images by means of an automated image analysis.
[0048] A device according to the invention preferably comprises a video stroboscopy camera with an image recording rate of 60 to 300 Hz as a means for recording a number of images. Preferably, no high-speed camera is present.
[0049] Preferably, the control unit is programmed to determine the fundamental frequency Fo of the oscillation of the vocal folds, preferably by automated image analysis and not by analysis of an audio signal.
[0050] Within the scope of the present invention, the determination of the fundamental frequency Fo of the oscillation of the vocal folds can be carried out exclusively by automated image analysis.
[0051] According to one embodiment, a device according to the invention has a database for storing the number of images, and the control unit is programmed to load the images from the database and to subject the images to an automated image analysis by means, preferably comprising a neural network, for predicting an oscillation state of the vocal cords and for predicting a degree of opening of the glottis.
[0052] Alternatively or additionally, the control unit can be programmed to create a function using the means for predicting a state of oscillation of the vocal folds and for predicting a degree of glottis opening, preferably based on the images loaded from the database, wherein the function represents an opening area (also known as a degree of opening) of the glottis as a function of time. Furthermore, the control unit can be programmed to detect a fundamental frequency Fo of the oscillation of the vocal folds using a means for compressed detection based on the number of images recorded by the means for recording a number of images. The number of images was preferably recorded at random times.
[0053] The control unit may therefore be programmed to control the means for recording a number of images to record one image at a time at random times within a recording time period.
[0054] According to one embodiment of the invention, the control unit is connected to a light source and programmed to control the light source to emit stroboscopic light flashes at the frequency Fo or a multiple thereof. Preferably, the control unit is further configured to simultaneously control the means for recording a number of images of a patient's larynx to record images, so that preferably the light source and the means for recording a number of images are controlled or switched at the same frequency, preferably the frequency Fo or a multiple thereof or a common control frequency. The light source can in principle also be controlled to emit light flashes at random times.
[0055] The control unit may also be programmed to control the light source to emit stroboscopic flashes of light at a control frequency created based on the frequency FO.
[0056] The control unit can comprise a means for compressed acquisition configured or programmed to use a number of data points to determine the fundamental frequency Fo, wherein each data point preferably corresponds to a degree of glottal opening identified in an image, in order to create a function that represents an opening area of the glottis as a function of time, wherein the number of data points can preferably be less than a number of data points corresponding to the Shannon-Nyquist criterion of the function. The degree of glottal opening can be determined by a neural network that acquires an opening area of the glottis, and / or the degree of glottal opening can be acquired by segmentation. In principle, the number of data points can also reach or exceed a number of data points corresponding to the Shannon-Nyquist criterion of the function.
[0057] It should be noted that, within the context of the present disclosure, the articles "a" and "an" are not to be interpreted as "exactly one." Thus, if a particular element in the present disclosure is disclosed in the singular, the disclosure is not limited to this, but also encompasses embodiments with this element in the plural, and vice versa.
[0058] Furthermore, it will be obvious to those skilled in the art that the present disclosure is not limited to the explicitly mentioned feature combinations. Rather, all features disclosed in the present disclosure can be singled out and claimed in any combination or even in isolation.
[0059] Further advantages, features, and effects of the present invention will become apparent from the following description of preferred embodiments of the invention with reference to the accompanying figures. Herein:
[0060] Fig. 1 is an overview illustrating the differences between a method according to the invention and a known method of laryngeal endoscopy. Fig. 2 shows exemplary images used within the scope of the present invention to determine the fundamental frequency (upper panel) and exemplary images recorded within the scope of a method according to the invention using stroboscopic control of a light source with the fundamental frequency Fo.
[0061] Fig. 3 is an overview of how a neural network used in the present invention can be trained.
[0062] As shown in the upper box in Fig. 1, audio and video data are recorded during a conventional laryngeal endoscopic examination.
[0063] The audio data, which includes the subject's signal (e.g., while the subject intones the vowel "e") and other signal sources, such as a doctor's voice and background noise, are conventionally used to determine the fundamental frequency (FO) of the vocal cord vibration behavior. A strobe unit or light source that illuminates the patient's larynx or vocal cords is then controlled with the recorded fundamental frequency (FO) or a multiple thereof.
[0064] A method according to the invention is shown in the lower box in Fig. 1. Audio and / or video data can also be recorded during a laryngeal endoscopic examination using a method according to the invention. Video data preferably includes images or image data.
[0065] As shown in Fig. 1, the endoscopic image shown as an example in the figure, which includes the vocal folds and the glottis, is used as input to a convolutional neural network (CNN) that estimates or detects the oscillation state of the vocal folds in the image, i.e., the relative opening state of the glottis. The glottal area (“glottal area” in Fig. 1) is used as a measure of the relative opening state of the glottis.
[0066] Based on a number of such endoscopic images, a number of relative opening states of the glottis can be recorded, which are in temporal context with each other.
[0067] Example images that sample relative glottal opening states are shown, for example, in the upper panel of Fig. 2.
[0068] Such images are preferably recorded within a short recording period, which may be, for example, between 20 ms and 600 ms, preferably between 50 ms and 200 ms, in particular between 50 ms and 100 ms, particularly preferably between 50 ms and 80 ms.
[0069] The number of images recorded and used as samples is at least 2, preferably at least 20. The number of images recorded and used as samples can be, for example, between 2 and 100 images, between 50 and 100, preferably between 50 and 80, and especially between 50 and 70 images. This absolute number depends on the hardware and algorithm used. New developments may allow for smaller samples.
[0070] It should be emphasized that one advantage of the present invention is that the images used can preferably be recorded, for example, using a conventional camera system of a conventional endoscopy device. The present invention therefore preferably does not require specialized hardware equipment, for example in the form of high-speed cameras, such as those used in videokymography or high-speed video endoscopy. According to the invention, high-speed cameras can therefore preferably be dispensed with, and a device according to the invention cannot have a high-speed camera, such as those commonly used in videokymography.
[0071] A compressed sensing algorithm uses acquired samples of the relative glottal opening state over time to reconstruct the time course of the oscillatory state of the vocal folds. The degree of glottal opening can be determined by a neural network (e.g., CNN) that captures the glottal opening area, and / or the degree of glottal opening can be captured by segmentation.
[0072] The time course of the oscillation state of the vocal folds is shown as a sinusoidal curve in Fig. 1. The time course of the relative opening state of the glottis corresponds to the time course of the oscillation state of the vocal folds.
[0073] The use of a compressed acquisition algorithm makes it possible to reliably reconstruct the time course of the oscillation state of the vocal folds from the “samples” even on the basis of only a few endoscopic images.
[0074] In particular, the use of a compressed acquisition algorithm makes it possible to determine the time course with sufficient accuracy even when fewer endoscopic images are available than appear necessary according to the Shannon-Nyquist criterion for describing the time course function and calculating the fundamental frequency Fo.
[0075] The Shannon-Nyquist sampling theorem states that for perfect signal reconstruction, a signal must be sampled at least twice its highest frequency (Equation 1). By sampling a continuous signal, it is converted into a sequence of discrete values. Based on these discrete values, the continuous signal can then be reconstructed. Equation 1
[0076] A disadvantage of applying this theorem is that the Nyquist sampling rate for many applications, such as predicting the fundamental frequency Fo in phonation, can be very high, technically demanding, and therefore expensive (requiring hardware such as high-speed cameras). Therefore, alternative methods to avoid Nyquist sampling have been sought. One of these is compressed acquisition or sampling. Its goal is to reconstruct a sparse signal from a few random, non-adaptive, linear measurements, preserving the signal structure through convex optimization. This can be more efficient than sampling according to the Shannon-Nyquist theorem.
[0077] Mathematically, compressed sensing exploits the sparsity of the signal in a domain to achieve a complete reconstruction of a signal from only a few measurements. The signal is preferably continuously sampled, resulting in its compressed version. Compressed sensing requires determining the sparsest vector s consistent with the measurements y (Equation 2).
[0078] Equation 2
[0079] In equation 2, y corresponds to a measurement vector that represents a compressed version of the original signal f. Since only a few samples are to be selected from the entire signal, the randomly generated sampling matrix C determines which samples are selected from all available sample points. The original signal f can be represented by * s, where MJ corresponds to the universal transformation basis, which can be, for example, a discrete cosine transform matrix known to the compressed sensing algorithm. 0 represents a diagonal matrix that is computed by 0 = 0 * ^. The vector s is unknown in this system of equations. The system of equations is underdetermined because there are infinitely many consistent solutions for "'s . . Minimizing the L1 norm yields the sparsest solution "s for s that satisfies the optimization problem. The convex optimization problem can be solved using linear algebra.
[0080] Equation 3
[0081] When using compressed acquisition, very low sampling rates are sufficient to reconstruct the entire signal, such as in the present case for the prediction of the fundamental frequency Fo. With this method, it is possible to reconstruct a temporally high-dimensional signal from a sparse set of random measurements
[0082] After reconstructing the time course of the relative opening state of the glottis by means of a compressed acquisition, as shown in Fig. 1, the fundamental frequency Fo is determined on the basis of the reconstructed time course, for example by a preferably fast Fourier transform (FFT).
[0083] A strobe unit is then controlled to emit light at the detected fundamental frequency Fo. This offers the advantage that the vocal cords are always illuminated in the same oscillation state (e.g., completely closed). Example images that arise when the strobe unit is controlled with the fundamental frequency Fo are shown, for example, in the lower panel of Fig. 2. In addition, a camera unit is preferably controlled to record an image at the detected fundamental frequency Fo. In this case, the camera always records an image when the strobe unit illuminates the vocal cords. Alternatively, the camera can also record continuously, for example, a video. Control with, for example, a multiple of Fo would also be conceivable.
[0084] Calculating the relative degree of vocal cord opening requires prior knowledge of the maximum extent of vocal cord opening and closing. High-speed recordings allow for very accurate estimation, as each opening and closing cycle is imaged with high temporal resolution. The glottal area (glottis) is preferably used as an approximation of vocal cord vibration. Detecting the glottis in an endoscopic image is a complex process. Deep neural networks, such as encoder-decoder CNNs, are preferably used for this purpose.
[0085] Fig. 3 illustrates the training of a neural network that can be used in the context of the present invention. For example, a set of endoscopic images (e.g., the BAGLS dataset) depicting the glottis is used as training data.
[0086] After the neural network is trained, each endoscopic input image is semantically segmented to capture the glottis opening area, as shown in Fig. 3. For each individual image, the pixels identified as belonging to the glottis opening area are summed, with a larger number of pixels indicating a more open glottis.
[0087] Based on each segmented endoscopic image, a glottis opening state can be determined, and thus a data point on the time course to be reconstructed. Using high-speed videos, the GAW can be constructed very accurately, as shown in Fig. 3. Using suitable methods, such as a fast Fourier transform (FFT), the fundamental frequency Fo can be calculated.
[0088] The time course of the GAW preferably has the general form according to equation 4:
[0089] Equation 4
[0090] The function GAW thus represents the sum of all segmented pixels at positions (i, j) in a given image with intensity I (1 for the glottis, 0 for the background) at time t.
[0091] The respective measured or calculated glottis opening area is normalized to its maximum opening, so that the opening area varies, for example, between 0 and 1. Therefore, in each endoscopic image, the oscillation state (essentially equivalent to the glottis opening area) is preferably predicted in the range from 0 to 1, where 0 represents completely closed and 1 represents completely open.
[0092] Within the scope of the invention, randomly acquired images can be used to predict the relative aperture angle, and thus subranges of the normalized GAW. Using the compressed acquisition algorithm, the GAW can be reconstructed in such a way that the fundamental frequency Fo can subsequently be determined.
[0093] Within the scope of the present invention, at least one of the following two neural networks can be used: 1. A first neural network designed to determine a time course of the oscillation state of the vocal cords (e.g., GAW function) based on a number of images. However, the GAW function can also be generated using other image processing methods.
[0094] Preferably, the relative degree of opening of the glottis is first determined (e.g. 0 closed, 1 open). If only individual images of the glottal area are viewed, it cannot be directly determined whether a recorded opening state corresponds to the maximum opening state or whether the glottis is possibly opening further. Therefore, the first neural network is preferably used to determine a minimum and a maximum opening state of the glottis based on recorded images of the glottis and thus to determine the time course (GAW). High-speed videos of the glottis can be analyzed here. In other words, a training data set for the first neural network comprises at least one image or video that was captured using a high-speed camera.
[0095] However, within the scope of the present invention, it is possible to dispense with high-speed recordings of the glottis during the examination.
[0096] 2. a second neural network designed to detect an oscillation state of the vocal cords shown in an image and to classify this in the time course of the oscillation state of the vocal cords (GAW function).
[0097] In other words, the second neural network is preferably configured to predict or determine the relative degree of opening (a number between 0 and 1 corresponding to the GAW function) of the glottis in this image based on an analysis of this image. For example, the second neural network detects a relative degree of opening of 0.75 in an image. The images accessed by the second neural network can originate from a high-speed camera. However, this is not required; the images can also originate from a conventional full-HD camera, for example, from a conventional endoscopy system.
[0098] Since conventional endoscopy systems usually do not have high-speed cameras, this significantly increases the versatility of the present invention's applicability in everyday clinical practice.
Claims
Patent claims Imaging method for visualizing a larynx and / or structures associated with the larynx, comprising the step: a) recording a number of images of a patient's larynx, in particular of the vocal cords and the glottis of the larynx, while the patient intones a sound; characterized in that, on the basis of the number of images, a fundamental frequency Fo of the oscillation of the vocal cords below the Shannon-Nyquist criterion is detected by automated image analysis.Imaging method according to claim 1, further comprising the steps of: b) storing the number of images in a database and automatically analyzing the images with a means, preferably comprising a neural network, for predicting an oscillation state of the vocal cords and for predicting a degree of opening of the glottis; and / or c) creating a function by means of the means for predicting an oscillation state of the vocal cords and for predicting a degree of opening. of the glottis, wherein the function represents an opening area of the glottis as a function of time. Imaging method according to claim 1 or 2, further comprising the step: d) determining a fundamental frequency Fo of the oscillation of the vocal folds by means of a means for compressed acquisition on the basis of the number of images recorded in step a). Imaging method according to one of the preceding claims, further comprising the step e) stroboscopically controlling a light source, preferably illuminating the vocal folds and the glottis, at the frequency Fo or a multiple thereof, preferably while the patient continues to intonate the sound and / or while images of the patient's larynx, in particular of the vocal folds and the glottis of the larynx, are being recorded.Imaging method according to one of the preceding claims, characterized in that the means for compressed acquisition is configured to use a number of data points to determine the fundamental frequency FO, wherein each data point preferably corresponds to a degree of opening of the glottis identified in an image, in order to create a function that represents an opening area of the glottis as a function of time, wherein the number of data points is or can be smaller than a number of data points corresponding to the Shannon-Nyquist criterion of the function. Imaging method according to one of the preceding claims, characterized in that the images recorded in step a) are recorded at random time points. A device for visualizing a larynx and / or structures associated with the larynx, comprising: - a means for recording a number of images of a patient's larynx, in particular of the vocal folds and the glottis of the larynx, and - a control unit which is programmed to, on the basis of the A number of images are used to detect a fundamental frequency Fo of the oscillation of the vocal cords below the Shannon-Nyquist criterion through automated image analysis. The device according to claim 7, further comprising: a database for storing the number of images, wherein the control unit is programmed to load the images from the database, to subject the images to automated image analysis by means, preferably comprising a neural network, for predicting an oscillation state of the vocal cords and for predicting a degree of opening of the glottis; and / or the control unit is programmed to create a function by means of the means for predicting an oscillation state of the vocal cords and for predicting a degree of opening of the glottis, wherein the function represents an opening area of the glottis as a function of time. The device according to claim 7 or 8, wherein the control unit is further programmed to detect a fundamental frequency Fo of the oscillation of the vocal folds by means of a compressed detection means based on the number of images recorded by the means for recording a number of images. The device according to any one of claims 7 to 9, wherein the control unit is further connected to a light source and programmed to drive the light source to emit stroboscopic light flashes at the frequency Fo or a multiple thereof, and preferably simultaneously drive the means for recording a number of images of a patient's larynx to record images.The device according to any one of claims 7 to 10, wherein the control unit further comprises a means for compressed acquisition configured or programmed to use a number of data points to determine the fundamental frequency FO, each data point preferably corresponding to a degree of glottis opening identified in an image, to create a function representing an opening area of the glottis as a function of time, the number of data points being or being able to be less than a number of data points corresponding to the Shannon-Nyquist criterion of the function. The device according to any one of claims 7 to 10, characterized in that the control unit is further programmed to control the means for recording a number of images of a patient's larynx to record images at random times.