A diagnostic aid method for laryngeal paralysis combined with audio and video processing

Through combined audio and video processing technology, effective vocal fragments in laryngoscopic videos are extracted and the vocal cord opening and closing angles are automatically marked, which solves the problem that existing diagnostic methods rely on subjective judgment and manual screening, and achieves a more efficient and accurate diagnosis of laryngo paralysis.

CN118743534BActive Publication Date: 2025-06-06DUKE KUNSHAN UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410849050.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-27
Publication Date
2025-06-06
Estimated Expiration
2044-06-27

AI Technical Summary

Technical Problem

The existing diagnosis methods for laryngeal paralysis rely on doctors’ subjective judgment of laryngoscopic videos, lack of objective indicators and data support, resulting in low credibility in misjudgment, misjudgment and diagnosis. At the same time, manual screening of video clips is time-consuming and labor-intensive, reducing diagnostic efficiency.

Method used

A method for diagnosis of laryngeal paralysis combined with audio and video processing is proposed. Effective vocal fragments are extracted through keyword recognition models, combined with glottic segmentation and recognition technology, and automatically label and calculate the vocal cord opening and closing angles to provide objective evaluation indicators of multiple angles.

Benefits of technology

It improves the efficiency and quality of laryngeal paralysis diagnosis, reduces the time for manual screening by doctors, provides more reliable diagnostic results, and enhances the objectivity and accuracy of the diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118743534B_ABST
    Figure CN118743534B_ABST
Patent Text Reader

Abstract

The present invention proposes a laryngeal paralysis diagnosis auxiliary method combined with audio and video processing, comprising the following steps: obtaining an original laryngoscope video with audio information and video information; predicting the audio information through a keyword recognition model to extract an effective phonation segment, and obtaining a first laryngoscope segment corresponding to the effective phonation segment from the original laryngoscope video segment according to the matching relationship between the audio information and the video information; detecting and identifying the glottis area in the first laryngoscope segment, and segmenting the glottis area from the first laryngoscope segment; marking the glottis area and calculating the physical properties of the glottis area, the physical properties including one or more of the glottis area area, vocal cord opening and closing angle, and vocal cord width. The present invention uses a multimodal analysis method combined with audio and video to segment the laryngoscope video, extract key video segments, and provide a variety of objective evaluation indicators for doctors to refer to, thereby improving diagnostic efficiency and quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of medical software, and in particular to a laryngeal paralysis diagnosis auxiliary method combining audio and video processing. Background Art

[0002] Laryngeal endoscopy (laryngoscopy) is a clinical method mainly used to diagnose vocal cord lesions. The doctor places a camera and a microphone close to the patient's vocal cords, and records the video and audio information of the vocal cords during the patient's vocalization by guiding the patient to speak. Since the vibration frequency of the vocal cords during the vocalization process is about 80 to 2000 times per second, such high-frequency vibration makes it difficult for doctors to observe the vibration of the vocal cords in detail. Therefore, many doctors now use stroboscopic laryngoscopes during examinations to better observe the high-speed vibrating vocal cords. The principle of stroboscopic laryngoscopes is to use a fixed flash frequency that is slightly different from the vibration frequency of the vocal cords, and use the 0.2-second visual afterglow of the human eye (Talbot's law) to cause an optical illusion and slow down the high-speed vibrating vocal cords.

[0003] The laryngoscopy video not only contains the patient's vocalization clips with the vocal cords and glottis of the patient, but also contains many useless clips, such as the clips in which the vocal cords and glottis were not photographed at the beginning of the laryngoscopy video recording, and the clips in which the patient did not make any sound. Therefore, in the process of analyzing the laryngoscopy video, doctors often need to manually select clear patient vocalization clips from a lengthy laryngoscopy video for further analysis, and this manual selection process requires a lot of time investment, which has a negative impact on the doctor's diagnostic efficiency. The second is the problem of diagnostic credibility. The diagnosis of laryngeal paralysis often depends on the doctor's subjective judgment of the laryngoscopy video, and the patient's paralysis is evaluated by observing the vibration law of the vocal cords and glottis when the patient speaks. Such a diagnostic method that relies on subjective evaluation lacks the support of some objective indicators and data. On the one hand, it may cause misjudgment and missed judgment, and on the other hand, it leads to low credibility of the patient in the doctor's judgment results.

[0004] Existing intelligent auxiliary methods for laryngeal paralysis diagnosis focus on image processing and do not use information about the patient's voice. These methods use deep learning methods to automatically capture the position of the glottis and vocal cords in the image, calculate the glottis area at each moment, and draw the glottis area curve (GAW). Although these intelligent auxiliary methods propose an objective evaluation index for doctors to refer to, there are still many shortcomings. First, the glottis area index alone is too single, and the vocal cord opening and closing angle, vocal cord width, etc. can all be used as references for diagnosing paralysis. Secondly, these methods are limited by image processing and do not use information about the patient's voice, so they cannot comprehensively and multi-angle evaluate the patient's laryngeal paralysis. Summary of the invention

[0005] In view of the shortcomings of the prior art, the present invention proposes a laryngeal paralysis diagnosis auxiliary method combined with audio and video processing, comprising:

[0006] Obtaining original laryngoscope video with audio information and video information;

[0007] The audio information is predicted by a keyword recognition model to extract a valid sound segment, and a first laryngoscope segment corresponding to the valid sound segment is obtained from the original laryngoscope video segmentation according to a matching relationship between the audio information and the video information;

[0008] Detecting and identifying the glottis region in the first laryngoscope segment, and segmenting the glottis region from the first laryngoscope segment;

[0009] Mark the glottis area and calculate the center point of the line connecting the glottis apex and make the midline of the glottis, draw a perpendicular line of N equally divided points on the midline, and record the intersection of the perpendicular line and the glottis boundary;

[0010] The angle changes of the left vocal cord opening and closing angle and the right vocal cord opening and closing angle formed by the line connecting the intersection and the glottis apex and the line connecting the glottis apex and the center point along the time series are calculated.

[0011] In a further embodiment, the method of predicting the audio information by a keyword recognition model to extract a valid sound segment, and obtaining a first laryngoscope segment from the original laryngoscope video segment according to the valid sound segment, comprises the following steps:

[0012] Extracting audio information of the original laryngoscope video by an audio processing tool, and converting the audio information into a Mel spectrum form;

[0013] Constructing a keyword recognition model, dividing the audio information in the form of Mel spectrum into several segments and inputting the segments into the keyword recognition model for training;

[0014] The segmented audio segments are input into a trained keyword recognition model to output the utterance probability of the segments, and valid utterance segments are searched according to the utterance probability and the first laryngoscope segment corresponding to the valid utterance segment is extracted.

[0015] In a further embodiment, the detecting and identifying the glottis region in the first laryngoscope segment, and segmenting the glottis region from the first laryngoscope segment, comprises the following steps:

[0016] Recognize the first laryngoscope segment frame by frame through the vocal cord recognition model and mark the confidence, sort the first laryngoscope segment according to the confidence to obtain the first K segments as the third laryngoscope segment;

[0017] A glottis region mask is generated by a glottis segmentation algorithm and the glottis region is segmented from the glottis image;

[0018] Physical properties of the glottis are calculated based on the glottis region mask.

[0019] In one embodiment, extracting the audio information of the original laryngoscope video by an audio processing tool and converting the audio information into a Mel spectrum format comprises the following steps:

[0020] The audio information is sampled 16,000 times per second and a sampling result is generated;

[0021] A 0.025 second window is taken along the time axis to perform spectrum analysis on the sampling results, and the number of windows for the spectrum analysis is set to 512;

[0022] The Mel spectrum of the audio information is calculated through 80 Mel filters.

[0023] In one embodiment, the keyword recognition model includes a two-dimensional convolutional layer, a maximum pooling layer, a residual module, an average pooling layer and two fully connected layers connected in sequence; after the audio clip is input into the keyword recognition model, the keyword recognition model outputs a two-dimensional probability vector, and by comparing the probability vectors, the utterance probability of the audio clip is inferred.

[0024] In one embodiment, when detecting and identifying the glottis area in the first laryngoscope segment and segmenting the glottis area from the first laryngoscope segment, the following steps are also included:

[0025] Performing image brightness analysis on the first laryngoscope segment, and intercepting the segment whose brightness changes more than a threshold value per unit time as a stroboscopic segment;

[0026] The stroboscopic segments are extracted and sequentially spliced ​​to form a second laryngoscope segment.

[0027] In one embodiment, the vocal cord recognition model is obtained by inputting a vocal cord recognition dataset into a YOLO-v5 model, and the construction of the vocal cord recognition dataset includes the following steps:

[0028] Obtain a glottis segmentation dataset from a public database and construct an image coordinate system in the glottis segmentation dataset;

[0029] Get the vertex coordinates of the vocal cords through the glottal segmentation label mask;

[0030] The vertex coordinates are expanded and connected with rectangles to obtain target detection labels, and the target detection labels are summarized as a vocal cord recognition dataset.

[0031] In one embodiment, the glottis segmentation algorithm is constructed by a diffusion model, comprising the following steps:

[0032] Obtain a glottal segmentation dataset from a public database, and randomly select sample images and corresponding segmentation labels from it;

[0033] Input the sampled image into two residual convolutional layers to obtain the original features;

[0034] The noisy sampled image after t+1 iterations is added to the time embedding vector obtained at time t+1, and the input is downsampled in the residual convolution layer to obtain the potential representation;

[0035] Through the self-attention mechanism, the latent representation is fused with the original features to obtain the combined features;

[0036] The combined features are added to the time embedding vector at time t+1 and input into the residual convolution layer for upsampling to obtain the denoised label image at time t+1.

[0037] In one embodiment, the step of identifying the first laryngoscope segment frame by frame through the vocal cord recognition model and marking the confidence, and sorting the first laryngoscope segment according to the confidence to obtain the first K segments as the third laryngoscope segment, comprises the following steps:

[0038] A sliding window of fixed length t is created along the time axis, the second laryngoscope segment is traversed by window movement, and the confidence mean of each window segment is calculated;

[0039] Generate a list for storing window fragments and preset a variable, wherein the variable is the window fragment with the smallest confidence in the list;

[0040] During the traversal process, if the confidence mean of the window fragment is greater than the variable, the window fragment with the smallest confidence mean in the list is deleted and the variable is updated; several window fragments with the largest confidence mean are selected and saved.

[0041] In one embodiment, the acquisition of the vocal cord opening and closing angle comprises the following steps:

[0042] The top coordinate U, bottom coordinate D, left coordinate R and right coordinate L of the glottis are obtained by the glottis mask, and the center point C thereof is calculated;

[0043] Establish a line CD passing through the center point C and the bottom coordinate D, calculate its function f(x) and find N equidistant points C on the function N-1 ;

[0044] Calculate the function f′ at equidistant points perpendicular to f(x) k (x) and calculate its intersection point L with the glottal mask k and R k ;

[0045] Rotate the coordinate axis so that f(x) is parallel to the y-axis and calculate the new coordinate L λ,k and R λ,k , the quadratic curve q is obtained by least squares fitting γ (x) and its lowest point D q ;

[0046] Establish the calibrated center line f″(x) connecting the center point C and the lowest point D q , the midline f″(x) is at the intersection point D ′ Intersecting with the glottal mask, by calculating L k and R k The distance from f″(x) and L k and R k The ratio of the distance to D′ gives the cosine value of the left glottis angle and the right glottis angle.

[0047] The embodiment of the present invention inputs the original laryngoscope video into the keyword recognition model for speech processing, obtains the first laryngoscope segment in which the patient pronounces the "e" sound in the original laryngoscope video, performs preliminary screening, and then performs multiple video processing on the first laryngoscope segment, further purifies the key laryngoscope segment therein, and further performs automatic annotation and calculation on it, so as to provide doctors with objective evaluation indicators from multiple angles. The present invention utilizes a multimodal analysis method combined with audio and video, applies multiple technologies such as speech keyword recognition, image segmentation, and image recognition, divides the laryngoscope video, extracts key video segments, and provides objective evaluation indicators from multiple angles for doctors to refer to, thereby improving diagnostic efficiency and diagnostic quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0049] Figure 1 A flowchart of a laryngeal paralysis diagnosis auxiliary method combining audio and video processing according to the present invention;

[0050] Figure 2 A schematic diagram of the architecture of a laryngeal paralysis diagnosis auxiliary method combining audio and video processing according to the present invention;

[0051] Figure 3 is a detailed flow chart of step S200 in an embodiment of the present invention;

[0052] Figure 4is a detailed flow chart of step S300 in an embodiment of the present invention;

[0053] Figure 5 Schematic diagram of the timing of image brightness analysis in an embodiment of the present invention;

[0054] Figure 6 A schematic diagram of establishing an image coordinate system when constructing a vocal cord recognition data set in an embodiment of the present invention;

[0055] Figure 7 A schematic diagram of the structure of a glottal segmentation algorithm in an embodiment of the present invention;

[0056] Figure 8 A schematic diagram of the measurement of calculating the vocal cord opening and closing angle in an embodiment of the present invention;

[0057] Fig. 9 A schematic diagram of the timing of the change in the opening and closing angle of the vocal cords in an embodiment of the present invention;

[0058] Fig.10 Schematic diagram of the steps for calculating the glottal angle in an embodiment of the present invention. DETAILED DESCRIPTION

[0059] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments, and the well-known modules, units and their connections, links, communications or operations are not shown or described in detail. In addition, the described features, architectures or functions can be combined in any way in one or more embodiments. It should be understood by those skilled in the art that the various embodiments described below are only for illustration and not for limiting the scope of protection of the present invention. It can also be easily understood that the modules or units or processing methods in the various embodiments described herein and shown in the drawings can be combined and designed according to various different configurations. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0060] In actual laryngoscopy diagnosis, doctors mainly analyze and diagnose by observing the laryngoscopy video data of the patient's phonation cycle. The video of the patient's phonation cycle must meet two requirements: 1) The video must contain a complete segment from the beginning to the end of the patient's phonation; 2) The vocal cord glottis area in the video must be clearly visible. However, there are a large number of video clips in the existing stroboscopic laryngoscope that do not meet the above conditions. Manual screening by doctors will greatly reduce the efficiency of diagnosis.

[0061] Please refer to Figures 1 to 10 As shown, the embodiment of the present invention discloses a laryngeal paralysis diagnosis auxiliary method combining audio and video processing, comprising the following steps:

[0062] 100 obtaining an original laryngoscope video with audio information and video information;

[0063] 200 Perform brightness analysis on the original laryngoscope video to detect and filter out the stroboscopic segments in the original laryngoscope video

[0064] 300 predicting the audio information through a keyword recognition model to extract a valid sound segment, and obtaining a first laryngoscope segment corresponding to the valid sound segment from the original laryngoscope video segmentation according to a matching relationship between the audio information and the video information;

[0065] 400 Detecting and identifying the glottis region in the first laryngoscope segment through video information, and segmenting the glottis region from the first laryngoscope segment;

[0066] 500: marking the glottis area and calculating the center point of the line connecting the glottis apex and making the midline of the glottis, making a perpendicular line of N equally divided points on the midline, and recording the intersection of the perpendicular line and the glottis boundary;

[0067] 600 calculates the angle change of the left vocal cord opening and closing angle and the right vocal cord opening and closing angle formed by the line connecting the intersection and the glottis apex and the line connecting the glottis apex and the center point along the time series.

[0068] The present invention analyzes the audio part of the original laryngoscope video to preliminarily screen out the effective first laryngoscope segment, analyzes the video part of the first laryngoscope segment, and further screens out a clear and effective third laryngoscope segment for doctors to perform laryngeal paralysis diagnosis and analysis. In addition, the second laryngoscope segment is obtained by extracting the stroboscopic segment in the original laryngoscope video, and the doctor does not need to manually screen, thereby achieving the purpose of improving the doctor's diagnostic efficiency.

[0069] The following are detailed descriptions:

[0070] like Figure 4 As shown, the present invention can be mainly divided into two parts: a speech signal processing module and a video processing module. The original laryngoscope video is processed by the speech signal processing module to obtain an effective sound segment of the patient's "e" sound in the video (in clinical practice, doctors usually ask patients to make an "e" sound. When making an "e" sound, the vocal cords of the larynx will move closer to the middle, the glottis will become smaller, and the glottis can be better exposed). The video corresponding to the patient's sounding part is further purified through the subsequent video processing module through the effective sound segment, thereby obtaining the first laryngoscope segment for the doctor to view. For the first laryngoscope segment, the video processing module further performs automatic annotation and calculation on it to provide doctors with objective evaluation indicators from multiple angles.

[0071] Furthermore, if Figure 3 As shown in the figure, the purpose of processing the audio of the original laryngoscope video is to extract the part of the entire audio in which the patient makes the "e" sound, and perform frame-level processing on it, so as to sample and analyze each frame of the original laryngoscope video, ensure that the probability of the patient's utterance in each frame is captured, and make the final analysis result fully show the glottis situation at each moment when the patient utters the sound, which specifically includes the following steps:

[0072] 201 extract the audio information of the original laryngoscope video through an audio processing tool, and convert the audio information into a Mel spectrum form; in order to analyze and process the voice signal in the video alone, the present invention uses the ffmpeg toolkit to extract the audio signal from the original laryngoscope video. In order to obtain a voice feature with better representation ability, the present application converts the one-dimensional audio information into a Mel spectrum, and the specific feature extraction method is as follows: the audio signal is sampled 16,000 times per second, and the sampling result is analyzed on the time axis using a window with a length of about 0.025 seconds. The jump length between windows is set to 64 sampling points (about 0.01 seconds of audio) to capture the short-term characteristics in the audio signal. When performing spectrum analysis, the Fourier transform window is set to 512, that is, the signal is divided into 512 different frequency bands to better understand the spectrum characteristics. Finally, 80 Mel filters are used to calculate the Mel spectrum. By converting the Mel spectrum, the information can be effectively reduced while retaining features with better representation ability.

[0073] 202 constructing a keyword recognition model, dividing the audio information in the form of Mel spectrum into a plurality of segments and inputting the segments into the keyword recognition model for training;

[0074] It is understandable that keyword spotting technology is intended to detect and identify whether there are predefined keywords or phrases in a piece of audio. The present invention uses keyword spotting technology to detect and identify whether there is an "e" sound made by the patient in the audio. In order to perform frame-level audio analysis, the Mel spectrum obtained after the above-mentioned feature extraction process is further divided into several segments along the time axis as training data, where each segment includes 60 frames of spectrum information. Each training data with a dimension of 1x60x80 is subjected to preliminary feature extraction through a two-dimensional convolutional layer and a maximum pooling layer, and the number of channels is increased to obtain 32x60x80 features. Subsequently, this feature is passed through four residual modules (each module consists of two residual networks and a maximum pooling layer) to obtain a feature with a dimension of 64x15x20. The purpose of increasing the number of channels and reducing the feature map is to further extract features in detail. Finally, the 64x15x20 features are flattened and sent to two fully connected layers after an average pooling layer to obtain the final 2D probability vector. One dimension represents the probability of recognizing the patient's voice in the segment, and the other dimension represents the probability of not recognizing the patient's voice in the segment. By comparing the probability vectors, the probability of the voice in the audio segment is inferred. The specific keyword recognition model architecture is shown in the following table:

[0075]

[0076]

[0077] 203: Input the segmented audio segments into the trained keyword recognition model to output the utterance probability of the segments, search for valid utterance segments according to the utterance probability, and extract the first laryngoscope segment corresponding to the valid utterance segment.

[0078] It is understandable that after training, the above model can analyze the patient's voice fragments and obtain the probability of whether the patient speaks in the frame-level audio information. After the test audio is converted into Mel spectrum in step 201 and framed and segmented, each fragment will be input into the trained model, and the keyword recognition model will infer the probability of the patient speaking or not speaking in the corresponding fragment, and then judge whether the patient in the fragment speaks by comparing the probabilities of the two. Since the keyword recognition model can only analyze the speech of the fragment, it is also necessary to splice the results of all fragments, filter out short fragments, and retain long fragments.

[0079] Furthermore, if Figure 4As shown, since the camera may not be able to clearly capture the vocal cord area during the patient's vocalization, the first laryngoscope segment cannot fully meet the doctor's diagnostic needs. For the convenience of the doctor's reference, the embodiment of the present application further selects the second laryngoscope segment and the third laryngoscope segment from the first laryngoscope segment based on the first laryngoscope segment screened based on the audio; wherein the second laryngoscope segment is a segment of the stroboscopic part, which can show the condition of the vocal cords in slow motion, and the third laryngoscope segment is a portion of the vocal cord glottis area that is clearly captured by the camera;

[0080] The acquisition of the second laryngoscope segment includes the following steps:

[0081] Identify the black screen frames in the original laryngoscopy video, and divide the original laryngoscopy video into several segments based on the black screen frames; perform image brightness analysis on the segments, calculate the number of times the brightness shakes up and down in each segment, and select the segments with brightness shake times greater than half of the total number of frames in the segment as stroboscopic segments; extract the stroboscopic segments and splice them in sequence to form a second laryngoscopy segment. It is understandable that since the vibration frequency of the vocal cords when making sounds is as high as 80 to 2000 times per second, it is difficult to visually see the performance of the vocal cords during vibration with the human eye, so the stroboscopic part in the laryngoscopy video is very important. In order to obtain this section, the doctor expert needs to manually find such a segment from the entire laryngoscopy video, which is not convenient. The present invention proposes an automated laryngoscope stroboscopic part capture method, which relies on the brightness analysis of the image to automatically capture the stroboscopic part in the laryngoscopy video. Since the brightness of the video image will be like this during stroboscopy. Figure 5 As shown in the figure, the video image shakes violently, so by analyzing the image brightness in the video, the stroboscopic part in the original laryngoscope video can be accurately located. The indicator function at the bottom of the figure represents the turning point between the stroboscopic part and the normal part (representing the starting and ending points of the stroboscopic video, respectively), and the curve with larger fluctuations shows the change of video image brightness over time.

[0082] The acquisition of the third laryngoscope segment includes the following steps:

[0083] 301 recognizes the first laryngoscope segment frame by frame through the vocal cord recognition model and marks the confidence, and sorts the first laryngoscope segment according to the confidence to obtain the first K segments as the third laryngoscope segment; the setting of K is flexibly set according to the demand.

[0084] It is understandable that the vocal cord recognition model is designed to identify the vocal cord part in the picture. During actual operation, by inputting the original laryngoscope video, the system directly outputs the second laryngoscope segment corresponding to the stroboscopic part and the third laryngoscope segment obtained by audio and video screening for the doctor's reference. The present invention utilizes a target detection algorithm based on deep learning to perform vocal cord target detection on frame-level images in laryngoscope videos. Through the target detection algorithm, the glottis area in the image can be identified and framed, and the recognition confidence is marked. Finally, the image with the highest confidence is selected based on the recognition confidence obtained above. The following is a description of the construction, training and reasoning of the target recognition model;

[0085] Specifically, the embodiment of the present invention adopts an image target detection model based on the Yolo-v5 model when constructing an image recognition algorithm. Due to the lack of open source recognition data sets, existing work requires professional practitioners to annotate the vocal cords and generate data sets manually. However, such an annotation method is time-consuming and labor-intensive. The embodiment of the present invention proposes a method for quickly obtaining a vocal cord recognition data set. Through the external open source glottal segmentation data set BAGLS, a vocal cord recognition data set that meets the target detection conditions is constructed. Figure 6 As shown in the figure, the specific annotation method is as follows: construct a rectangular coordinate system based on the image, and use the upper left corner of the image as the origin (0, 0) to establish the x-axis and y-axis; use the label mask of the glottis segmentation in the glottis segmentation dataset to obtain the coordinates of the upper, lower, left, and right vertices of the vocal cords, which are recorded as U(x 1 ,y 1 )、D(x 2 ,y 2 )、L(x 2 ,y 2 )、R(x 2 ,y 2 ); Expand the four vertex coordinates by n pixels to obtain new vertex coordinates, namely U′(x 1 -n,y 1 )、D′(x 2 +n,y 2 )、L′(x 3 ,y 3 -n)、R′(x 4 ,y 4 +n); connect the four expanded vertices U′, D′, L′, and R′ with a rectangle to frame the glottis area and obtain the label required for target detection.

[0086] Through the above process, the dataset originally used for glottal segmentation is converted into a dataset for vocal cord recognition. The newly generated recognition data can be input into the YOLO-V5 model for training to obtain the final vocal cord target recognition model.

[0087] Furthermore, the trained vocal cord recognition model is used to perform vocal cord recognition reasoning on the laryngoscope video clips obtained after screening the speech module. Since the analysis of laryngeal paralysis requires video clips rather than single pictures, and emphasizes the continuity of the vocalization process, it is not possible to select only a few pictures with the best reasoning effect during the reasoning process. The present invention analyzes the video clips frame by frame according to the vocal cord recognition model, and obtains a target detection frame for vocal cord recognition and its recognition confidence for each frame. Then, the frame-level confidence is obtained to further infer the best clips. The specific reasoning process is as follows:

[0088] A sliding window of fixed length t is created along the time axis, and the video frames are traversed with t / 2 as the window shift.

[0089] Calculate the mean of the built-in confidence for each window.

[0090] Maintain a list L, which is used to store the start and end timestamps of the K segments with the highest placement confidence.

[0091] Set a variable that is used to calculate the value in list L that has the smallest mean confidence.

[0092] During the traversal process, if the mean confidence value of the current window is greater than the variable, the segment with the smallest mean confidence value in L is recalculated and removed.

[0093] If the two obtained fragments have 50% duplication, the one with the larger mean confidence score is retained and the other is discarded.

[0094] Finally, select the K window segments with the largest confidence mean and save them.

[0095] Through the above method, the present invention realizes the extraction of the first. It should be noted that the processing object of the above method is the first laryngoscopic segment obtained after preliminary screening by voice analysis, and the third laryngoscopic segment refers to the video segment in which the patient makes the "e" sound and the vocal cord part can be clearly captured by the camera, that is, the video content is further analyzed on the basis of the first laryngoscopic segment to obtain a more effective third laryngoscopic segment.

[0096] 302 generating a glottis region mask by a glottis segmentation algorithm and segmenting the glottis region from the glottis image;

[0097] It is understandable that the glottis segmentation algorithm aims to output the shape and size of the glottis in the image in a masked manner. The mask obtained by glottis segmentation can automatically measure and calculate objective indicators such as the boundary point set of the glottis, the area size of the glottis, and the opening and closing angle of the glottis, providing a basis for subsequent intelligent measurement. The present invention completes the operation of glottis segmentation on the image by constructing a diffusion model. The diffusion model is an emerging generative model. During training, the original image is denoised and then denoised to complete the reconstruction of the image, emphasizing that each step of denoising is reversible. The present invention uses the above-mentioned BAGLS glottis segmentation data set to construct a diffusion model for training. Through the glottis segmentation model, the approximate shape of the glottis can be outlined, paving the way for subsequent measurement and parameter calculation.

[0098] The architecture of the glottal segmentation algorithm is as follows Figure 7 As shown, the skeleton of the model is composed of the UNet model, taking the t+1 moment as an example:

[0099] The original laryngoscope image is passed through two residual convolution layers to obtain the original image features, which provide effective information for vocal cord segmentation. The noisy label image of the original laryngoscope image after t+1 iterations is added to the time embedding vector obtained at time t+1, and the features obtained by downsampling through the residual convolution layer are used to obtain the potential representation of the information at time t+1.

[0100] Using the self-attention mechanism, the latent representation of the information at time t+1 is fused with the original image features to obtain feature e; the fused feature e is added to the time embedding vector obtained at time t+1, and upsampled through the residual convolution layer to obtain the denoised label image at time t+1. During the upsampling process, the features obtained by the convolutional network will be merged with the features obtained in the previous downsampling process.

[0101] During the training process, in order to cover all the BAGLS data, randomly sample pictures and their segmentation labels. In addition to the original pictures and labels, it is also necessary to sample from a uniform distribution to obtain an iteration number t. In order to be able to iterate different noises t, the model must be able to eliminate the noise in this step, so that each iteration of the noise addition operation is reversible and can be denoised by the model.

[0102] During the inference process, since the diffusion model is a generative model, each model inference requires not only an original vocal cord image as a conditional prior, but also an input of a Gaussian noise image for the diffusion model to process and generate the final result. Since the model has learned each inverse process from Gaussian noise to the final segmentation label during training, with the vocal cord image to be segmented as a prior, the model can gradually reduce the noise of the Gaussian noise image to generate the final required vocal cord segmentation label.

[0103] The second laryngoscope segment and the third laryngoscope segment are both obtained by processing the first laryngoscope segment, and the two are in a parallel relationship. By performing diversified processing on the screened first laryngoscope segment, the doctor is provided with multi-angle vocal cord information as a reference.

[0104] 303 Calculate the physical properties of the glottis according to the glottis area mask.

[0105] It can be understood that the glottis region mask is obtained through the glottis segmentation in step 303, and the approximate outline of the glottis region can be obtained. Based on the outline of the glottis, the various judgment indicators of the glottis region are calculated using the pixel coordinates. Specifically, the glottis region area can be directly obtained by counting the number of pixels of the glottis part in the glottis region mask, and the curve of the glottis region size changing with time is obtained by analyzing and calculating the glottis area of ​​each frame in the video.

[0106] Since the laryngoscope analysis method in the prior art can usually only calculate the angle between the left and right vocal cords of the patient at a certain moment, and since it is only the information of the angle between the left and right vocal cords at a single moment, it cannot help the doctor to judge the condition of the patient's vocal cords well, the embodiment of the present invention also proposes a method for automatically identifying and measuring the change of the opening and closing angle of the left and right vocal cords over time, such as Fig. 9 The time series shown shows the changes in the opening and closing angles of the left and right vocal cords of the patient, providing doctors with multi-angle indicators for laryngeal paralysis diagnosis: Figure 8 and Fig.10 As shown in the figure, the vocal cord opening and closing angle needs to be intelligently marked and measured through the glottal area mask, including the following steps:

[0107] First, the algorithm obtains the coordinates of the top, bottom, left, and right vertices from the glottal mask and calculates their center points, which are recorded as U, D, R, L, and C (x c ,y c ). Then, a straight line connecting points C and D is established, and the function f(x) passing through C and D is calculated.

[0108] Then, the equidistant point C 1 ,C 2 ,…,C N-1 Placed on the line segment intercepted by the glottal mask. For each Ck, calculate its function f′ perpendicular to f(x) k (x), and determine its intersection point L with the glottal mask k and R k For k∈[1,N-1].

[0109] To ensure consistency, the coordinate system is rotated by an angle γ to make f(x) parallel to the y-axis, and the new coordinates of all intersection points after rotation are obtained, denoted as L λ,k and R λ,kUsing these points, we can approximate a quadratic curve q using the least squares method in the rotated coordinate system. γ (x). Calculate q γ The lowest point D of (x) q , and map it back to the original coordinate system

[0110] Then, establish a calibrated center line f″(x) connecting points C and D q , and intersects with the glottal mask at point D′. Finally, by connecting the points L k and R k and D′, respectively, to obtain the glottal angle ∠L k D′C and ∠CD′R k Specifically, by calculating L k and R k The distance from f″(x) and L k and R k The ratio of the distance to D′ is used to obtain the cosine value of the left and right glottal angles, and then the arc cosine is used to obtain the final left and right vocal cord opening and closing angles.

[0111] The specific steps are as follows:

[0112]

[0113]

[0114] The beneficial effects of the present invention are:

[0115] 1) A method for extracting video clips of the stroboscopic part in a laryngoscope video, that is, a method for extracting the stroboscopic part in a laryngoscope video by using video image brightness analysis.

[0116] 2) A multimodal method for extracting the best video clips, i.e., a laryngoscope video analysis method based on speech and image processing. Specifically, a method for processing laryngoscope videos using speech keyword recognition combined with image segmentation and recognition.

[0117] 3) A method for extracting intelligent auxiliary diagnostic indicators for laryngeal paralysis, that is, a method for automatically identifying and measuring the changes in the opening and closing angles of the left and right vocal cords over time.

[0118] 4) Propose a method to quickly obtain vocal cord recognition dataset.

[0119] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0120] The above embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for those of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. A method for assisting the diagnosis of laryngeal paralysis by combining audio and video processing, characterized in that: The following steps are involved: Obtaining original laryngoscope video with audio information and video information; Performing brightness analysis on the original laryngoscope video to detect and filter out stroboscopic segments in the original laryngoscope video; The audio information is predicted by a keyword recognition model to extract a valid sound segment, and a first laryngoscope segment corresponding to the valid sound segment is obtained from the original laryngoscope video segmentation according to a matching relationship between the audio information and the video information; Detecting and identifying the glottis region in the first laryngoscope segment based on the video information, and segmenting the glottis region from the first laryngoscope segment; Mark the glottis area and calculate the center point of the line connecting the glottis apex and make the midline of the glottis, draw a perpendicular line of N equally divided points on the midline, and record the intersection of the perpendicular line and the glottis boundary; Calculate the angle changes of the left vocal cord opening and closing angle and the right vocal cord opening and closing angle formed by the line connecting the intersection point and the glottis apex and the line connecting the glottis apex and the center point along the time series; The method of predicting the audio information by a keyword recognition model to extract a valid sound segment, and obtaining a first laryngoscope segment from the original laryngoscope video segment according to the valid sound segment, comprises the following steps: Extracting audio information of the original laryngoscope video by an audio processing tool, and converting the audio information into a Mel spectrum form; Constructing a keyword recognition model, dividing the audio information in the form of Mel spectrum into several segments and inputting the segments into the keyword recognition model for training; The segmented audio segments are input into a trained keyword recognition model to output the utterance probability of the segments, and valid utterance segments are searched according to the utterance probability and the first laryngoscope segment corresponding to the valid utterance segment is extracted.

2. The method according to claim 1, characterized in that The detecting and identifying the glottis region in the first laryngoscope segment based on the video information, and segmenting the glottis region from the first laryngoscope segment, comprises the following steps: Recognize the first laryngoscope segment frame by frame through the vocal cord recognition model and mark the confidence, sort the first laryngoscope segment according to the confidence to obtain the first K segments as the third laryngoscope segment; A glottis region mask is generated by a glottis segmentation algorithm and the glottis region is segmented from the glottis image; Physical properties of the glottis are calculated based on the glottis region mask.

3. The method according to claim 1, characterized in that The step of extracting the audio information of the original laryngoscope video by an audio processing tool and converting the audio information into a Mel spectrum form comprises the following steps: The audio information is sampled 16,000 times per second and a sampling result is generated; A 0.025 second window is taken along the time axis to perform spectrum analysis on the sampling results, and the number of windows for the spectrum analysis is set to 512; The Mel spectrum of the audio information is calculated through 80 Mel filters.

4. The method according to claim 1, characterized in that The keyword recognition model includes a two-dimensional convolutional layer, a maximum pooling layer, a residual module, an average pooling layer and two fully connected layers connected in sequence; after the audio clip is input into the keyword recognition model, the keyword recognition model outputs a two-dimensional probability vector, and by comparing the probability vectors, the utterance probability of the audio clip is inferred.

5. The method according to claim 1, characterized in that The step of performing brightness analysis on the original laryngoscope video to detect and filter out the stroboscopic segments in the original laryngoscope video further includes the following steps: Identify black screen frames in the original laryngoscope video, and divide the original laryngoscope video into a plurality of segments based on the black screen frames; Performing image brightness analysis on the segments, calculating the number of times the brightness in each segment jitters up and down, and selecting the segments whose brightness jitters are greater than half of the total number of frames of the segments as stroboscopic segments; The stroboscopic segments are extracted and sequentially spliced ​​to form a second laryngoscope segment.

6. The method according to claim 2, characterized in that The vocal cord recognition model is obtained by inputting the vocal cord recognition data set into the YOLO-v5 model, and the construction of the vocal cord recognition data set includes the following steps: Obtain a glottis segmentation dataset from a public database and construct an image coordinate system in the glottis segmentation dataset; Get the vertex coordinates of the vocal cords through the glottal segmentation label mask; The vertex coordinates are expanded and connected with rectangles to obtain target detection labels, and the target detection labels are summarized as a vocal cord recognition dataset.

7. The method according to claim 2, characterized in that The glottis segmentation algorithm is constructed by a diffusion model and includes the following steps: Obtain a glottal segmentation dataset from a public database, and randomly select sample images and corresponding segmentation labels from it; Input the sampled image into two residual convolutional layers to obtain the original features; The noisy sampled image after t+1 iterations is added to the time embedding vector obtained at time t+1, and the input is downsampled in the residual convolution layer to obtain the potential representation; Through the self-attention mechanism, the latent representation is fused with the original features to obtain the combined features; The combined features are added to the time embedding vector at time t+1 and input into the residual convolution layer for upsampling to obtain the denoised label image at time t+1.

8. The method according to claim 2, characterized in that The method of identifying the first laryngoscope segment frame by frame by using the vocal cord recognition model and marking the confidence, and sorting the first laryngoscope segment according to the confidence to obtain the first K segments as the third laryngoscope segment, comprises the following steps: A sliding window of fixed length t is created along the time axis, the first laryngoscope segment is traversed by window movement, and the confidence mean of each window segment is calculated; Generate a list for storing window fragments and preset a variable, wherein the variable is the window fragment with the smallest confidence in the list; During the traversal process, if the confidence mean of the window fragment is greater than the variable, the window fragment with the smallest confidence mean in the list is deleted and the variable is updated; several window fragments with the largest confidence mean are selected and saved.

9. The method according to claim 2, characterized in that The acquisition of the vocal cord opening and closing angle comprises the following steps: The top coordinate U, bottom coordinate D, left coordinate R and right coordinate L of the glottis are obtained by the glottis mask, and the center point C thereof is calculated; Establish a line CD passing through the center point C and the bottom coordinate D, calculate its function f(x) and find N equidistant points C on the function N-1 ; Calculate the function f' at equidistant points perpendicular to f(x) k (x) and calculate its intersection point L with the glottal mask k and R k ; Rotate the coordinate axis so that f(x) is parallel to the y-axis and calculate the new coordinate L λ,k and R λ,k , the quadratic curve q is obtained by least squares fitting γ (x) and its lowest point D q ; Create the calibration center line f" (x) connecting the center point C and the lowest point D q , the midline f" (x) intersects the glottal mask at the intersection point D', and by calculating L k and R k The distance from f”(x) and L k and R k The ratio of the distance to D' gives the cosine value of the left glottis angle and the right glottis angle.

Citation Information

Patent Citations

  • Semi-automatic labeling method, medium, equipment and device for laryngoscope image

    CN112634266A

  • Voice recognition method based on glottis wave information

    CN112735386A