Video analysis method and system for stroboscopic laryngoscope, terminal and storage medium
Through the video analysis method, the features in the stroboscopic video signal are extracted and fused with the TDD algorithm and ResNeXt model, and trajectory pooling and Fisher Vector encoding are solved, which solves the problem of long-term stroboscopic examination, and achieves rapid and accurate video behavior recognition and diagnostic efficiency.
Patent Information
- Application Number
- CN202510135944.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-30
AI Technical Summary
In actual examination, strobe laryngoscopy requires doctors to observe different vocal cord vibration patterns, which takes a long time and requires high medical standards.
Using video analysis method, by obtaining video signals, extracting feature points to form feature point trajectories, using TDD algorithm and ResNeXt model for feature extraction and fusion, trajectory pooling and Fisher Vector encoding, and finally behavior classification is performed through a linear support vector machine.
It realizes the rapid and accurate identification of video behavior, reduces the observation time of doctors, improves diagnostic efficiency, and can identify vibrating videos in different states, expanding the scope of application.
Smart Images

Figure CN120070362A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video analysis methods, and particularly to a video analysis method, system, terminal and storage medium for a stroboscopic laryngoscope. Background Art
[0002] The greatest advantage of a stroboscopic laryngoscope is that, by using the principle of physics, it replaces the flat light with a stroboscopic light source to make the rapidly vibrating vocal cords become a slow movement visible to the naked eye, so that we can observe the minute changes on the vocal cord mucosa. It is a non-invasive, non-damaging and less painful examination method. The stroboscopic laryngoscope adopts the heterodyne principle to slow down the vibration of the vocal cords and is widely used in the examination of laryngeal diseases and the observation of curative effects, etc.
[0003] When performing an examination with a stroboscopic laryngoscope, it is necessary to respectively shoot a dynamic image, a static image, a phase from 0 to 360 degrees, and the vibration patterns of the vocal cords in true voice, false voice, low pitch, high pitch, soft voice, and strong voice according to needs. In actual examinations, doctors may need to observe the corresponding different vibration patterns of the vocal cords, which is time-consuming and also places certain requirements on the medical level of doctors. Summary of the Invention
[0004] The technical problem to be solved by the present invention is that in the actual examination of a stroboscopic laryngoscope, doctors may need to observe the corresponding different vibration patterns of the vocal cords, which is time-consuming. In view of the above-mentioned defects of the prior art, a video analysis method, system, terminal and storage medium for a stroboscopic laryngoscope are provided.
[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0006] Construct a video analysis method for a stroboscopic laryngoscope, including the following steps:
[0007] Obtain a video signal;
[0008] Extract feature points in the video signal to form a feature point trajectory;
[0009] Perform feature extraction on the feature points through the TDD algorithm to obtain a feature mapping to form a feature map;
[0010] Map and project the trajectory of the feature points onto the feature map for trajectory pooling;
[0011] Fuse the trajectory of the feature points and the feature map for feature fusion and perform encoding to obtain a feature vector of the global dynamics and structure of the video;
[0012] Perform behavior classification through the application of a linear support vector machine to obtain a video analysis result.
[0013] Preferably, in the process of performing feature extraction on the feature points through the TDD algorithm to obtain a feature mapping to form a feature map, it further includes:
[0014] Use the ResNeXt model to extract feature points, and calculate the feature map of the feature points to obtain the feature map.
[0015] Preferably, in the process of fusing the trajectory of the feature points and the feature map and performing feature encoding to obtain the feature vector of the global dynamics and structure of the video, it further includes:
[0016] Perform feature encoding through the Gaussian mixture model and Fisher Vector to collect the depth features of all feature point trajectories.
[0017] Preferably, in the process of performing feature encoding through the Gaussian mixture model, it further includes:
[0018] Calculate the probability of each feature point under the k-th Gaussian component through the Gaussian probability density, calculate the posterior probability of each feature belonging to the k-th Gaussian classification based on the calculated probability, and calculate the Fisher Vector to aggregate the features into a global video description.
[0019] Preferably, in the process of calculating the Fisher Vector, it further includes:
[0020] Obtain the Fisher Vector aggregation by the mean partial derivative of the mean part and the variance partial derivative of the variance part, and summing all the mean partial derivatives and variance partial derivatives of each Gaussian component.
[0021] Preferably, in the process of performing behavior classification through the application of the linear support vector machine to obtain the video analysis result, it further includes:
[0022] When training the linear support vector machine, find the optimized weights and biases to maximize the margin and calculate each video feature.
[0023] Preferably, in the process of extracting the feature points in the video signal to form the feature point trajectory, it further includes:
[0024] Calculate the motion vector of each feature point through the optical flow method, and update its position along the direction of the motion vector for each feature point to form the trajectory of the feature points.
[0025] Construct a video analysis system for a stroboscope, including:
[0026] A video acquisition module for acquiring video signals;
[0027] A video analysis module, which is used to extract feature points in a video signal to form a feature point trajectory, perform feature extraction on the feature points through a TDD algorithm to obtain a feature map, project the trajectory of the feature points onto the feature map for trajectory pooling, fuse the trajectory of the feature points and the feature map, and perform feature encoding to obtain a feature vector of the global dynamics and structure of the video, and classify and identify the video features;
[0028] A video report module, which outputs the conclusion analyzed by the video analysis module.
[0029] Construct an intelligent terminal, characterized in that the intelligent terminal includes at least one memory, at least one processor, and a video analysis program of a stroboscopic laryngoscope stored on the memory and executable on the processor. When the video analysis program of the stroboscopic laryngoscope is executed by the processor, the steps of the video analysis method of the stroboscopic laryngoscope as described above are implemented.
[0030] Construct a computer-readable storage medium, characterized in that a video analysis program of a stroboscopic laryngoscope is stored on the computer-readable storage medium. When the video analysis program of the stroboscopic laryngoscope is executed by a processor, the steps of the video analysis method of the stroboscopic laryngoscope as described above are implemented.
[0031] The beneficial effects of the present invention are as follows: First, obtain feature points and obtain the trajectory of the feature points through an optical flow method, and then perform feature extraction on the feature points through a TDD algorithm to form a feature map. The ResNeXt model is used to improve the TDD algorithm, making the TDD algorithm more efficient in calculation and faster in operation speed. The efficiency and accuracy of the model are also improved through grouped convolution and feature fusion mechanisms. Then, project the feature point trajectory onto the feature map for trajectory pooling, and aggregate the features along the trajectory pooling to form a global description, and calculate the Gaussian probability density function and posterior probability to implement Fisher Vector, so as to perform Fisher Vector encoding on the features, realize aggregating the features into a global video description, and perform behavior classification by applying a linear support vector machine in the subsequent process, so that the video behavior can be completely recognized, and the recognition speed is faster and more accurate. At the same time, vibration videos in different states can also be recognized, with a wider application range and higher and faster recognition efficiency. Description of the Drawings
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will further illustrate the present invention in conjunction with the drawings and embodiments. The drawings in the following description are only partial embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings:
[0033] Figure 1 It is a schematic flowchart of the video analysis method of a preferred embodiment of the present invention;
[0034] Figure 2 Schematic structural diagram of the video analysis system according to a preferred embodiment of the present invention;
[0035] Figure 3 Schematic structural diagram of the intelligent terminal according to a preferred embodiment of the present invention. Specific embodiments
[0036] In order to make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0037] A video analysis method for a stroboscopic laryngoscope according to a preferred embodiment of the present invention; as Figure 1 shown, it is a flowchart of a video analysis method for a stroboscopic laryngoscope provided by an embodiment of the present invention. This method can be executed by a device, and the device can be implemented by software and / or hardware.
[0038] Specifically, in this embodiment, the video analysis method for the stroboscopic laryngoscope includes:
[0039] S10: Obtain a video signal;
[0040] In a preferred embodiment of the present invention, a video image of the larynx is obtained through a stroboscopic laryngoscope. Different modes can be selected as needed to obtain, store, and mark the video image. For example, when the larynx emits a true voice, a corresponding true voice analysis model is selected; when emitting a false voice, a corresponding false voice analysis model is selected, so as to obtain a corresponding analysis model for video analysis according to the voice emitted by the laryngeal video image.
[0041] S20: Extract the dense trajectories of feature points in the video signal to form a feature point trajectory;
[0042] In a preferred embodiment of the present invention, dense trajectory extraction needs to be performed on the obtained video first. First, feature points need to be obtained from the video as the first step to determine the dense trajectory for video analysis, and these feature points are tracked on consecutive video frames to form a feature point trajectory.
[0043] Further, in the process of obtaining feature points, initial feature point detection is first performed. Assume that the input video frame sequence is {I t}, for each initial frame I t , a set of feature points {P 1 , P 2......P N}. Then, optical flow calculation is performed to calculate the optical flow field of the frame time step t through Formula ①, and the optical flow method provides a motion vector from frame t to t+1 for each pixel.
[0044] V t (x,y) = (μ t (x,y), v t (x,y)) ①
[0045] Then, for each feature point (x t , y t ), its position is updated through Formula ② to form a feature point trajectory τ:
[0046] (x t+1 , y t+1 ) = (x t + μ t (x t , y t ), y t + v t (x t , y t )) ②
[0047] It will form a series of points on the time window, thus mapping out the trajectory: τ = {(x t , y t ) t , (x t+1 , y t+1 ) t+1 , …}, and the time window can be selected as 15 frames.
[0048] S30: Feature maps of each feature point are obtained by extracting the feature points through depth convolution to form a feature map;
[0049] In the preferred embodiment of the present invention, the TDD algorithm is selected to extract the feature points through depth convolution. In the original TDD algorithm, the VGG model is often used to extract the feature map. In the present invention, the ResNeXt model is used to replace the VGG model in the original TDD algorithm. Compared with the VGG model, the ResNeXt model has fewer parameters and higher computational efficiency, and improves the efficiency and accuracy of the model through grouped convolution and feature fusion mechanisms, and is more suitable for scenarios that require real-time processing in video processing and real-time image recognition. The processed feature map is: F t = ResNeXt(I t ), that is, each F t is the feature map obtained from the network of frame I t .
[0050] S40: Project the trajectory of each feature point onto the feature map for trajectory pooling;
[0051] In a preferred embodiment of the present invention, the trajectory of each feature point is projected onto the feature map for trajectory pooling. First, the trajectory of each feature point is mapped to the corresponding coordinates of the feature map F t as follows:
[0052]
[0053] where s is the scaling ratio.
[0054] Then, perform a max pooling operation on the features within the mapped coordinate region: to achieve mapping and projection of the trajectory of each feature point onto the feature map for trajectory pooling.
[0055] S50: Collect the depth features of all trajectories, and perform Fisher Vector calculation on the features pooled along the trajectories to obtain a global video description;
[0056] In a preferred embodiment of the present invention, first aggregate the TDD features extracted along the trajectories to form a global description, collect the depth features of all trajectories for feature fusion. Then perform feature encoding on the fused features.
[0057] Furthermore, feature encoding can be performed through a Gaussian mixture model (GMM) and Fisher Vector. First, assume that the pooled features are {X 1 , X 2 ......X N}, and the model parameters λ include the weights, means, and variances of K Gaussian components, i.e.: Then, through the Gaussian probability density function: calculate the probability of each feature x n under the k-th Gaussian component:
[0058]
[0059] where D is the dimension of the feature. Then, use the above probability to find the posterior probability that each feature belongs to the k-th Gaussian component:
[0060]
[0061] After obtaining the posterior probability of each feature relative to the Gaussian components, the Fisher Vector can be calculated to aggregate the features into a global video description. The Fisher Vector also includes the mean partial derivative and the variance partial derivative. The mean partial derivative, i.e., the part of the FV related to the mean:
[0062]
[0063] where σ k is the square root of the diagonal element of the diagonal variance matrix Σ k ;
[0064] The variance partial derivative is the part of FV with respect to variance:
[0065]
[0066] Finally, through Fisher Vector aggregation, for each Gaussian component k, the partial derivatives of the means and variances of all samples are summed:
[0067]
[0068] Ultimately, these partial derivative contributions constitute the Fisher Vector of the entire video description, which is composed of the partial derivatives of each Gaussian component and provides rich information about the changes in the data distribution. In this way, TDD integrates local trajectory feature information into a feature vector representing the global dynamics and structure of the video, which can be directly used in subsequent classification tasks.
[0069] S60: Classify and identify the encoded video features and output the analysis conclusion;
[0070] In a preferred embodiment of the present invention, a linear support vector machine (SVM) is applied for behavior classification. During training, the optimized weight w and bias b are found to maximize the margin:
[0071]
[0072] For each video feature φ(x i ),
[0073] y i (w · φ(x i ) + b) ≥ 1 - ξ i , ξ i ≥ 0
[0074] where φ(x i ) is the feature encoded by the Fisher Vector, thereby realizing video feature classification and obtaining the corresponding analysis conclusion for analysis.
[0075] Through the above method, first obtain the feature points and get the trajectories of the feature points through the optical flow method, and then extract the feature points through the TDD algorithm to form a feature map. Among them, the ResNeXt model is used to improve the TDD algorithm, making the TDD algorithm more efficient in calculation and faster in operation speed. The efficiency and accuracy of the model are also improved through grouped convolution and feature fusion mechanisms. Then, map and project the feature point trajectories onto the feature map for trajectory pooling. Aggregate the features along the trajectory pooling to form a global description, calculate the Gaussian probability density function and posterior probability to implement the Fisher Vector, so as to perform FisherVector encoding on the features, realize aggregating the features into a global video description, and perform behavior classification by applying a linear support vector machine in the subsequent process, so as to completely identify the video behavior, with faster and more accurate recognition speed. At the same time, it can also identify vibration videos in different states, with a wider application range and higher and faster recognition efficiency.
[0076] Corresponding to a video analysis method for a stroboscopic laryngoscope, the present invention also provides a video analysis system for a stroboscopic laryngoscope. Specifically, as Figure 2 shown, the video analysis system for the stroboscopic laryngoscope includes: a video acquisition module 100, a video analysis module 200, and a video reporting module 300.
[0077] The video acquisition module 100 is used to obtain video signals;
[0078] Specifically, obtain the video image of the larynx through a stroboscopic laryngoscope, and different modes can be selected according to needs to obtain the video image and perform storage and marking. For example, when the larynx emits a true voice, select the corresponding true voice analysis model; when emitting a false voice, select the corresponding false voice analysis model, so as to obtain the corresponding analysis model for video analysis according to the voice emitted from the video image of the larynx.
[0079] The video analysis module 200 is used to extract the feature points in the video signal to form feature point trajectories, perform feature extraction on the feature points through the TDD algorithm to obtain a feature map by forming a feature map, map and project the trajectories of the feature points onto the feature map for trajectory pooling, fuse the trajectories of the feature points and the feature map, and perform feature encoding to obtain the feature vectors of the global dynamics and structure of the video, and classify and identify the video features.
[0080] Specifically, it is necessary to first perform dense trajectory extraction on the obtained video. First, it is necessary to obtain the feature points from the video as the first step to determine the dense trajectory for video analysis, and track these feature points on consecutive video frames to form trajectories. In the process of obtaining the feature points, first perform initial feature point detection, and then perform optical flow calculation. The optical flow method provides a motion vector for each pixel from frame t to t+1 to form a trajectory.
[0081] Further, the TDD algorithm of the ResNeXt model is used for feature point extraction to obtain the feature map of each feature point. Then, the trajectory map of each feature point is projected onto the feature map for trajectory pooling, and the TDD features extracted along the trajectory are aggregated to form a global description. The depth features of all trajectories are collected for feature fusion, and then the fused features are encoded to calculate the posterior probability of each feature relative to the Gaussian mixture to calculate the Fisher Vector, and the mean partial derivative and variance partial derivative are calculated to obtain the Fisher Vector of the entire video description, so as to realize that TDD integrates the local trajectory feature information into a feature vector representing the global dynamics and structure of the video, which can be directly used in the subsequent classification task. Finally, a linear support vector machine (SVM) is applied for behavior classification to achieve video feature classification and corresponding analysis conclusions can be obtained for analysis.
[0082] The video reporting module 300 outputs the conclusions analyzed by the video analysis module;
[0083] Specifically, the conclusions analyzed are output through the video reporting module.
[0084] Based on the above embodiments, the present invention also provides an intelligent terminal, the principle block diagram of which is as Figure 3 shown. The above intelligent terminal includes a processor, a memory, a network interface, and a display screen connected through a system bus. The processor of the intelligent terminal is used to provide computing and control capabilities. The memory of the intelligent terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a video analysis program for a stroboscopic laryngoscope. The memory provides an environment for the operation of the operating system and the path generation program in the non-volatile storage medium. The network interface of the intelligent terminal is used to communicate with an external terminal through a network connection. When the video analysis program for the stroboscopic laryngoscope is executed by the processor, the steps of any one of the above video analysis methods for the stroboscopic laryngoscope are implemented. The display screen of the intelligent terminal can be a liquid crystal display screen or other display screens.
[0085] Those skilled in the art can understand that Figure 3 the principle block diagram shown only shows the block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the intelligent terminal to which the solution of the present invention is applied. The specific intelligent terminal may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0086] In an embodiment of the present invention, an intelligent terminal is provided. The above intelligent terminal includes a memory, a processor, and a video analysis program for a stroboscopic laryngoscope stored on the above memory and executable on the above processor. When the above video analysis program for the stroboscopic laryngoscope is executed by the above processor, the following operation instructions are performed:
[0087] Obtain the laryngeal video signal through a stroboscopic laryngoscope;
[0088] Perform initial feature point detection on the video signal to generate a set of feature points;
[0089] Calculate the optical flow field at frame time step t and calculate the motion vector to form the trajectory of each feature point;
[0090] Extract the feature points using deep convolution to obtain a feature map;
[0091] Map the trajectory of each feature point to the feature map for trajectory pooling;
[0092] Aggregate the TDD features extracted along the trajectory to form a global description, and collect the depth features of all trajectories;
[0093] Perform feature encoding through a Gaussian mixture model and a Fisher Vector, and aggregate the features into a global video description;
[0094] Apply a linear support vector machine for video classification and recognition;
[0095] Output the conclusion of the recognition.
[0096] An embodiment of the present invention also provides a computer-readable storage medium, on which a video sub-program of a stroboscopic laryngoscope is stored. When the video sub-program of the stroboscopic laryngoscope is executed by a processor, the steps of any one of the video sub-methods of the stroboscopic laryngoscope provided by the embodiment of the present invention are implemented.
[0097] It should be understood that the present invention is described through some embodiments. Those skilled in the art know that without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. In addition, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.
Claims
1. A video analysis method for stroboscopic laryngoscope, characterized in that: The steps include: Acquire video signal; Extracting feature points from the video signal to form a feature point trajectory; The feature points are extracted by TDD algorithm to obtain feature mapping to form a feature map; Project the trajectory map of the feature points onto the feature map for trajectory pooling; The trajectory of the feature points and the feature map are fused and encoded to obtain the feature vector of the global dynamics and structure of the video; The video analysis results were obtained by applying a linear support vector machine for behavior classification.
2. The video analysis method according to claim 1, characterized in that: In the process of extracting features from feature points by using the TDD algorithm to obtain feature mapping to form a feature graph, the process also includes: The ResNeXt model is used to extract feature points, and the feature mapping of the feature points is calculated to obtain the feature map.
3. The video analysis method according to claim 1, characterized in that: In the process of fusing the trajectory of the feature point with the feature map and performing feature encoding to obtain the feature vector of the global dynamics and structure of the video, the following is also included: Feature encoding is performed through Gaussian mixture model and Fisher Vector to collect deep features of all feature point trajectories.
4. The video analysis method according to claim 3, characterized in that: In the process of encoding features by using a Gaussian mixture model, the following is also included: The probability of each feature point under the kth Gaussian component is calculated through the Gaussian probability density, and the posterior probability of each feature belonging to the kth Gaussian classification is calculated, and the Fisher Vector is calculated to aggregate the features into a global video description.
5. The video analysis method according to claim 4, characterized in that: The process of calculating Fisher Vector also includes: The Fisher Vector aggregation is obtained by taking the mean partial derivative of the mean part and the variance partial derivative of the variance part, and summing up all the mean partial derivatives and variance partial derivatives of each Gaussian component.
6. The video analysis method according to claim 1, characterized in that: In the process of obtaining the video analysis result by applying the linear support vector machine to classify the behavior, the process further includes: When applying linear support vector machine training, find the optimized weights and biases to maximize the margin of error in computing per-video features.
7. The video analysis method according to claim 1, characterized in that: The process of extracting feature points from the video signal to form feature point trajectories also includes: The motion vector of each feature point is calculated by the optical flow method, and the position of each feature point is updated along the direction of the motion vector to form the trajectory of the feature point.
8. A video analysis system for stroboscopic laryngoscope, characterized in that: include: A video acquisition module, used for acquiring video signals; The video analysis module is used to extract feature points from the video signal to form feature point trajectories, extract the feature points through the TDD algorithm to obtain feature mapping to form a feature graph, project the trajectory mapping of the feature points onto the feature graph for trajectory pooling, fuse the trajectory of the feature points and the feature graph and perform feature encoding to obtain the feature vector of the global dynamics and structure of the video, and classify and identify the video features; The video report module outputs the conclusions of the video analysis module.
9. An intelligent terminal, characterized in that: The intelligent terminal includes at least one memory, at least one processor, and a video analysis program of the stroboscopic laryngoscope stored in the memory and executable on the processor. When the video analysis program of the stroboscopic laryngoscope is executed by the processor, the steps of the video analysis method of the stroboscopic laryngoscope as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a stroboscopic laryngoscope video analysis program, and when the stroboscopic laryngoscope video analysis program is executed by a processor, the steps of the stroboscopic laryngoscope video analysis method according to any one of claims 1 to 7 are implemented.