Vocal cord problem identification feedback system for ophthalmology and otorhinolaryngology department
Through the multimodal data acquisition and fusion of vocal cord problems identification and feedback system, combined with AI diagnosis and rehabilitation guidance, the limitations and subjective problems of diagnosis in the existing technology are solved, and efficient and accurate diagnosis and treatment of vocal cord problems are achieved.
Patent Information
- Application Number
- CN202510657309.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-08-26
AI Technical Summary
The existing vocal cord diagnosis technology for ventricular vocal cord problems has limitations in single-modal detection, static evaluation system and subjective dependence, resulting in the high rate of misdiagnosis and high rate of misdiagnosis in early cancers.
Multimodal data acquisition and fusion technology is adopted, combining sound, imaging and biofeedback modules, and dynamically updated deep learning model diagnosis is carried out through the AI diagnosis module, and real-time feedback and remote diagnosis and treatment support are provided in combination with the rehabilitation guidance module.
The early cancer detection rate has been improved, the misdiagnosis rate has been reduced, the equipment portability and diagnostic efficiency have been improved, and the equipment cost has been reduced.
Smart Images

Figure CN120531329A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of ENT vocal cord problem identification equipment, and specifically to an ENT vocal cord problem identification feedback system. Background Art
[0002] In recent years, vocal cord problem identification technology has seen significant development. Currently, ENT diagnoses primarily rely on electronic dynamic laryngoscopes and acoustic analysis systems, but these systems have significant limitations: Single-modality detection limitations: Traditional methods often utilize a single technique (such as white-light laryngoscopes or acoustic parameter analysis) that is unable to simultaneously integrate multi-dimensional data such as vocal cord vibration dynamics, tissue elastic modulus, and muscle coordination, resulting in a high rate of missed diagnosis of early-stage cancers.
[0003] Static Assessment System: While existing devices such as stroboscopic laryngoscopes can observe vocal cord mucosal waves, they lack the ability to dynamically correlate with disease databases and are unable to update diagnostic models in real time. For example, the German XION dynamic laryngoscope, while supporting electronic shutter stroboscopic technology, still relies on manual experience to assess non-periodic vocal cord movement.
[0004] Subjective reliance: Diagnosis of conditions like vocal cord paralysis relies heavily on physician experience, resulting in a high rate of misdiagnosis. Even using the ELS voice assessment standard, subjective judgment must be combined with imaging data. To address this, we propose a vocal cord problem identification and feedback system for ENT. Summary of the Invention
[0005] The purpose of the present invention is to provide an ENT vocal cord problem identification and feedback system to solve the problems raised in the above background technology.
[0006] To achieve the above-mentioned object, the present invention provides the following technical solutions: a vocal cord problem identification and feedback system for ENT, comprising a sound acquisition module, an image acquisition module, a biofeedback module, a data processing module, an AI diagnosis module, and a rehabilitation guidance module; The sound collection module includes a microphone array and an adaptive noise reduction unit; The image acquisition module is equipped with an endoscope camera and an image enhancement processor; The biofeedback module integrates a laryngeal electromyographic sensor and a three-dimensional motion simulator; The data processing module performs multimodal feature extraction and fusion; The AI diagnostic module includes a dynamically updated deep learning model; The rehabilitation guidance module generates a visual biofeedback training program; The data processing modules are connected to other modules via data buses.
[0007] Preferably, the biofeedback module includes: Annular array sensor to monitor thyroarytenoid / cricothyroid coordination; A fluid dynamics simulator that uses the immersed boundary method to generate 3D vocal fold vibration animations.
[0008] Preferably, the image enhancement processor includes: HSV color space conversion unit for blood vessel texture enhancement; UNet segmentation unit, which performs lesion area identification; Image correction unit, which performs motion compensation based on the vocal cord vibration modal equation.
[0009] Preferably, the deep learning model includes: Backbone network, used for feature extraction; Bidirectional gated recurrent unit to process time series data; Dynamic update unit, synchronously update the vocal cord disease atlas database every week.
[0010] Preferably, the rehabilitation guidance module includes: A virtual reality interface displays an animated simulation of vocal cord movement; Real-time threshold alarm unit, activated when the myoelectric intensity exceeds the preset value; Remote diagnosis and treatment interface, connected to the hospital's central server.
[0011] Preferably, it also includes a hardware structure, which includes: Removable silicone sensing ring with integrated biofeedback module; Portable processing terminal with built-in modular slots for expansion of functional units.
[0012] Compared with the prior art, the beneficial effects of the present invention are: the ENT vocal cord problem identification and feedback system has the following advantages: Accurate diagnosis: Through the cooperation of the sound acquisition module, image acquisition module, biofeedback module, and data processing module, data can be accurately collected. Multimodal fusion improves the detection rate of early cancer and reduces the misdiagnosis rate. Dynamic optimization: The data processing module analyzes the data collected from the multimodal structure, and the model is combined with the weekly updated disease map to improve the accuracy of recurrence prediction; Portable and efficient: The detachable silicone sensing ring and modular slots in the hardware structure are both detachable. The modular hardware design shortens the time required for a single test and reduces equipment costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 Schematic diagram of the control principle of the present invention.
[0014] Figure 2 It is a structural schematic diagram of the present invention.
[0015] Figure 3 It is a structural schematic diagram of the present invention.
[0016] In the figure: sound acquisition module 100; microphone array 101; adaptive noise reduction unit 102; image acquisition module 200; endoscopic camera 201; image enhancement processor 202; HSV color space conversion unit 2021; UNet segmentation unit 2022; image correction unit 2023; biofeedback module 300; laryngeal electromyography sensor 301; ring array sensor 3011; three-dimensional motion simulator 302; fluid dynamics simulator 3021; data processing module 400; AI diagnosis module 500; deep learning model 501; backbone network 5011; bidirectional gated recurrent unit 5012; dynamic update unit 5013; vocal cord disease atlas database 502; rehabilitation guidance module 600; virtual reality interface 601; real-time threshold alarm unit 602; remote diagnosis and treatment interface 603; central server 604; data bus 700; detachable silicone sensor ring 800; portable processing terminal 900; modular slot 901. DETAILED DESCRIPTION
[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention. Example
[0018] See also Figure 1-3 The present invention provides a technical solution: a vocal cord problem identification and feedback system for ENT, comprising a sound acquisition module 100, an image acquisition module 200, a biofeedback module 300, a data processing module 400, an AI diagnosis module 500 and a rehabilitation guidance module 600; The sound collection module 100 includes a microphone array 101 and an adaptive noise reduction unit 102. The microphone array 101 uses a 32-channel high-sensitivity microphone distributed in a ring, supports 70-1600Hz sound wave collection, and is combined with an improved LMS adaptive noise reduction algorithm. Adaptive noise reduction unit 102: generates a noise reduction coefficient using a real-time noise reference signal; The weight update adopts the gradient descent method of symbolic function optimization to ensure that the speech signal-to-noise ratio is ≥25dB in complex environments.
[0019] The image acquisition module 200 is equipped with an endoscope camera 201 and an image enhancement processor 202; Endoscopic camera 201: equipped with a 4K resolution optical lens, supporting 0.1mm-level vocal cord mucosal wave capture; The image enhancement processor 202 includes: HSV color space conversion unit 2021: Enhances vascular texture contrast and optimizes blood flow display through H channel histogram equalization; UNet segmentation unit 2022: uses a skip connection structure to achieve higher lesion recognition accuracy; The biofeedback module 300 integrates a laryngeal electromyographic sensor 301 and a three-dimensional motion simulator 302; Laryngeal EMG Sensor 301: Annular array design (8-channel electrodes), dynamic spacing adjustment range to adapt to different laryngeal anatomical structures; 3D Motion Simulator 302: A fluid dynamics model based on the Immersed Boundary Method (IBM) to simulate the phase difference of vocal cord vibration; The data processing module 400 performs multimodal feature extraction and fusion; Multimodal feature fusion uses principal component analysis (PCA) to reduce the dimension to 512, integrating acoustic features (fundamental frequency perturbation, harmonic-to-noise ratio), imaging features (glottal closure rate, mucosal wave symmetry), and bioelectric features (myoelectric coordination index); The AI diagnostic module 500 includes a dynamically updated deep learning model 501; an EfficientNet-B4 backbone network that extracts 2048-dimensional vocal cord feature vectors; BiGRU time series processing unit 5012: predict the development trend of vocal cord lesions; Dynamic Update Unit 5013: Weekly update of the vocal cord disease atlas database 502 containing over 3,000 clinical cases. The rehabilitation guidance module 600 generates a visual biofeedback training program; the virtual reality interface 601 provides a 6-DOF interactive training scene and real-time feedback on the deviation angle of vocal posture; Real-time threshold alarm unit 602: triggers an audible and visual alarm when the myoelectric intensity of the thyroarytenoid muscle exceeds a safety threshold.
[0020] The data processing module 400 is connected to other modules via a data bus 700. Detachable silicone sensor ring 800: Made of medical-grade silicone, it can be sterilized at high temperature and high pressure for 200 times. Portable processing terminal 900: Modular slot 901 supports expansion of GPU accelerator cards, resulting in higher inference speed.
[0021] In the above solution, the biofeedback module 300 includes: Annular array sensor 3011, monitoring thyroarytenoid / cricothyroid muscle coordination; Fluid Dynamics Simulator 3021 uses the immersed boundary method to generate 3D vocal fold vibration animation.
[0022] In the above solution, the image enhancement processor 202 includes: HSV color space conversion unit 2021, used for blood vessel texture enhancement; UNet segmentation unit 2022, performs lesion region identification; The image correction unit 2023 performs motion compensation based on the vocal cord vibration modal equation.
[0023] In the above solution, the deep learning model 501 includes: Backbone network 5011, used for feature extraction; Bidirectional gated recurrent unit 5012, processing time series data; The dynamic updating unit 5013 synchronously updates the vocal cord disease atlas database 502 every week.
[0024] In the above solution, the rehabilitation guidance module 600 includes: Virtual reality interface 601, showing a simulated animation of vocal cord movement; A real-time threshold alarm unit 602 is activated when the myoelectric intensity exceeds a preset value; Remote diagnosis and treatment interface 603 is connected to the hospital central server 604.
[0025] The above solution also includes a hardware structure, which includes: Removable silicone sensing ring 800, integrated biofeedback module 300; The portable processing terminal 900 has a built-in modular slot 901 for expanding functional units.
[0026] In this ENT vocal cord problem recognition and feedback system, the endoscopic camera 201 captures the vocal cord mucosal waves, and the image enhancement processor 202 performs three-stage optimization: HSV Color Enhancement: Convert RGB to HSV space, perform histogram equalization on the H (hue) channel, dynamically adjust the S (saturation) value to enhance the contrast of vascular texture, and strengthen the color difference between the vocal cord edge and the lesion area.
[0027] UNet lesion segmentation: The encoder uses ResNet-34 to extract multi-scale features, and the decoder fuses shallow texture information through skip connections; it outputs a lesion area mask with an accuracy of 0.08mm.
[0028] Vibration modal compensation: The vocal cord vibration model is established based on the Navier-Stokes equation, and the amplitude deviation is calculated in real time; Correct image distortion and eliminate motion artifacts through affine transformation.
[0029] Principles of biofeedback and three-dimensional dynamic modeling: The annular array sensor 3011 monitors the thyroarytenoid / cricothyroid myoelectric signals (sampling rate 10kHz) and is linked to the fluid dynamics simulator 3021: Myoelectric-vibration synergistic analysis: Calculate the time-frequency characteristics of the electromyographic signal (wavelet energy entropy, Hjorth parameter) to identify abnormal muscle coordination; trigger the threshold alarm: when the electromyographic intensity exceeds 50μV, activate the sound and light warning.
[0030] Immersed boundary method simulation: The vocal cords are modeled as elastic membrane structures, and 3D vocal cord vibration animation is output, with simultaneous display of phase difference and vorticity distribution.
[0031] Multimodal feature fusion and AI diagnosis: Data processing module 400 integrates acoustic, imaging, and bioelectric trimodal data: Feature dimensionality reduction and alignment: Principal component analysis (PCA) compresses 2048-dimensional acoustic features (fundamental frequency perturbation, harmonic-to-noise ratio) to 512 dimensions; dynamic time warping (DTW) aligns imaging features (mucosal wave symmetry) with bioelectric time series data.
[0032] Dynamic deep learning model: EfficientNet-B4 backbone network 5011: extracts vocal cord surface texture and morphological features; BiGRU timing unit 5012: processes vocal cord vibration cycle sequence; incremental update mechanism 5013: loads new cases from the disease atlas database 502 every week and uses transfer learning to optimize the model.
[0033] Closed-loop rehabilitation training and remote collaboration: Virtual reality interface 601 combined with real-time biofeedback to build a closed-loop training system; Freedom interactive training: Patients observe their own vocal cord movement simulation through a VR helmet and adjust their vocal posture according to deviation prompts; the system records training data and generates a personalized plan.
[0034] 5G remote diagnosis and treatment: The remote diagnosis and treatment interface 603 connects to the hospital server 604 via the SRv6 protocol to transmit 4K image streams; it supports multi-expert consultation mode, improving the diagnostic accuracy of grassroots hospitals.
[0035] Hardware implementation and optimization: Detachable Silicone Sensing Ring 800: Medical-grade silicone conforms to the curve of the throat, and built-in flexible circuit enables 8-channel electromyography acquisition; Portable processing terminal 900: Modular slot 901 supports expansion of GPU accelerator cards, which increases inference speed and adapts to different working environments.
[0036] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A vocal cord problem recognition and feedback system for ENT, characterized by: It includes a sound acquisition module (100), an image acquisition module (200), a biofeedback module (300), a data processing module (400), an AI diagnosis module (500) and a rehabilitation guidance module (600); The sound collection module (100) comprises a microphone array (101) and an adaptive noise reduction unit (102); The image acquisition module (200) is equipped with an endoscope camera (201) and an image enhancement processor (202); The biofeedback module (300) integrates a laryngeal electromyographic sensor (301) and a three-dimensional motion simulator (302); The data processing module (400) performs multimodal feature extraction and fusion; The AI diagnostic module (500) includes a dynamically updated deep learning model (501); The rehabilitation guidance module (600) generates a visual biofeedback training program; The data processing module (400) is connected to other modules via a data bus (700).
2. The ENT vocal cord problem identification and feedback system according to claim 1, characterized in that: The biofeedback module (300) comprises: Annular array sensor (3011), monitoring thyroarytenoid / cricothyroid coordination; A fluid dynamics simulator (3021) that uses the immersed boundary method to generate three-dimensional vocal fold vibration animations.
3. The ENT vocal cord problem identification and feedback system according to claim 1, characterized in that: The image enhancement processor (202) includes: HSV color space conversion unit (2021) for vascular texture enhancement; UNet segmentation unit (2022), which performs lesion area identification; The image correction unit (2023) performs motion compensation based on the vocal cord vibration modal equation.
4. The ENT vocal cord problem identification and feedback system according to claim 1, characterized in that: The deep learning model (501) comprises: Backbone network (5011), used for feature extraction; Bidirectional gated recurrent unit (5012), processing time series data; The dynamic updating unit (5013) synchronously updates the vocal cord disease atlas database (502) every week.
5. The ENT vocal cord problem identification and feedback system according to claim 1, characterized in that: The rehabilitation guidance module (600) comprises: A virtual reality interface (601) displays a simulated animation of vocal cord movement; A real-time threshold alarm unit (602) is activated when the myoelectric intensity exceeds a preset value; Remote diagnosis and treatment interface (603), connected to the hospital central server (604).
6. The ENT vocal cord problem identification and feedback system according to claim 1, characterized in that: Also included is a hardware structure, which includes: a detachable silicone sensing ring (800) with an integrated biofeedback module (300); A portable processing terminal (900) has a built-in modular slot (901) for expanding functional units.
Citation Information
Cited By
Multi-modal learning throat rehabilitation action quantitative evaluation system and method thereof
CN120827371A
Otological disease prediction method based on multi-modal data fusion and confidence evaluation
CN121096599A
Ear disease prediction method based on multi-modal data fusion and confidence evaluation
CN121096599B