A method and system for monitoring otolaryngological symptoms based on multimodal data fusion
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-20
- Publication Date
- 2026-08-14
AI Technical Summary
[0002]在耳鼻喉疾病辅助诊断场景中,临床常需通过多场景多类型设备采集内窥镜图像、生理信号、环境数据及患者个体数据以实现症状评估,但这些设备及采集数据存在显著的异构性问题,主要体现为不同场景设备的测量精度差异可达 ±15% 以上(如医院 4K 内窥镜与家用微型探头内窥镜的图像清晰度偏差),同一参数的多源数据偏差最高达 20%(如不同品牌血氧仪的血氧测量值差异)
1、通过构建 三维可信度评估模型,结合初步可信度等级校准,实现对多模态数据质量的精准分层,能有效区分高、中、低可信数据并赋予差异化融合权重,避免低质量数据干扰融合结果,确保高可信数据(如医院专业内窥镜图像)在后续融合中发挥核心作用,为多模态数据融合提供可靠的质量分层依据;同时,采用 DTW 算法结合个体时序偏差修正因子,适配不同患者(如儿童、老年人)的生理节律差异,实现多模态数据精准时序同步,再通过医学术语词典映射消除跨模态语义偏差,解决多模态数据 “时间不同步、语义不统一”的关键问题,为深度融合提供时空一致、语义统一的特征基础,显著提升后续特征融合的准确性与有效性。
Smart Images

Figure CN121512476B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical monitoring technology, and in particular to a method and system for monitoring otolaryngological symptoms based on multimodal data fusion. Background Technology
[0002] In auxiliary diagnostic scenarios for ENT diseases, clinicians often need to collect endoscopic images, physiological signals, environmental data, and individual patient data using various devices across multiple scenarios to assess symptoms. However, these devices and the collected data exhibit significant heterogeneity, primarily manifested in measurement accuracy differences of over ±15% between devices in different scenarios (e.g., image clarity discrepancies between hospital 4K endoscopes and home-use miniature probe endoscopes), and deviations of up to 20% for the same parameter across multiple sources (e.g., differences in pulse oximeter measurements from different brands). Since device performance and acquisition scenarios directly determine data reliability and feature validity, the consistency and stability of heterogeneous data vary greatly under the same diagnostic needs, making accurate feature extraction extremely difficult.
[0003] Traditional diagnostic aids often employ a standardized data processing workflow, neglecting the characteristics of equipment environments and individual patient differences. This can lead to issues such as distorted data due to excessive errors (e.g., pollen concentration data from outdoor environmental sensors affected by electromagnetic interference) and ineffective correlation of data due to lack of calibration (e.g., semantic discrepancies between home endoscopic images and clinical terminology). Furthermore, the complexities of patient environments, including temperature and humidity fluctuations in hospitals, homes, and outdoor settings, variations in electromagnetic interference from different devices (e.g., interference from mobile phone signals to physiological sensors in home settings), changes in patient circadian rhythms, and differences in operational procedures, introduce signal noise and parameter measurement biases, further degrading data quality.
[0004] Furthermore, existing systems lack the ability to adapt to the individual pathological characteristics of patients in real time, cannot dynamically compensate for feature shifts caused by genetic history and underlying diseases, and feature extraction relies heavily on fixed model parameters, resulting in a lag in response to individual differences. This makes it difficult to achieve accurate extraction of 1024-dimensional deep fusion feature vectors and disease judgment in complex diagnostic and treatment scenarios, which seriously affects the accuracy of ENT disease diagnosis and the pertinence of clinical recommendations. Summary of the Invention
[0005] This invention effectively improves the accuracy of diagnosing ear, nose, and throat diseases by using noise stratification suppression, individual correlation redundancy feature screening, and dynamic evaluation of multimodal data credibility.
[0006] The technical solution proposed in this invention is: a method for monitoring otolaryngological symptoms based on multimodal data fusion, the method comprising: High-precision endoscopic images, physiological signals, environmental data, and individual patient data are collected to form a multi-scenario, multi-modal raw dataset. After noise identification, suppression, and credibility pre-evaluation, a denoised dataset and a preliminary credibility level are obtained. By combining the denoised dataset with individual patient data, an individual-adaptive core feature set is obtained through three-dimensional feature evaluation and dynamic threshold screening. After credibility stratification and attention weight allocation, a hierarchical core feature set and credibility evaluation matrix are obtained. Based on the hierarchical core feature set and individual patient data, after time synchronization and semantic alignment processing, combined with individual data and initial weights, a deep fusion feature vector and GNN association map are generated through feature embedding, GNN construction and GAN training. If the feature vector is missing, the data is completed by combining the confidence assessment matrix and individual patient data through VAE. The clinical gold standard and individual data are combined, and the clinically calibrated feature vector is obtained through three-dimensional mapping and Kappa coefficient verification. By combining clinical calibration feature vectors, efficacy data, GNN correlation graphs, and model parameters, and through difference comparison and transfer learning optimization, based on calibration features, optimized models, and diagnostic criteria, the system outputs symptom assessment reports, clinical recommendations, and visualization results.
[0007] Preferably, the specific process for obtaining the denoised dataset is as follows: To address interference, corresponding noise type identification algorithms are selected based on the characteristics of different modal data. For image data, a combination of peak signal-to-noise ratio (PSNR) and noise histogram analysis is used. For physiological signal data, a combination of PSNR and temporal feature determination is used. For environmental data, data fluctuation coefficient analysis is employed. Algorithm parameters are adjusted to adapt to multiple scenarios based on data characteristics. Modality-specific suppression algorithms are selected based on the noise characteristics of each modality. For electronic noise in images, a dark current correction algorithm is used, and a calibration curve is established based on scene temperature differences. For motion artifacts in physiological signals, an adaptive Kalman filter algorithm is used, and the motion noise variance is adjusted based on the motion monitoring characteristics of wearable devices. For electromagnetic interference in environmental data, an adaptive recursive filter algorithm is used, and the filter coefficients are adjusted based on data fluctuation characteristics.
[0008] Preferably, the specific process for obtaining the individual-adaptive core feature set is as follows: After acquiring the denoised multimodal dataset and individual patient data, the SHAP value analysis algorithm was used to convert different modal features into numerical feature vectors. The SHAP value of each feature was calculated, and the average of the absolute values was taken to obtain the feature contribution. Features with high contribution were initially screened. Combined with the feature-symptom association table annotated by clinical experts, the Pearson correlation coefficient was used to analyze the association strength between the screened features and individual patient data, retaining features with strong association and clinical annotations indicating strong or moderate association. A dynamic threshold algorithm was introduced to determine the screening threshold based on the feature variance contribution, eliminating redundant features below the threshold. All screening results were integrated to generate corresponding core feature sets for different individual patients.
[0009] Preferably, the specific process for obtaining the hierarchical core feature set and the credibility evaluation matrix is as follows: The system acquires individual-adapted core feature sets and preliminary credibility levels for each modality. It then calculates scores based on data acquisition scenarios, equipment quality levels, and operator qualifications to assess data source reliability. Data integrity is assessed by calculating scores based on the total number of core features a modality should contain versus the actual number of features retained, and the total data recording duration versus the effective duration. Data consistency is assessed by calculating scores based on the number of other modalities associated with the current modality and the clinical logical matching degree of each associated modality. The final credibility score is calculated by combining the scores from these three dimensions with their corresponding weights. This final credibility score is then calibrated and labeled with the final credibility level. Initial fusion weights are assigned to different credibility levels. These initial weights are calculated using a linear mapping formula and then normalized to ensure the sum of all modality weights equals 1.
[0010] Preferably, the specific process for obtaining the deep fusion feature vector is as follows: The process involves acquiring a multimodal feature set after temporal synchronization and semantic alignment, patient individual data, and initial weights for each modality. Patient individual data is then transformed into feature embedding vectors via a fully connected neural network. These feature embedding vectors are concatenated with the aligned multimodal features to form a total feature matrix, which is then input into a graph neural network (GNN) constructed from the association graph. Edge weights are calculated using the initial weights for each modality, and nonlinear relationships between nodes are mined and node features are updated via the GNN message passing mechanism. The node features output by the GNN are concatenated into an initial fused feature vector, which is then input into a generative adversarial network (GAN) for adversarial training. The generator and discriminator are trained alternately. When the discriminator accuracy stabilizes within a preset range, the generator outputs a feature vector with enhanced discriminative power. Finally, the GNN association graph and the GAN training results are integrated to obtain a deep fused feature vector containing multimodal association information and individual difference information.
[0011] Preferably, the specific process of VAE data completion is as follows: The process involves acquiring deep fusion feature vectors, a credibility assessment matrix, and individual patient data. Missing data in the deep fusion feature vectors are detected, the total number of missing features is counted, and the modality to which the missing features belong is located. Based on the credibility assessment matrix, and combining the normalized weights of each modality with the final credibility score, data compensation weights for each modality are calculated. The complete feature vectors and individual patient data are embedded into a variational autoencoder (VAE). The encoder maps the input to a latent distribution, and the decoder reconstructs the complete feature vectors from the latent variables. Individual difference compensation constraints are introduced into the VAE loss function. The VAE model is trained using an optimizer. After the reconstruction error on the validation set stabilizes, the input deep fusion feature vectors containing missing data are given, and the decoder outputs compensation data corresponding to the missing positions. The compensation data is embedded into the original deep fusion feature vectors to complete the data dimensions, resulting in the completed deep fusion feature vectors.
[0012] Preferably, the specific process for obtaining the clinical calibration feature vector is as follows: We acquire the completed deep fusion feature vector, ENT clinical gold standard data, and individual patient data to construct a mapping relationship between the fusion features and the clinical gold standard. Individual feature embedding vectors are introduced as adjustment factors, and the mapping is implemented through a fully connected neural network. The fusion features and individual embedding vectors are concatenated into a 1280-dimensional vector, which is then processed through the input layer to the hidden layer and the hidden layer to the output layer to output the symptom prediction probability distribution. A weighted Kappa coefficient is introduced to evaluate the mapping consistency. Based on the confusion matrix, the weighted observation consistency rate and the weighted expected consistency rate are calculated using linear weights to obtain the weighted Kappa coefficient. With the goal of maximizing the weighted Kappa coefficient, the Adam optimizer is used to adjust the mapping function parameters. The loss function is set as the negative weighted sum of cross-entropy loss and the Kappa coefficient. Training stops when the Kappa coefficient stabilizes above 0.8. The hidden layer output of the optimized mapping function is used as the clinically calibrated fusion feature vector.
[0013] Preferably, the specific process for obtaining the symptom assessment report and the clinical recommendations is as follows: The process involves obtaining clinically calibrated fusion feature vectors, optimized model parameters, and ENT symptom diagnostic criteria. The fusion feature vectors are then input into the optimized GNN classification model. Combining these with the symptom diagnostic criteria, a Softmax output layer is used to obtain the probability distribution of each disease type, determining whether the patient has an ENT disease and its specific type. Disease severity scores are calculated by integrating disease-related key indicators from the fusion features with individual patient data, and the severity level is determined based on these scores. Adaptability scores for different treatment plans are calculated using associated data from the optimized model, and recommendations are prioritized based on these scores. Finally, supplementary nursing recommendations are provided based on individual patient data.
[0014] The present invention also provides an ENT symptom monitoring system based on multimodal data fusion, the system being used to execute the aforementioned ENT symptom monitoring method based on multimodal data fusion.
[0015] The present invention also provides a computer-readable storage medium storing a computer program, which is executed by a processor to implement the aforementioned method for monitoring otolaryngological symptoms based on multimodal data fusion.
[0016] The beneficial effects of this invention are: 1. By constructing a three-dimensional credibility assessment model and combining it with preliminary credibility level calibration, we can achieve precise stratification of multimodal data quality. This effectively distinguishes between high, medium, and low credibility data and assigns differentiated fusion weights, avoiding interference from low-quality data in the fusion results. This ensures that high-credibility data (such as professional endoscopic images from hospitals) plays a core role in subsequent fusion, providing a reliable basis for quality stratification in multimodal data fusion. Simultaneously, by employing the DTW algorithm combined with individual temporal deviation correction factors, we adapt to the physiological rhythm differences of different patients (such as children and the elderly), achieving precise temporal synchronization of multimodal data. Furthermore, by mapping through a medical terminology dictionary, we eliminate cross-modal semantic deviations, solving the key problems of "time asynchrony and semantic inconsistency" in multimodal data. This provides a spatiotemporally consistent and semantically unified feature foundation for deep fusion, significantly improving the accuracy and effectiveness of subsequent feature fusion.
[0017] 2. Using the clinical gold standard as a benchmark, a dynamic mapping relationship between fusion features and the gold standard is constructed. Individual characteristics are introduced as adjustment factors, and the mapping consistency is verified by weighted Kappa coefficient. This transforms the technically-based fusion features into data that conforms to clinical diagnostic logic, ensuring that the consistency between the fusion features and the clinical gold standard remains stable above 0.8, meeting the data adaptability requirements for clinical applications. Simultaneously, by comparing the disease progression trend predicted by the fusion features with the actual efficacy data, the difference value is calculated. For cases where the deviation exceeds the standard, transfer learning is used to fine-tune the weights of the GNN association graph and the VAE to compensate for the model parameters. This allows the model to continuously adapt to the pathological characteristics of different patients (such as patients sensitive to specific drugs or patients with underlying diseases), continuously improving the model's adaptability and accuracy for individual patients, and ensuring that the model can consistently output reliable results in clinical applications.
[0018] 3. For different scenarios such as hospitals, homes, and outdoors, scene-adaptive acquisition devices (such as 4K ultra-high-definition endoscopes for hospitals and miniature probe endoscopes for home use) are adopted. Combined with acquisition parameter optimization algorithms (parameters for low-brightness pediatric endoscopes and anti-electromagnetic interference design for outdoor use) and data synchronization triggering mechanisms, efficient and high-quality acquisition of multimodal data in multiple scenarios is achieved, providing comprehensive and scenario-specific raw data for the solution. At the same time, differentiated noise identification and suppression algorithms are adopted for different modal data characteristics (images, physiological signals, environmental data) to accurately eliminate interference from equipment noise, environmental noise, and human noise, thereby improving data quality. In addition, based on SHAP values, clinical correlation annotations, and dynamic threshold screening, redundant features irrelevant to individual conditions (such as pollen concentration data of non-allergic patients) are eliminated, reducing data dimensionality and computational complexity, ensuring that the features input to the fusion model are highly correlated with individual conditions, and laying a high-quality and highly correlated feature foundation for subsequent deep fusion and symptom judgment. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart of an ENT symptom monitoring method based on multimodal data fusion according to the present invention. Detailed Implementation
[0021] The following description is intended to disclose the present invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art. The basic principles of the invention defined in the following description can be applied to other embodiments, modifications, improvements, equivalents, and other technical solutions that do not depart from the spirit and scope of the invention.
[0022] It is understood that the term "a" should be understood as "at least one" or "one or more," that is, in one embodiment, the number of an element can be one, while in another embodiment, the number of the element can be multiple, and the term "a" should not be understood as a limitation on the number.
[0023] like Figure 1As shown, scene-adaptive acquisition devices are used (professional hospital endoscopes use 4K ultra-high-definition optical endoscopes equipped with 5-10mm adjustable depth-of-field lenses, supporting precise focusing on the nasal mucosa / tympanic membrane area, with a frame rate ≥30fps to capture dynamic lesions; simple home endoscopes use a 0.3-inch miniature probe design with a diameter ≤5mm to reduce user discomfort, and have a built-in LED supplementary lighting module with adjustable brightness (10-100 lux), suitable for non-professional home operation scenarios; wearable physiological sensors include a wrist-worn ECG and blood oxygen monitoring device (sampling rate 250Hz, blood oxygen measurement range 70%-100%, heart rate measurement range 30-250bpm), a chest-attached respiratory sensor (detection accuracy ±1 time / minute), and a head sleep posture sensor (using a triaxial accelerometer, resolution 0.01g)), combined with acquisition parameter optimization algorithms (in pediatric scenarios, the parameters of the professional hospital endoscope are adjusted to...). Brightness 50 lux, exposure time 0.02s, strong light mode disabled to avoid irritating children's delicate mucous membranes; in outdoor scenarios, the environmental sensor uses an anti-electromagnetic interference algorithm, and reduces the interference of complex outdoor electromagnetic environment on data through electromagnetic shielding coating and digital filtering technology. The temperature and humidity sensor measurement accuracy is improved to ±0.5℃, ±3%RH. The pollen concentration sensor adopts laser scattering method, detection range 0-5000 grains / m³, resolution 1 grain / m³) and data synchronization trigger mechanism (in hospital scenarios, endoscope and physiological monitoring equipment achieve millisecond-level synchronization through wired Ethernet, timestamp error ≤10ms; in home / outdoor scenarios, devices are linked through low power Bluetooth 5.2 protocol, synchronization delay ≤50ms, and timed collection reminders can be set on the patient terminal (customizable interval 1-24 hours). If the collection is missed, a supplementary collection prompt is triggered, and the supplementary collection data is automatically marked with "supplementary collection").
[0024] In a hospital clinic setting, professional medical staff operate the equipment to collect: high-precision endoscopic images (nasal mucosa images focus on the nasal septum and inferior turbinate region, with 3-5 images from different angles collected from each nasal cavity; tympanic membrane images focus on the umbilicus and light reflex region of the tympanic membrane to ensure the complete presentation of the tympanic membrane morphology), and simultaneously collect physiological signals such as electrocardiogram (continuous collection for ≥5 minutes, including at least 3 complete cardiac cycles), blood oxygen (continuous monitoring for ≥10 minutes, recording data once every 10 seconds), and respiratory rate (continuous monitoring for ≥5 minutes). During the collection process, the patient's immediate symptoms at the time of consultation (such as whether they are accompanied by nasal congestion or ear pain) are recorded simultaneously.
[0025] In a home setting, guide patients or their families on how to operate the device to collect: home endoscopic images (collected at fixed times every day, with standardized operating guidelines provided before collection, such as the nasal endoscope needing to be inserted parallel to the nasal septum), sleep posture data (continuous collection at night, with a sampling interval of 5 seconds, recording the duration and frequency of switching between supine, lateral, and prone positions), and record daily symptom changes (such as the degree of sore throat and the amount of nasal secretions) and medication information (drug name, dosage, and time of administration) through the patient's terminal.
[0026] In outdoor settings, portable devices are used to collect real-time data on: pollen concentration (collected once per hour, with simultaneous recording of weather conditions (sunny / rainy / cloudy)), temperature and humidity (collected once every 30 minutes), and air quality (focusing on PM2.5 (range 0-500 μg / m³, accuracy ±5 μg / m³), PM10 (range 0-1000 μg / m³, accuracy ±10 μg / m³), and SO2 (range 0-500 μg / m³, accuracy ±5 μg / m³). Real-time location data is obtained through the patient's terminal's positioning function (positioning accuracy ≤10 meters, using BeiDou and GPS dual-mode positioning to ensure positioning stability in complex outdoor environments). At the same time, the patient's outdoor activity type (such as walking, sitting, and exercise) and exposure duration are recorded.
[0027] Meanwhile, structured self-reported data is collected through patient terminals (smartphones / tablets supporting iOS 13.0 and above, Android 9.0 and above): Symptom descriptions use a grading scale (e.g., sore throat is divided into 0-4 levels, 0 no pain, 1 mild pain, 2 moderate pain, 3 severe pain, 4 excruciating pain), medication records require the generic name of the drug, manufacturer, single dose, and time of administration (accurate to the minute); individual baseline data is collected in a form, including age (accurate to years), gender (male / female / other), genetic history (e.g., whether there is a family history of allergic rhinitis or otitis media, the kinship and age of onset must be filled in), underlying diseases (e.g. whether there is diabetes or hypertension, the diagnosis time and current treatment plan must be filled in), and allergy history (e.g. whether there is an allergy to pollen, dust mites, or specific drugs).
[0028] The final output is a multi-scenario, multi-modal raw dataset. Each data file is accompanied by two labels: "Collection Scene Label" indicates the specific scene type (hospital clinic - ENT / otolaryngology / laryngology, home - bedroom / living room, outdoor - park / street) and collection time (accurate to the second); "Raw Quality Label" is rated from 1 to 5 points (5 points is the best) based on three dimensions: data integrity (e.g., whether the endoscope image fully presents the target area, whether there are any breaks in physiological signals), clarity (e.g., whether the endoscope image is blurry or has noise), and accuracy (e.g., whether the self-reported data is complete, whether the positioning data is valid). At the same time, a data collection quality report is generated, indicating the compliance rate of the data in each scene (e.g., hospital endoscope image compliance rate ≥95% for 5 points, home scene data integrity compliance rate ≥85%). This provides a quality assessment basis for the subsequent raw data noise stratification and suppression stage, forming the data source foundation for the entire solution process.
[0029] After obtaining a multi-scenario, multimodal raw dataset containing hospital / home / outdoor scene labels, raw quality labels (1-5 points), and covering endoscopic images (nasal mucosa / tympanic membrane, 4K / 0.3-inch miniature probe acquisition), physiological signals (ECG / blood oxygen / respiration, sampling rate 250Hz / detection accuracy ±1 time / minute), environmental data (temperature and humidity / pollen concentration / air quality, temperature and humidity accuracy ±0.5℃ / ±3%RH, pollen concentration detection range 0-5000 grains / m³), individual basic data (age / genetic history, etc.), and self-reported data, different noise type identification algorithms were first selected for interference localization based on the characteristics of different modalities of data. Among them, the peak signal-to-noise ratio (PSNR) and noise histogram analysis method was used for image data, the signal-to-noise ratio (SNR) and time domain feature determination method was used for physiological signal data, and the data fluctuation coefficient (CV) analysis method was used for environmental data.
[0030] In image-based data, the PSNR calculation for professional hospital endoscope images uses standard laboratory mucosa / tympanic membrane images as noise-free reference images. For home-use simple endoscope images, due to their non-professional operation characteristics, the reference images are adjusted to qualified demonstration images acquired in a home setting (critical structural integrity ≥80%). PSNR calculation parameters are defined as follows: Peak signal-to-noise ratio (in dB). The mean square error between the original image and the reference image (unit: squared gray value). The maximum grayscale value of the image (255 for an 8-bit image, unitless) is calculated using the following formula: This formula can be used to quantify image noise intensity, such as the noise level calculated from nasal mucosa images obtained using a home endoscope. 30dB, combined with the Gaussian distribution characteristics of the histogram, indicates "medium noise - mainly electronic noise".
[0031] Based on a large amount of clinical image data statistics, PSNR is divided into three noise intensity ranges: low noise range. ≥35dB (noise intensity ≤10%), medium noise range 25dB≤ <35dB (noise intensity 10%-20%), high noise range <25dB (noise intensity >20%) provides a basis for subsequent noise type determination. Next, based on the different mechanisms by which different noise types affect image grayscale values, a correlation rule between PSNR range and noise type is established: Electronic noise, manifested as subtle grayscale fluctuations across the entire image domain (affecting an amplitude of ±5 grayscale values), corresponds to a PSNR mostly in the low noise range (≥35dB), such as the PSNR of nasal mucosa images acquired by professional hospital endoscopes, which is mostly 38-45dB with MSE <15; Salt-and-pepper noise, manifested as isolated pure white / pure black pixels (dispersed influence range), corresponds to a PSNR mostly in the medium noise range (25-35dB), such as the PSNR of home endoscopes during Bluetooth transmission due to interference from mobile phone signals, which is mostly 28-32dB with MSE approximately 20-50; Halo noise, manifested as large-area gradient grayscale changes (grayscale difference between edge and center 30-50), corresponds to a PSNR mostly in the high noise range (<25dB), such as the PSNR of home scenes with tilted lighting angles, which is mostly around 20. -24dB, MSE>50; then considering special cases (such as low-intensity salt-and-pepper noise entering the low-noise range), it is necessary to combine the noise histogram shape to complete the accurate judgment: the electronic noise histogram has a symmetrical Gaussian distribution (gray values are concentrated near the mean), such as a hospital tympanic membrane image PSNR=40dB, the histogram peak is located near the gray value 120, and the 115-125 range has a smooth bell-shaped curve; the salt-and-pepper noise histogram has a "main distribution + discrete peaks" shape (peaks at 0 and 255), such as a home nasal mucosa image PSNR=30dB, the main body of the histogram is in the 80-150 range, and the peak height at 0 and 255 is 1 / 5 of the main peak value; the halo noise histogram has an asymmetrical gradient distribution ("left low right high" or "left high right low"), such as an outdoor nasal mucosa image PSNR=22dB, the histogram 50-80 range pixels account for 60%, and the 150-180 range pixels account for 30%.
[0032] To address potential "mixed noise" situations in clinical practice (such as the simultaneous presence of electronic noise and salt-and-pepper noise within the medium noise range), it is necessary to use " The principle of "priority + primary noise determination" is used to process the mixed noise. First, the magnitude of the mixed noise's impact is determined, with the core basis being "primary noise intensity + secondary noise proportion": if the primary noise is in the low noise range ( If the noise level is ≥35dB and the proportion of secondary noise is <10% (e.g., the proportion of features corresponding to secondary noise in the histogram is <10%), then the mixed noise has a minimal impact and can be ignored. Only the primary noise needs to be processed. For example, in an endoscopic image from a hospital... =38dB (main noise is electronic noise), and the salt-and-pepper noise peak accounts for only 5% of the histogram. At this point, the mixed noise has no impact on the recognition of key structures in the image (such as nasal mucosal lesions), and the dark current correction algorithm can be directly applied as electronic noise; if the main noise is in the medium noise range (25dB≤...), If the noise level is <35dB and the secondary noise accounts for 10%-20%, then the mixed noise has a relatively small impact. When processing the primary noise, the secondary noise only needs to be slightly considered; there is no need to design complex algorithms specifically for the secondary noise. For example, consider a home endoscope image. =32dB (the main noise is electronic noise, accounting for 85%), and the secondary salt-and-pepper noise accounts for 10%. At this point, a light median filter with a 3×3 window can be superimposed on the dark current correction, and there is no need to use a strong filter with a 5×5 window (to avoid excessive smoothing and loss of details). If the main noise is in the medium noise range and the secondary noise accounts for >20%, or the main noise is in the high noise range ( If the noise level is <25dB, the mixed noise has a significant impact and cannot be ignored. A combined strategy of "dedicated processing of primary noise + auxiliary suppression of secondary noise" is required, for example, in an outdoor endoscopic image. =24dB (the main noise is halo noise, accounting for 70%), and the secondary electronic noise accounts for 25%. At this point, halo noise needs to be corrected first through gradient inverse compensation (formula). , =0.4), and then corrected by dark current (based on the ambient temperature of 28℃). =15) Suppress electronic noise and ensure that after noise reduction Improved to ≥28dB, key structure recognition rate ≥90%.
[0033] After determining the noise type and the impact of mixed noise on image data, noise identification is performed using a combination of signal-to-noise ratio (SNR) and temporal features for physiological signal data, and data fluctuation coefficient (CV) analysis for environmental data. In physiological signal data, blood oxygen signals in outdoor scenes are subject to motion artifact interference; therefore, the effective frequency band needs to be adjusted from 0.1-10Hz to 0.1-5Hz during SNR calculation (to reduce the impact of motion noise on the SNR value). The SNR calculation formula is defined as follows: ( For effective frequency band power, (Power in the noise frequency band); In the environmental data category, the pollen concentration sensor in a rainy scenario, The calculation time window was extended from 1 minute to 3 minutes (to avoid instantaneous fluctuations caused by raindrop impacts being misinterpreted as electromagnetic interference), and the definition was... The calculation formula is ( For standard deviation, (mean).
[0034] After noise type identification, mode-specific suppression algorithms are selected based on the noise characteristics of different modes. For electronic noise in image data, a dark current correction algorithm is used. Considering the temperature difference between hospital and home environments (hospital clinic temperature 22±2℃, home environment temperature 18-28℃), a temperature-dark current calibration curve is first established (temperature... The temperature range is 15-30℃, with one point taken for every 1℃, corresponding to the dark current grayscale value. Define the correction parameters: The original pixel grayscale value (0-255). The corrected pixel grayscale value (0-255) is obtained using the following formula: For example, when the temperature in a home environment is 25℃, the calibration curve can be found... =12, original pixel grayscale value =150, then after correction =150 12=138; For motion artifacts in physiological signal data (easily generated during walking in outdoor scenes), an adaptive Kalman filter algorithm is used to adapt to the motion monitoring characteristics of wearable devices, and the filter parameters are defined as follows: The variance of motion noise (unit: ), =0.01 (basic noise variance, fixed value). =0.005 (motion influence coefficient, fixed value) Motion acceleration obtained from a triaxial accelerometer (unit: ),but When walking =1.2 Calculated =0.01+0.005×1.2=0.016, and the filtering matrix is dynamically adjusted by this variance to achieve precise suppression of motion artifacts; for electromagnetic interference of environmental data (which is easily generated when outdoors near high-voltage power lines), an adaptive recursive filtering algorithm is adopted to adapt to the fluctuation characteristics of environmental data, and the filtering parameters are defined as follows: α is the filtering coefficient (0-1). The current data fluctuation coefficient (%) =5% (low noise threshold) =15% (high noise threshold), then such as pollen concentration data =10%, calculated as follows =0.5, then use the filtering formula ( This is the current sampled value. (The filtered value from the previous moment) is used to obtain the denoised environmental data. .
[0035] After denoising the data for each modality, the noise suppression effect index for each modality is calculated to generate a "noise suppression effect label," and image class effect parameters are defined: ( The peak signal-to-noise ratio after denoising. (Peak signal-to-noise ratio before denoising) The structural similarity index (0-1) is used, such as the tympanic membrane image of a hospital before denoising. =28dB, after noise reduction =36dB, then =8dB, through ( , The mean grayscale value after denoising. , The standard deviation of the front grayscale after noise reduction. For covariance, , Calculated =0.93, then the label is "endoscopic image - =8dB-SSIM=0.93-Noise Removal Rate 85%-Good Detail Preservation; Define physiological signal effect parameters: ( The signal-to-noise ratio after denoising. (Signal-to-noise ratio before denoising) ( , (Noise power before and after denoising), such as blood oxygen signal before denoising. =15dB, after noise reduction =26dB, =11dB, =0.02W =0.0018W, then =91%, tagged as "blood oxygen signal" =11dB- =91% - Motion artifact removal rate 91%; Define environmental data class effect parameters: ( (The fluctuation coefficient after denoising) For example, temperature and humidity data before noise reduction =12%, after noise reduction =4%, =8%, ≈66.7%, labeled as "temperature and humidity data - =8%- =66.7% - Electromagnetic interference removal rate 66.7%.
[0036] The credibility of the denoised data is then evaluated using a data credibility pre-evaluation module, defining the evaluation parameters: The total credibility score is 1-5. The noise reduction effect is scored (out of 5 points: ≥90% noise reduction rate gets 5 points, 80%-90% gets 3 points, <80% gets 1 point). Data integrity is scored (out of 5 points: 5 points for missing data < 5%, 3 points for missing data between 5% and 10%, and 1 point for missing data > 10%). Scores are awarded for key information retention (out of 5 points: 5 points for complete key features, 3 points for partially complete features, and 1 point for missing features), with the following weights: =0.5、 =0.3、 =0.2, then the formula for the total credibility score is: Based on this, the credibility levels are specifically divided as follows: the high credibility level score range is 4.0 ≤ ≤5.0 points, the criterion for judgment is that the noise reduction effect score must be met. ≥3 points and at least one dimension scored 5 points, key information retention score =5 points, and then directly enter the individual correlation-based redundant feature screening stage and be assigned a high weight (e.g., 0.7), for example, ECG signal data from a hospital. =5 points =5 points =5 points, calculated as follows =5.0 points; Medium confidence level score range A score of <4.0 is determined based on the noise reduction performance score. ≥1 point for no dimension, and points for retaining key information. =3 points, and subsequent cross-validation with high-confidence data of the same modality is required (bias ≤5% to pass) and assigned a medium weight (e.g., 0.3), for example, a pollen concentration data =3 points =5 points =5 points, calculated as follows =3.5 points; Low credibility level score range <2.5 points, the criterion is the score for meeting the noise reduction effect. ==1 point, data integrity score =1 point, Key Information Retention Score A score of 1 is awarded for any condition, and the data is subsequently marked as "data to be reviewed." If the data can be repaired after manual inspection, it is returned for reprocessing; otherwise, it is discarded. (Example: an image from a home endoscope.) =1 point =3 points =3 points, calculated as follows =2.0 points.
[0037] Through the above process, a denoised multimodal dataset containing denoised data for each modality, "noise suppression effect labels", and "preliminary confidence level" is finally obtained, providing a high-quality, hierarchical data foundation for the subsequent screening of individual correlation-based redundant features.
[0038] In obtaining images containing denoised endoscopic images (with "noise suppression effect label", such as "...") =8dB - Noise removal rate 85%); denoised physiological signals (e.g., ECG SNR=26dB); denoised environmental data (e.g., pollen concentration). =4%), and a denoised multimodal dataset with each modality data labeled with a preliminary confidence level (high / medium / low, low confidence data has been marked and removed), and individual patient data covering patient age (A, unit: years), genetic history (H, 0=no related genetic history, 1=family history of allergic rhinitis / otitis media), underlying diseases (D, 0=no underlying diseases, 1=diabetes / hypertension, etc.), and medication history (M, such as "long-term use of antihistamines" is marked as 1, otherwise 0). First, the contribution of each modality feature to ENT symptoms is calculated using the SHAP value analysis algorithm. This algorithm quantifies the marginal contribution of each feature to symptom judgment through game theory principles. First, different modal features are uniformly converted into numerical feature vectors. The denoised endoscopic images are then processed through a ResNet50 network to extract 2048-dimensional texture features (denoted as ). Denoising physiological signals to extract mean heart rate ( (Unit: bpm), Standard deviation of blood oxygen saturation ( 10-dimensional time-series features (unit: %), etc. ), daily average pollen concentration extracted from noise-reduced environmental data ( Unit: grains / Temperature and humidity fluctuation values ( Unit: ℃; Five-dimensional environmental characteristics (unit: %RH) and others (denoted as %RH). At the same time, patient individual data is transformed into 4-dimensional individual characteristics (denoted as...). To form the total feature matrix (Total 2048+10+5+4=2067 dimensions).
[0039] definition Value calculation parameters: For the first In the nth sample The SHAP value of each feature (unit: none; the larger the absolute value, the greater the contribution of the feature to the symptoms). For feature subset The predicted values (based on a trained ENT symptom classification model, such as XGBoost, which outputs symptom probability values). For excluding the first A subset of features If the total number of features is , then the number of individual features is . The formula for calculating the value is: This formula is used to calculate all features. After setting the value, for each feature across all samples The absolute values are taken as the mean to obtain the feature contribution. ( (This refers to the total number of samples), for example, calculating the "nasal mucosal congestion texture features" of endoscopic images. of 0.85, mean heart rate of =0.62, pollen concentration of =0.78, patient's genetic history of =0.91, preliminary screening Features with a value ≥0.5 (high contribution) are excluded, while features with a value < (such as features with small fluctuations in temperature and humidity) are removed. , =0.32).
[0040] After calculating the feature contribution, and combining the "feature-symptom association table" annotated by clinical experts (e.g., "nasal mucosal congestion texture feature - allergic rhinitis" association strength is marked as "strong", "mean heart rate - snoring" association strength is marked as "medium", "ambient temperature and humidity - otitis media" association strength is marked as "weak"), the association strength between the filtered features and individual patient data is analyzed using Pearson correlation coefficient analysis, and the association strength parameter is defined: For the first The first feature and the second Pearson correlation coefficient for individual data (such as age A, genetic history H) (range [-1, 1], the closer the absolute value is to 1, the stronger the association). Features The mean, For individual data The mean of the two values is calculated using the following formula: For example, calculating the "nasal mucosal congestion texture features". With genetic history H =0.72 (strong correlation), with age A =0.23 (weak correlation); pollen concentration With genetic history H =0.81 (strong association), r=0.15 (weak association) with the underlying disease D, combined with the "strong association" annotation in the "feature-symptom association table", retain the strong association with individual data (| Features with a correlation ≥ 0.4 and clinically labeled as "strong / moderate association" are excluded; weak associations are eliminated. Features with a correlation of |<0.4 and clinically labeled as "weakly correlated" (such as "mean heart rate") With all individual data | | < 0.3, and clinically labeled "weak association").
[0041] To accurately remove redundant features, a dynamic thresholding algorithm is introduced. The filtering threshold is determined based on the feature variance contribution, and redundancy judgment parameters are defined. For the first The variance contribution of each feature (i.e., the proportion of the total variance of all features contributed by that feature, in %). Features variance ( (the number of features currently retained), then By calculating all retained features Sort by size from largest to smallest and set dynamic thresholds. (The mean of the variance contributions of all features), excluding < Redundant features, for example, currently retaining 100 features. =500, including "nasal mucosal congestion texture features" of =35, =7%; "Pollen concentration" of =42, =8.4%; "Mean Heart Rate" of =12, =2.4%, calculated as follows =1% (The mean is low here due to the large number of features; it needs to be adjusted based on clinical practice, and the final value should be set at 1%). If the percentage is 3%, then remove it. =2.4% < Redundant features, retain Its characteristics.
[0042] Finally, by integrating all the above screening results, an individual-fit core feature set is generated—for example, for patients with a family history of allergic rhinitis (H=1), the core feature set includes "nasal mucosal congestion texture features". "Daily average pollen concentration" "History of antihistamine use" (M), "Standard deviation of blood oxygen saturation" ( =3.5%≥ With genetic history The core feature set includes 15 features such as H=0.65; for patients with chronic pharyngitis without a family history of the disease (H=0), the core feature set includes "pharyngeal mucosal texture features". "Dietary habit data" (Extracted from self-reported data), 12 features including "Underlying Disease D" are selected, and a feature screening report is generated. The report should record the screening process for each feature in detail, including: Values and individual data value, Values and reasons for removal / retention (e.g., mean heart rate) because < (Features deemed redundant are removed) to ensure that the screening process is traceable and verifiable. The resulting individual-adaptive core feature set and feature screening report will be directly used for subsequent dynamic evaluation and stratification of multimodal data credibility, laying the foundation for improving the accuracy of multimodal fusion.
[0043] In order to obtain a feature set (such as "nasal mucosal congestion texture features") that includes different patient individuals (such as allergic rhinitis patients with a genetic history of H=1, chronic pharyngitis patients with H=0, etc.) "Daily average pollen concentration" Medication history (M), etc., each characteristic is accompanied by... Value, correlation strength r-value with individual data, and variance contribution. After establishing the individual-adaptive core feature set and the preliminary credibility level (high / medium / low) of each modality data, the credibility of each modality data in the core feature set is evaluated from three dimensions: data source reliability, data integrity, and data consistency. The quantitative indicators and calculation methods of the three dimensions are first clarified.
[0044] For the data source reliability dimension, define the evaluation parameters: The source reliability score is 0-10, with a higher score indicating a more reliable source. For data collection scenarios (3 points for collection using professional hospital equipment, 2.5 points for collection using community medical point equipment, 2 points for collection using certified home equipment, and 1 point for collection using uncertified outdoor equipment), The equipment quality level is calculated as follows: 4 points for meeting medical-grade standards (such as hospital endoscopes, certified as Class III medical devices by NMPA), 3 points for meeting quasi-medical-grade standards (such as portable ECG monitors in community health points, certified as Class II medical devices by NMPA), 2 points for meeting consumer-grade health standards (such as home pulse oximeters, registered as medical devices), and 1 point for having no clear standard (such as ordinary outdoor environmental sensors). The calculation formula is as follows: (3 points for professional medical personnel, 2.5 points for community medical personnel, 2 points for trained users, and 1 point for ordinary users). For example, 4K endoscopic images in hospitals ( =3 points =4 points =3 points) =3+4+3=10 points; Data from portable ECG equipment at community health points ( =2.5 points =3 points =2.5 points) =2.5+3+2.5=8 points; Home-based certified pulse oximeter data ( =2 points =2 points =2 points) =2+2+2=6 points; Outdoor ordinary pollen concentration sensor data ( =1 point =1 point =1 point) =1+1+1=3 points.
[0045] For the data integrity dimension, an evaluation is conducted based on the missing data of each modality in the core feature set, and evaluation parameters are defined as follows: Completeness score (0-10 points). This represents the total number of core features that this mode should include. The number of core features that are actually completely retained (features without missing values or outliers are considered complete). The total data recording time is in hours, such as 0.5 hours for a single acquisition of endoscopic images and 24 hours of continuous monitoring of physiological signals. The effective data duration (duration of records without missing or broken records) is calculated using the following formula: For example, the endoscopic image modality of a patient (for allergic rhinitis) should include eight core features such as "nasal mucosal texture, degree of congestion, and type of secretion". =8), actually 7 were completely retained ( =7), single data collection duration 0.5 hours and no disconnection ( = =0.5), then =5×(7 / 8)+5×(0.5 / 0.5)=4.375+5=9.375 points; Physiological signal modalities (ECG and blood oxygen) should include 5 core characteristics such as "mean heart rate, standard deviation of blood oxygen, and heart rate variability". =5), actually 3 were completely retained ( =3), with an effective duration of 18 hours in 24-hour monitoring ( =24, =18), then =5×(3 / 5)+5×(18 / 24)=3+3.75=6.75 points.
[0046] For the data consistency dimension, the evaluation parameters are defined by assessing the clinical logical matching degree between different modalities of the same patient: Consistency score (0-10 points). The number of other modalities that are clinically associated with the current modality (e.g., endoscopic images and symptom self-reports, environmental data and allergy history are all associated) is recorded. =2), For the first The matching degree of each associated modality (1 point for complete clinical logic, 0.5 points for partial logic, and 0 points for no logic) is calculated using the following formula: For example, an endoscopic image of a patient with allergic rhinitis showed "severe nasal mucosal congestion + watery discharge" ( =2), self-reported data record "nasal congestion and runny nose symptoms" ( =1), environmental data shows "pollen concentration > 3000 grains / (Matched with allergy history, =1), then =10×(1+1) / 2=10 points; Another patient's laryngoscopic image showed "redness and swelling of the pharyngeal mucosa" ( =1), but self-reported "no sore throat or foreign body sensation" (which does not conform to clinical logic). =0), then =10×0 / 1=0 points.
[0047] The final credibility score is calculated by combining the scores from the three dimensions, and the parameters are defined as follows: The final credibility score (0-10 points, with a maximum of 10 points) is calculated. =0.4、 =0.3、 =0.3 represent the weights for source reliability, integrity, and consistency, respectively (determined based on the Delphi method of 50 ENT specialists, with source reliability having the greatest impact on data quality, hence the highest weight). The calculation formula is as follows: The initial confidence level is used for calibration (0.5 points are added to the calculated result if the initial level is "high"; no adjustment is made if the initial level is "medium"; 0.5 points are deducted if the initial level is "low"; and 0 points are taken if the score is lower than 0 after calibration). For example, hospital endoscopic images. =10 points =9.375 points =10 points, preliminary grade "high", then =0.4×10+0.3×9.375+0.3×10+0.5=4+2.8125+3+0.5=10.3125 points (rounded to 10 points); Community medical center electrocardiogram data =8 points =8.5 points (4 / 5 core features retained, effective duration 22 / 24 hours) =7 points (partially matching the hospital's basic ECG data), preliminary grade "medium", then =0.4×8+0.3×8.5+0.3×7=3.2+2.55+2.1=7.85 points; Home pulse oximeter data =6 points =6.75 points =7.5 points (matching the heart rate data trend), initial level "medium", then =0.4×6+0.3×6.75+0.3×7.5=2.4+2.025+2.25=6.675 points; Outdoor pollen concentration data =3 points =5.2 points (features retained 3 / 4, effective duration 16 / 24 hours) =4 points (partially matches allergy symptoms), preliminary grade "low", then =0.4×3+0.3×5.2+0.3×4 0.5 = 1.2 + 1.56 + 1.2 0.5 = 3.46 points.
[0048] Based on the final credibility score Final credibility level: 8≤ ≤10 points is "highly reliable" (good data quality, can be directly used for core fusion computing), 5 ≤ A score of <8 indicates "moderately reliable" (data quality is good, but needs to be used after cross-validation with other modalities). A score of <5 indicates "low reliability" (poor data quality, requiring manual review or restricted use). For example, the hospital endoscopy images mentioned above are "high reliability," the community medical center electrocardiogram data and home pulse oximeter data are "medium reliability," and the outdoor pollen concentration data are "low reliability."
[0049] Assign initial fusion weights to different levels, and define parameters: The initial fusion weights (0-1, the total weights must satisfy the condition that the sum of all modal weights is 1, after normalization) are used. For the "high confidence" level, a linear mapping formula is adopted. (Mapped to 0.7-1.0, higher scores have higher weights); "Medium Confidence" level, formula is... Mapped to 0.4-0.6); "Low Trust" level, formula is... (Mapped to 0.1–0.3), e.g., “high-confidence” endoscopic images =10 points =0.7 + 0.3 × 210 8 = 0.7 + 0.3 = 1.0; "Zhongxin Trust" community healthcare electrocardiogram data. =7.85 points, Wmid=0.4+0.2×37.85 5 = 0.4 + 0.2 × 0.95 ≈ 0.59; "Zhongkexin" pulse oximeter data =6.675 points, =0.4 + 0.2 × 36.675 5 = 0.4 + 0.2 × 0.558 ≈ 0.51; "Low reliability" pollen concentration data ==3.46 points, =0.1+0.2×53.46=0.1+0.138≈0.24. Normalize the above weights (1.0+0.59+0.51+0.24=2.34, divide each weight by 2.34) to get the normalized weights: endoscopic image 0.43, ECG data 0.25, blood oxygen data 0.22, pollen concentration data 0.10.
[0050] The final core feature set after credibility stratification is formed (each feature is labeled with the final credibility level, initial fusion weight, and normalized weight, such as "nasal mucosal congestion feature - high credibility - =0.43") and a credibility assessment matrix, which is arranged with "modal name" as the row and "assessment dimension ( / / / The column is named " / level / weight", for example, the row record for "Hospital Endoscopy" is " =10, =9.375, =10, =10, level = high reliability, normalized weight = 0.43” provides a clear hierarchical weight basis for subsequent multimodal data time-series synchronization and deep fusion, ensuring that high reliability data plays a core role in fusion, while low reliability data is only used as an auxiliary reference.
[0051] After obtaining a core feature set containing various modal features (such as "nasal mucosal congestion features" from hospital endoscopy, "mean heart rate" from community medical electrocardiogram data, "standard deviation of blood oxygen" from home oxygen saturation, and "daily average" of outdoor pollen concentration) and each feature labeled with a final confidence level (high / medium / low) and normalized fusion weights (such as 0.43 for endoscopy and 0.25 for electrocardiogram), and individual patient data covering patient age (A, unit: years), circadian rhythm type (R, 0 = early to bed and early to rise, 1 = late to bed and late to rise), and the influence coefficient of underlying diseases (D, such as D=0.15 for diabetic patients and D=0 for those without underlying diseases), the Dynamic Time Warping (DTW) algorithm is first used to align the timestamps of different modal data to adapt to the temporal differences in the multimodal data of this scheme (such as endoscopy being a single acquisition (timestamp)). (Unit: seconds), physiological signals are continuously acquired (timestamps) Environmental data is collected periodically (at 1-second intervals) using timestamps. (30-minute intervals) Define the core parameters of the DTW algorithm: For a time series of data of a certain modality (such as electrocardiogram data, (Number of data points) The time series of reference modal data (selecting the physiological signal with the highest acquisition frequency as the reference) (Number of data points) for The Middle Data points and The Middle The time difference distance of each data point, among which for Timestamp (unit: seconds). for The timestamp (unit: seconds) is used to calculate the time difference distance formula. DTW(X,Y) is the minimum time-normalized distance between two sequences, calculated using the following formula: in To normalize the path length, it must satisfy the following conditions: =1, =1, = , = The boundary conditions are solved by dynamic programming to obtain the cumulative distance matrix. ( ), ultimately resulting in Minimal time alignment path, such as timestamps of a single endoscope acquisition. =3600 seconds (1 hour) and ECG sequence =3600 seconds of data point alignment, environmental data =3600 seconds (1 hour on the hour) is also aligned to this time point, achieving initial time synchronization.
[0052] Calculate the individual time series deviation correction factor by combining individual patient data, adjust the alignment threshold of the DTW algorithm, and define the correction parameters: This is the individual time series deviation correction factor (unitless, value range 0.8-1.2). The factor is determined by the patient's age (A), circadian rhythm (R), and underlying disease coefficient (D), and the calculation formula is as follows: 30 is the baseline age, when When the age is greater than 30, increasing age leads to greater fluctuations in the timing of physiological signals (such as increased delay in heart rate acquisition in elderly patients). Follow Increased and elevated; for patients with R=1 (late to bed and late to rise type), the physiological rhythm deviation needs to be reduced, so multiplied by (1 R); underlying diseases can exacerbate the time-series bias, hence the addition of D. For example, a 35-year-old (A=35), an early-to-bed, early-to-rise type (R=0), and a diabetic patient (D=0.15) =1+0.005×(35 30)×1+0.15=1+0.025+0.15=1.175; 25 years old (A=25), late-to-bed, late-to-rise type (R=1), no underlying diseases (D=0) =1+0.005×(25 30)×0+0=1. Based on Adjust DTW alignment threshold (Original threshold) =5 seconds, after adjustment When the time difference between two data points ≤ If the alignment is successful, the normalization path is determined to be alignable; otherwise, it needs to be re-optimized and normalized. For example... =1.175 patients =5×1.175=5.875 seconds, allowing for a larger time difference alignment, adapting to the acquisition delay characteristics of elderly patients.
[0053] After completing temporal alignment, visual and textual features are linked through a medical terminology dictionary to eliminate cross-modal semantic bias. First, an ENT-specific medical terminology dictionary is constructed, establishing a mapping relationship between visual feature descriptions (such as "pale and edematous nasal mucosa" in endoscopic images) and standardized medical terms (such as "typical signs of allergic rhinitis" in ICD-11), defining semantic association parameters. Visual features Text features The semantic similarity (0-1, 1 indicating a perfect match) is calculated using cosine similarity. and Term vectors (trained based on Word2Vec, dimension) =100) is denoted as , The calculation formula is: when When the value is ≥0.8, it is determined to be a semantic match, and an association is established; when... When <0.8, it is necessary to use a "feature-symptom association table" annotated by clinical experts for further verification; when A value <0.5 indicates semantic bias, requiring correction of the textual feature description (e.g., correcting the patient's self-reported "itchy nose" to the standardized term "itchy nose symptom"). For example, the visual feature "tympanic membrane retraction" (…). ) and text features “signs related to otitis media” )of =0.92, judged as semantic match; patient self-reported "dry throat" ( ) and the visual feature "dry throat mucosa" )of =0.65, which is determined to be a match after verification using the related table.
[0054] Finally, the temporal alignment and semantic association results are integrated to generate a multimodal feature set after temporal synchronization and semantic alignment. Each feature in this feature set is accompanied by a unified timestamp after alignment (e.g., "t=3600 seconds"), a semantic matching label (e.g., "semantic matching - similarity 0.92"), and a confidence weight (using the previous normalized weight). A synchronization alignment report is also generated, which should record in detail the DTW alignment results for each modality (e.g., "endoscopic and ECG data"). =2.3 seconds, less than =5.875 seconds, alignment successful); Individual timing deviation correction factor and calculation basis, semantic similarity The distribution (e.g., "85% of feature pairs have a semantic similarity ≥ 0.8") ensures that the entire process is traceable and verifiable. This feature set will be directly used in subsequent individual difference feature embedding and deep feature fusion stages, providing a spatiotemporally consistent and semantically unified data foundation for improving the accuracy of multimodal data fusion.
[0055] After obtaining a multimodal feature set containing various modal features (such as "nasal mucosal congestion features", "mean heart rate", and "daily average pollen concentration"), each feature accompanied by a unified timestamp (such as "t=3600 seconds"), semantic matching labels (such as "semantic matching - similarity 0.92"), and normalized initial weights (such as endoscopy 0.43, ECG 0.25, blood oxygen 0.22, pollen concentration 0.10), and after obtaining individual patient data covering patient age (A, unit: years), genetic history (H, 0=none, 1=yes), underlying diseases (D, 0=none, 1=yes), and medication history (M, 0=none, 1=yes), the individual patient data is first converted into feature embedding vectors to adapt to the input requirements of the GNN algorithm. Define the individual feature embedding parameters: The original individual feature vector (dimension 4, , The embedded individual feature vector (with the same dimension as the multimodal features, denoted as ) =256), embedded using a fully connected neural network, with ReLU activation function for the hidden layer and the output layer calculation formula as follows: in This is the weight matrix from the input layer to the hidden layer (dimension 4×128). This is the hidden layer bias vector (dimension 128). This is the weight matrix from the hidden layer to the output layer (dimension 128×256). This is the output layer bias vector (dimension 256). For example, data from a specific patient. (35 years old, with a family history of genetic disorders, underlying medical conditions, and a history of medication use), calculated as follows: (Each element takes values from 0 to 1), achieving dimensional unification between individual features and multimodal features.
[0056] The embedded individual feature vector Aligned multimodal features ( Features of endoscopic images, dimension 256; Physiological signal characteristics, dimension 256; The environmental data features (dimension 256) are concatenated to obtain the total feature matrix. (Dimension 4×256), then input into a Graph Neural Network (GNN) to construct an association graph between modal features and individual features. Define the parameters of the GNN association graph: For the correlation map, where A set of nodes (containing individual feature nodes) and multimodal feature nodes , , (4 nodes in total) A set of edges (representing the relationships between nodes, such as...) (Indicates the relationship between individual features and image features) This is the edge weight matrix (reflecting the correlation strength). Edge weight calculation incorporates the initial weights for each mode. (Endoscopy) =0.43, ECG =0.25, blood oxygen =0.22, pollen concentration =0.10), the formula is: in For feature vectors and Cosine similarity (measures the linear relationship between features). , For nodes , The initial weights corresponding to the modality, such as individual feature nodes. Image feature nodes edge weight (Individual features have no initial weight, set to 1), if =0.85, then =20.85×1.43≈0.607. The nonlinear relationships between nodes are mined using the message passing mechanism of GNN (emphasizing GAT attention mechanism), and node features are updated accordingly. ,in For the updated node Features For nodes The set of neighboring nodes, Attention weights (derived from edge weights) (obtained by normalization) is the activation function (ReLU), and W is the message passing weight matrix (dimension 256×256). As the bias vector (dimension 256), the updated association graph is finally obtained. This enables a deep correlation between modal features and individual features.
[0057] To enhance the discriminative power of the fused features, the node features output by the GNN are concatenated into an initial fused feature vector. (Dimension 1024), input to Generative Adversarial Network (GAN) for adversarial training. Define GAN training parameters: generator G input random noise z (dimensional 128) and initial fused features Output fake fusion features Discriminator D inputs true fused features (Generated from "feature-symptom" matching samples annotated by clinical experts) and The generator loss function outputs the probability P of the feature's authenticity (0-1, 1 representing authenticity). and discriminator loss function They are respectively: ; ; in This is the regularization coefficient (set to 0.5 to ensure consistency between the generated features and the initial fused features). The mean squared error is used. The generator and discriminator are trained alternately (Adam optimizer, learning rate 1e-4). Training stops when the discriminator's accuracy stabilizes at around 50% (unable to distinguish between real and fake features). At this point, the generator outputs... That is, the deep fusion feature vector after enhancing discriminativity. (Dimension 1024).
[0058] Finally, by integrating the GNN association graph with the GAN training results, a deep fusion feature vector is obtained. (Each element reflects the feature's contribution to the discrimination of ear, nose, and throat symptoms) and the updated GNN association graph. (Including node update features, edge weights, and attention weights), and simultaneously recording the loss curve during training. , The convergence trend and feature discrimination accuracy (such as improving the discrimination accuracy of allergic rhinitis symptoms to 92%), this deep fusion feature vector will be directly used in the subsequent intelligent compensation of missing data, providing core feature support for improving the accuracy of ENT symptom judgment.
[0059] The goal is to obtain a deep fusion feature vector containing deep fusion information from various modalities (such as "association features between nasal mucosal congestion and genetic history" and "interaction features between heart rate fluctuations and pollen concentration," with 1024 dimensions). Record the confidence scores for each modality (e.g., endoscopy). =10 points, ECG =7.85 points) and a credibility assessment matrix with normalized weights (endoscopy 0.43, ECG 0.25, blood oxygen 0.22, pollen 0.10), and individual patient data covering patient age (A), circadian rhythm (R), and underlying disease coefficient (D), firstly, data missingness in the deep fusion feature vector is detected, and missing data labeling parameters are defined: For the first Missing markers for each feature ( =1 indicates missing. =0 indicates completeness), count the total number of missing features. And locate the mode to which the missing feature belongs (e.g., if the 32nd to 64th dimension features are missing, the corresponding ECG mode).
[0060] Data compensation weights are assigned to each modality based on the credibility assessment matrix, and the compensation weight parameters are defined as follows: For the first Compensation weights for each modality ( =1 corresponds to an endoscope. =2 corresponds to electrocardiogram (ECG). =3 corresponds to blood oxygen saturation. =4 corresponds to pollen), the calculation formula is: ,in For modality Normalized weights, For modality The final credibility score. For example, endoscopes. =0.43、 =10, ECG =0.25、 =7.85, blood oxygen =0.22、 =6.675, pollen =0.10、 =3.46, then the numerators are 0.43×10=4.3, 0.25×7.85=1.9625, 0.22×6.675=1.4685, and 0.10×3.46=0.346, respectively. The sum is 4.3+1.9625+1.4685+0.346=8.077. Therefore, =4.3 / 8.077≈0.532, =1.9625 / 8.077≈0.243, =1.4685 / 8.077≈0.182, =0.346 / 8.077≈0.043, ensuring that high-confidence modalities (such as endoscopes) play a leading role in the compensation process.
[0061] The complete feature vector and individual patient data are input into the variational autoencoder (VAE) model. Missing data is generated in the latent space by incorporating individual variability compensation constraints. The VAE model parameters are defined as follows: the encoder E receives the input features... (Complete feature set) and individual embedding vectors (Using the 256-dimensional embedding results from the previous section) Mapping to the latent distribution ,in It is a mean vector (dimension 128). The variance vector (dimension 128); the decoder D extracts data from the latent variables. Reconstruct the complete feature vector The loss function includes reconstruction loss and KL divergence loss, and also introduces an individual difference compensation constraint term: ; in, ( Features Belonging mode, For modality The feature mean is used to ensure that the generated values of missing features converge to the mean of the same mode and are controlled by the weight of the high-confidence mode. The constraint latent distribution approximates the standard normal distribution; ( (These are feature vectors generated using only individual data), ensuring that the generated features reflect individual differences. =1、 =0.5 is a hyperparameter.
[0062] The VAE model is trained using the Adam optimizer (learning rate 5e-5). When the reconstruction error on the validation set decreases to a stable value (e.g., <0.01), the input feature vector containing missing features is used. The decoder output The corresponding missing position in the middle ( The value of =1) is the compensation data. The compensated data is embedded into the original fused feature vector to obtain the completed deep fused feature vector. The complete features retain their original values ( when Missing features are replaced with compensation values ( when =1).
[0063] Finally, a missing data compensation report is generated, recording the location of the missing features and their respective modalities (e.g., "Features in dimensions 32-64 (ECG modality) are missing, totaling 32 features"), and the compensation weight for each modality. The VAE model training metrics (e.g., the final value of the reconstruction loss is 0.008), the deviation between the compensated data and the mean of the same modality (e.g., the mean deviation of the ECG compensated features is 0.03), and the integrity verification results of the completed feature vector (e.g., "all 1024-dimensional features are completed, and the integrity is 100%). This completed deep fusion feature vector will be directly used in the subsequent intelligent assessment model of ENT symptoms, providing complete feature input to improve the accuracy of the assessment.
[0064] After obtaining complete deep fusion feature vector information containing 1024-dimensional complete feature information (such as "interaction features between nasal mucosal congestion and pollen concentration" and "association features between heart rate fluctuations and genetic history"), the result is a complete version of the deep fusion feature vector information (such as "interaction features between nasal mucosal congestion and pollen concentration" and "association features between heart rate fluctuations and genetic history"). This includes clinical gold standard data covering diagnostic results of typical ENT symptoms (e.g., allergic rhinitis: 0 = none, 1 = mild, 2 = moderate, 3 = severe; otitis media: 0 = none, 1 = acute, 2 = chronic). After obtaining individual patient data including patient age (A), symptom frequency (F, times / week), and treatment history (T, 0 = no treatment, 1 = drug treatment, 2 = surgical treatment), a mapping relationship between fusion features and clinical gold standards is first constructed. Individual features are then introduced as moderating factors to optimize the mapping process. Mapping function parameters are defined as follows: The mapping function from fusion features and individual characteristics to the clinical gold standard (where C is the number of symptom categories, e.g., C=4 corresponds to the 4 levels of allergic rhinitis) is expressed as follows: ; in, To fuse features with individual embedding vectors (using the 256-dimensional approach mentioned earlier) The concatenated vector (dimension 1280). This is the weight matrix from the input layer to the hidden layer (1280×512). For hidden layer bias (512-dimensional). This is the weight matrix (512×C) from the hidden layer to the output layer. For output layer bias (C-dimensional). The symptom prediction probability distribution (C-dimensional, element-wise sum of 1) is mapped to the output. The moderating effect of individual characteristics is achieved through the embedding vector. Achievement, for example, in elderly patients (A=60) It will increase the weight of features related to "chronic pharyngitis", making the mapping more consistent with age-related clinical characteristics.
[0065] To quantify the consistency between the mapping relationship and the clinical gold standard, a weighted Kappa coefficient is introduced as an evaluation indicator, and a consistency parameter is defined: The weighted Kappa coefficient (ranging from -1 to 1, where 1 indicates complete agreement) is calculated based on the confusion matrix M (C×C dimension, where M[i][j] represents the number of samples whose actual class is i and whose predicted class is j): ; in, For weighted observation consistency rate ( The total number of samples, The weight matrix uses linear weights. (emphasizing that minor differences between adjacent levels have little impact on consistency). Weighted expected consistency rate ( For the first Total of lines, For the first (The sum of the columns). For example, for allergic rhinitis (C=4), if a certain sample set... =0.85、 =0.5, then (0.85 0.5) / (1 0.5) = 0.7, indicating that the height is consistent.
[0066] To maximize Optimize the mapping function parameters for the target (using the Adam optimizer, learning rate 1e-4, and loss function is cross-entropy loss and summation). (negative weighted sum), when Training stops when the consistency stabilizes above 0.8 (clinically acceptable high consistency). At this point, the mapping function outputs... The corresponding category (the category with the highest probability) is the clinically calibrated prediction result. The output of the hidden layer of the optimized mapping function is used as the clinically calibrated fusion feature vector. (512-dimensional) This vector retains the discriminative information of the original fusion features and incorporates clinical diagnostic logic.
[0067] Finally, a mapping consistency report is generated, recording the calculation process of the weighted Kappa coefficients (such as the specific values of the confusion matrix M). and The calculation results), and the consistency scores of different symptom categories (such as allergic rhinitis). =0.82, otitis media =0.79), the moderating effect of individual features on mapping (e.g., the predictive weight of chronic symptoms increases by 5% for every 10 years of age), and the difference in feature vectors before and after calibration (e.g., the cosine similarity increases from 0.75 to 0.88). This clinically calibrated fusion feature vector will be directly used for intelligent assessment and auxiliary diagnostic decision-making of ENT symptoms, providing clinicians with feature support that conforms to both multimodal data patterns and clinical logic.
[0068] After obtaining a clinically calibrated fusion feature vector containing 512 clinically adapted features (such as "correlation factors between calibrated nasal mucosal features and allergy history" and "efficacy-sensitive physiological signal features"), Record the degree of symptom improvement after treatment (e.g., nasal congestion score for allergic rhinitis: 3 points before treatment, 1 point after treatment, with a range of 0-4 points) as efficacy data. This includes node update features and edge weights (such as...) GNN association graph (=0.607) and the encoder weights of the VAE compensation model Decoder weights After obtaining the parameters, the patient's disease progression trend is first predicted based on the clinically calibrated fused feature vector, and the trend prediction parameters are defined as follows: To predict disease progression trends (e.g., "symptom score drops to 1.2 points one month after treatment," continuous values), a time series prediction model (based on LSTM, with input as...) is used. Characteristics of the treatment plan (e.g., drug type, dosage)), the output layer calculation formula is: ; in, This is the weight matrix of the LSTM output layer (dimension 128×1). For output layer bias (dimension 1). This is a concatenated vector of fusion features and treatment plan features (dimensions 512 + 32 = 544). For example, a patient... The "allergy-associated characteristic" value was 0.85. The predicted treatment was "antihistamine + saline flushing". =1.2 points (nasal congestion score).
[0069] Calculate the difference between the predicted trend and the actual efficacy data, and define the difference assessment parameters: This represents the trend difference value (percentage, reflecting the degree of prediction deviation). For actual post-treatment efficacy data (e.g., "actual nasal congestion score of 1.8 points one month after treatment"), the calculation formula is: ; in, The larger of the predicted and actual values should be used to avoid distorting the difference due to an excessively small denominator. For example... =1.2 points =1.8 points, =∣(1.2 1.8) / 1.8|×100%≈33.3%, if a difference threshold is preset. =20%, then > If the prediction deviation is deemed to be excessive, the model optimization process needs to be initiated.
[0070] When the difference exceeds a threshold, the patient's individual characteristics are extracted. (e.g., age 35, family history of disease 1, underlying disease 1) and fusion features Transfer learning is used to fine-tune the weights of the GNN association graph and the parameters of the VAE compensation model. For GNN association graph fine-tuning, the weight update parameters are defined as follows: This represents the GNN edge weight update amount. For difference loss (based on) calculate, Let ηG be the learning rate of the GNN (set to 0.01), then the update formula is: ; in, The edge weights before fine-tuning. For nodes and Cosine similarity of features (ensuring the weight update direction aligns with the feature association logic). For example, the edge weights in a GNN for "individual feature - ECG feature". =0.45, =0.333, =0.7, then =0.45+0.01×0.333×0.7≈0.452, which enhances the correlation between individual characteristics and electrocardiogram characteristics, and is suitable for the patient's pathological features.
[0071] For fine-tuning the parameters of the VAE compensation model, the parameter update objective is defined as minimizing the correlation error between the compensation data and the actual efficacy data. The VAE loss update parameters are defined as follows: ( =0.5 (where 0.5 is the difference loss weight), and the encoder weights of the VAE are updated using gradient descent. With decoder weights The updated formula is: ; in, Set the VAE learning rate to 0.005. For loss function pairs The gradient ensures that the compensation data generated by VAE after fine-tuning is more in line with the pathological characteristics corresponding to the actual therapeutic effect of the patient.
[0072] Repeat the above difference calculation and parameter fine-tuning process until five consecutive patients exceed the target. All < =20%, stop fine-tuning. At this point, the parameters of the iteratively optimized GNN fusion model (including the updated edge weight matrix) are obtained. Node features ), optimized VAE compensation model parameters ( , ).
[0073] Finally, a model optimization log is generated, which records in detail the patient information for each round of fine-tuning (e.g., "Patient ID: P001, age 35, underlying disease: diabetes") and the results of the difference calculation (e.g., " =33.3%, exceeding the standard), parameter update amount of GNN and VAE (such as "GNN edge weights") Updated from 0.45 to 0.452" "VAE decoder weights" The gradient value is -0.02”), and the change in the difference value after fine-tuning (e.g., “after fine-tuning”). =18.5%, meeting the standard), ensuring that the model optimization process is traceable and reproducible, and providing data support for subsequent model generalization improvement and clinical application.
[0074] After obtaining a clinically calibrated fusion feature vector containing 512 clinically adapted features (such as "post-calibrated nasal mucosal hyperemia-allergy association factor" and "tympanic membrane morphology-inflammatory risk feature"), Optimized GNN classification model parameters (including the updated edge weight matrix) Node features After configuring the VAE compensation model parameters and diagnostic criteria covering diseases such as "allergic rhinitis (ICD-11 code J30), chronic pharyngitis (J31.0), and secretory otitis media (H65.3)" (e.g., allergic rhinitis requires meeting two or more of the three criteria of "nasal itching + nasal congestion + pale and edematous nasal mucosa"), the clinically calibrated fusion feature vector is first input into the optimized GNN classification model to determine the presence and type of disease. The output parameters of the GNN classification model are defined as follows: The probability of the presence of a disease is 0-1, and >0.5 indicates the presence of a disease. For the first Probability distribution of disease types ( =1 corresponds to allergic rhinitis. =2 corresponds to chronic pharyngitis, =3 corresponds to secretory otitis media), and the classification output layer uses the Softmax function, calculated as follows:
[0075] in, For the first Classification weight vector for diseases (512 dimensions). For the first The classification bias (scalar) for each disease category is derived from the optimized GNN model parameters. For example, a patient... After input, the result is calculated. =0.82、 =0.15、 =0.03, > Based on the diagnostic criteria for otolaryngological symptoms (the patient's characteristics include "nasal itching correlation value of 0.9 and nasal congestion correlation value of 0.85", which meet the diagnostic criteria for allergic rhinitis), it was determined that "there is an otolaryngological disease, specifically allergic rhinitis".
[0076] Further assess the severity of the disease and define severity parameters: The severity score is calculated (0-10 points, 0-3 points for mild, 4-7 points for moderate, and 8-10 points for severe), combined with disease-related key indicators in the fusion features (such as the "nasal mucosal congestion degree feature" for allergic rhinitis). Symptom frequency characteristics The formula for calculating the relationship between the patient's individual data (age A, underlying disease coefficient D) and the patient's individual data (age A, underlying disease coefficient D) is as follows: ; in, , The value ranges from 0 to 1 (derived from the normalized value of the fused features). The underlying disease coefficient (0 = no underlying disease, 1 = with underlying disease). This is an age-adjusted term (0 for those under 60, linearly increasing to 1 for those 60-80, and 1 for those over 80). For example, a patient... =0.7、 =0.6、 =1、 =55, then =3×0.7+3×0.6+2×1+2×0=2.1+1.8+2=5.9 points, which is judged as "moderate allergic rhinitis".
[0077] Based on disease type, severity, and individual patient data, targeted clinical recommendations are generated, and recommendation adaptation parameters are defined: The recommendations are prioritized (levels 1-5, with level 1 being the core recommendation), and combined with the "feature-treatment response" correlation data in the optimized model (such as "antihistamine response features"). saline flushing response characteristics ), calculate the fit score for different treatment options: ; in, For the first The response features of the treatment options (from the edge weights of "treatment options - fusion features" in the GNN association graph, 0-1). This represents the predictive bias of the treatment regimen in similar patients (historical bias recorded by the optimized model, 0-0.2). For example, "antihistamines". =0.85、 =0.08, score =0.6×0.85+0.4×(1 0.08) = 0.51 + 0.368 = 0.878; "Surgical saline flushing" =0.7、 =0.1, score =0.6×0.7+0.4×0.9=0.42+0.36=0.78, suggestions are generated according to the score: Level 1 suggestion "Oral second-generation antihistamine (such as loratadine), 10mg once a day"; Level 2 suggestion "Nasal irrigation with saline twice a day, 200ml each time"; At the same time, considering the age (55 years old) and underlying diseases (none), the nursing suggestion "Avoid contact with allergens such as pollen, and monitor the condition of the nasal mucosa once a week" is added.
[0078] The final product generates a symptom assessment report and multimodal data visualization results. The report includes "disease presence (present / absent), specific type, severity score and basis, and targeted clinical recommendations (including priority)", such as "the patient has allergic rhinitis (moderate, 5.9 points), the core recommendation is oral antihistamines, based on an antihistamine suitability score of 0.878 (8% deviation for similar patients)". The visualization results use heatmaps to show the correlation strength between fusion features and disease types (e.g., "the correlation strength between nasal mucosal congestion features and allergic rhinitis is 0.82"), line graphs to show symptom trend predictions (e.g., "the severity score is expected to drop to 3.5 points after 1 month of treatment"), and radar charts to show the distribution of multimodal features (comparison of normalized values of endoscopic, physiological signals, and environmental data features), ensuring that clinicians can intuitively understand the judgment basis. This report and visualization results can be directly used for clinical auxiliary diagnosis and treatment plan formulation.
[0079] The processes described above with reference to the flowcharts in the embodiments disclosed in this invention can be implemented as computer software programs. The embodiments disclosed in this invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), it performs the functions defined in the methods of this application. It should be noted that the computer-readable medium described above in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wire segments, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless segments, wire segments, optical fibers, RF, etc., or any suitable combination thereof.
[0080] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0081] Those skilled in the art should understand that the embodiments of the present invention described above and shown in the accompanying drawings are merely examples and do not limit the present invention. The purpose of the present invention has been fully and effectively achieved. The functions and structural principles of the present invention have been shown and explained in the embodiments. Without departing from the stated principles, the implementation of the present invention may have any variations or modifications.
Claims
1. A method for monitoring otolaryngological symptoms based on multimodal data fusion, characterized in that, The method includes: High-precision endoscopic images, physiological signals, environmental data, and individual patient data are collected to form a multi-scenario, multi-modal raw dataset. After noise identification, suppression, and credibility pre-evaluation, a denoised dataset and a preliminary credibility level are obtained. By combining the denoised dataset with individual patient data, an individual-adaptive core feature set is obtained through three-dimensional feature evaluation and dynamic threshold screening. After credibility stratification and attention weight allocation, a hierarchical core feature set and credibility evaluation matrix are obtained. Based on the hierarchical core feature set and individual patient data, after time synchronization and semantic alignment processing, combined with individual data and initial weights, a deep fusion feature vector and GNN association map are generated through feature embedding, GNN construction and GAN training. If the feature vector is missing, the data is completed by combining the confidence assessment matrix and individual patient data through VAE. The clinical gold standard and individual data are combined, and the clinically calibrated feature vector is obtained through three-dimensional mapping and Kappa coefficient verification. By combining clinical calibration feature vectors, efficacy data, GNN correlation graphs and model parameters, and through differential comparison and transfer learning optimization, based on calibration features, optimized models and diagnostic criteria, symptom judgment reports, clinical suggestions and visualization results are output. The specific process for obtaining the hierarchical core feature set and the credibility evaluation matrix is as follows: The system acquires individual-adapted core feature sets and preliminary credibility levels for each modality. It then calculates scores based on data acquisition scenarios, equipment quality levels, and operator qualifications to assess data source reliability. Data integrity is assessed by calculating scores based on the total number of core features a modality should contain versus the actual number of features retained, and the total data recording duration versus the effective duration. Data consistency is assessed by calculating scores based on the number of other modalities associated with the current modality and the clinical logical matching degree of each associated modality. The final credibility score is calculated by combining the scores from these three dimensions with their corresponding weights. This final credibility score is then calibrated against the preliminary credibility level, and the final credibility level is labeled based on the final credibility score. Initial fusion weights are assigned to different credibility levels. These initial weights are calculated using a linear mapping formula and then normalized to ensure the sum of all modality weights equals 1. The specific process for obtaining the deep fusion feature vector is as follows: The process involves acquiring a multimodal feature set after temporal synchronization and semantic alignment, individual patient data, and initial weights for each modality. Individual patient data is then transformed into feature embedding vectors using a fully connected neural network. These feature embedding vectors are concatenated with the aligned multimodal features to form a total feature matrix, which is then input into a graph neural network (GNN) constructed from the association graph. Edge weights are calculated using the initial weights for each modality, and nonlinear relationships between nodes are mined and node features are updated via the GNN message passing mechanism. The node features output by the GNN are concatenated into an initial fused feature vector, which is then input into a generative adversarial network (GAN) for adversarial training. The generator and discriminator are trained alternately. When the discriminator accuracy stabilizes within a preset range, the generator outputs a feature vector with enhanced discriminative power. Finally, the GNN association graph and the GAN training results are integrated to obtain a deep fused feature vector containing multimodal association information and individual difference information. The specific process of VAE data completion is as follows: The process involves acquiring deep fusion feature vectors, a credibility assessment matrix, and individual patient data. Data gaps in the deep fusion feature vectors are detected, the total number of missing features is counted, and the modality to which each missing feature belongs is identified. Based on the credibility assessment matrix, combined with the normalized weights of each modality and the final credibility score, data compensation weights for each modality are calculated. The complete feature vectors and individual patient data are embedded into a variational autoencoder (VAE). The encoder maps the input to a latent distribution, and the decoder reconstructs the complete feature vectors from the latent variables. Individual difference compensation constraints are introduced into the VAE loss function. The VAE model is trained using an optimizer. After the reconstruction error on the validation set stabilizes, the input deep fusion feature vectors containing missing data are given, and the decoder outputs compensation data corresponding to the missing positions. The compensation data is then embedded into the original deep fusion feature vectors to complete the data dimensions, resulting in the completed deep fusion feature vectors. The specific process for obtaining the clinical calibration feature vector is as follows: We acquire the completed deep fusion feature vector, ENT clinical gold standard data, and individual patient data to construct a mapping relationship between the fusion features and the clinical gold standard. Individual feature embedding vectors are introduced as adjustment factors, and the mapping is implemented through a fully connected neural network. The fusion features and individual embedding vectors are concatenated into a 1280-dimensional vector, which is then processed through the input layer to the hidden layer and the hidden layer to the output layer to output the symptom prediction probability distribution. A weighted Kappa coefficient is introduced to evaluate the mapping consistency. Based on the confusion matrix, the weighted observation consistency rate and the weighted expected consistency rate are calculated using linear weights to obtain the weighted Kappa coefficient. With the goal of maximizing the weighted Kappa coefficient, the Adam optimizer is used to adjust the mapping function parameters. The loss function is set as the negative weighted sum of cross-entropy loss and the Kappa coefficient. Training stops when the Kappa coefficient stabilizes above 0.
8. The hidden layer output of the optimized mapping function is used as the clinically calibrated fusion feature vector. The specific process for obtaining the symptom assessment report and the clinical recommendations is as follows: The process involves obtaining clinically calibrated fusion feature vectors, optimized model parameters, and ENT symptom diagnostic criteria. The fusion feature vectors are then input into the optimized GNN classification model. Combining these with the symptom diagnostic criteria, a Softmax output layer is used to obtain the probability distribution of each disease type, determining whether the patient has an ENT disease and its specific type. Disease severity scores are calculated by integrating disease-related key indicators from the fusion features with individual patient data, and the severity level is determined based on these scores. Adaptability scores for different treatment plans are calculated using associated data from the optimized model, and recommendations are prioritized based on these scores. Finally, supplementary nursing recommendations are provided based on individual patient data.
2. The method for monitoring otolaryngological symptoms based on multimodal data fusion according to claim 1, characterized in that, The specific process of obtaining the denoised dataset is as follows: To address interference, corresponding noise type identification algorithms are selected based on the characteristics of different modal data. For image data, a combination of peak signal-to-noise ratio (PSNR) and noise histogram analysis is used. For physiological signal data, a combination of PSNR and temporal feature determination is used. For environmental data, data fluctuation coefficient analysis is employed. Algorithm parameters are adjusted to adapt to multiple scenarios based on data characteristics. Modality-specific suppression algorithms are selected based on the noise characteristics of each modality. For electronic noise in images, a dark current correction algorithm is used, and a calibration curve is established based on scene temperature differences. For motion artifacts in physiological signals, an adaptive Kalman filter algorithm is used, and the motion noise variance is adjusted based on the motion monitoring characteristics of wearable devices. For electromagnetic interference in environmental data, an adaptive recursive filter algorithm is used, and the filter coefficients are adjusted based on data fluctuation characteristics.
3. The method for monitoring otolaryngological symptoms based on multimodal data fusion according to claim 2, characterized in that, The specific process for obtaining the individual-adaptive core feature set is as follows: After acquiring the denoised multimodal dataset and individual patient data, the SHAP value analysis algorithm was used to convert different modal features into numerical feature vectors. The SHAP value of each feature was calculated, and the average of the absolute values was taken to obtain the feature contribution. Features with high contribution were initially screened. Combined with the feature-symptom association table annotated by clinical experts, the Pearson correlation coefficient was used to analyze the association strength between the screened features and individual patient data, retaining features with strong association and clinical annotations indicating strong or moderate association. A dynamic threshold algorithm was introduced to determine the screening threshold based on the feature variance contribution, eliminating redundant features below the threshold. All screening results were integrated to generate corresponding core feature sets for different individual patients.
4. An ENT symptom monitoring system based on multimodal data fusion, characterized in that, The system is used to execute the ENT symptom monitoring method based on multimodal data fusion as described in any one of claims 1-3.
5. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that is executed by a processor to implement the method for monitoring otolaryngological symptoms based on multimodal data fusion as described in any one of claims 1-3.
Citation Information
Patent Citations
Ear-nose-throat department symptom monitoring method and system based on multi-modal data fusion
CN119480157A
Computer-aided medical diagnosis system and method
CN120636760A
Multimodal computing-based early intelligent graded screening system for brain disease
WO2025175424A1