A health screening system and method based on multi-modal ear image dynamic timing
By employing multimodal image acquisition and data fusion technologies, combined with adaptive light source control and AI analysis, the system addresses the shortcomings of existing health screening systems in terms of single-modal imaging and insufficient data integration. This enables accurate and reliable health assessment and prediction, while enhancing user experience and testing convenience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG CANCER HOSPITAL
- Filing Date
- 2026-04-03
- Publication Date
- 2026-07-03
AI Technical Summary
Existing health screening systems rely on single-modal imaging, lack adaptive light source control, are easily affected by ambient light and user skin color, have insufficient data integration capabilities, cannot eliminate individual differences, have low reliability of assessment results, prominent false positive and false negative problems, poor user experience, cumbersome testing procedures, and are difficult to achieve at-home self-operation.
The system employs a multimodal image acquisition module combined with a high-resolution color camera and a high-speed multispectral imager, and is equipped with an adaptive light source control unit. It extracts microvascular pulsation waveforms and time-frequency characteristics, integrates ear images, wearable device data, and health questionnaire information, establishes an individualized health baseline, and performs comprehensive evaluation and prediction through an AI dynamic analysis engine to generate a visual report.
It has improved the accuracy and early identification capabilities of health screening, eliminated interference from individual differences and environmental factors, enhanced the credibility and accuracy of assessment results, realized the transformation from passive screening to proactive prevention, lowered the operational threshold, and enhanced the user experience.
Smart Images

Figure CN122337598A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of health and medical testing technology, and in particular to a health screening system and method based on dynamic temporal sequence of multimodal ear images. Background Technology
[0002] In the field of health screening, ear image analysis is widely used due to its non-invasiveness and ease of operation; however, existing technologies still have many significant shortcomings. Traditional ear health testing relies heavily on manual observation of the static morphology and color characteristics of acupoints, which is greatly influenced by subjective experience and can only capture static surface information, failing to reflect dynamic physiological changes such as subcutaneous microvascular pulsation, making it difficult to achieve early identification of subclinical conditions such as decreased vascular elasticity. While some image-based ear screening technologies have introduced imaging equipment, they often employ single-modal imaging methods, lacking adaptive light source control mechanisms. This makes them susceptible to the influence of ambient light and user skin color, resulting in low image signal-to-noise ratios, poor microvascular visualization, and difficulty in ensuring the accuracy of subsequent feature extraction. Furthermore, they have not established differentiated assessment standards for different age groups and regional populations, making it impossible to effectively eliminate interference from individual differences.
[0003] Existing health screening systems generally suffer from insufficient data integration capabilities and poor adaptability of assessment systems, making it difficult to meet the needs of precise and personalized health management. Most systems rely on only a single type of test data for risk assessment, failing to effectively integrate imaging features, wearable device physiological parameters, and health questionnaire information. This makes them prone to biased assessment results due to limited data dimensions and data distortion, and lacks data consistency verification and cross-checking mechanisms, resulting in low reliability. Furthermore, traditional systems often use population statistics as the baseline for health assessment, failing to distinguish between normal physiological fluctuations and abnormal pathological changes, leading to significant false positives and false negatives. The models also lack feedback and iteration mechanisms based on clinical diagnoses, making it difficult to continuously improve assessment accuracy. In addition, some screening systems present analysis results in complex formats, lacking visualized risk warnings and tiered alerts, resulting in a poor user experience. The cumbersome testing procedures also hinder home-based self-operation, limiting the widespread application of the technology. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention discloses a health screening system and method based on dynamic temporal sequence of multimodal ear images, which establishes differentiated assessment standards for people of different ages and regions and can effectively eliminate interference caused by individual differences.
[0005] This invention discloses a health screening system based on dynamic temporal sequence of multimodal ear images, comprising:
[0006] The multimodal image acquisition module is used to acquire visible light image sequences and multispectral image sequences of the user's ear area over a continuous time period.
[0007] The data processing and feature extraction module, whose input end is connected to the output end of the multimodal image acquisition module, is used to register and segment the received image sequence, extract the static features of each auricular acupoint region and the temporal features reflecting the dynamic changes of subcutaneous microvessels;
[0008] The multi-source data fusion module has a first input end connected to the output end of the data processing and feature extraction module, a second input end for accessing data from user-authorized external health monitoring devices, and a third input end for accessing health questionnaire information, which is used to perform consistency assessment and weighted fusion of multi-source data.
[0009] The personal health record database is bidirectionally connected to the data processing and feature extraction module and the multi-source data fusion module. It is used to store multimodal image data collected in each iteration, extracted feature vectors, external device synchronization data and user feedback information, and to establish an individualized health baseline.
[0010] The AI dynamic analysis engine module has its input end connected to the output end of the multi-source data fusion module and bidirectionally connected to the personal health record database. It is used to read individualized health baselines and write analysis results, including at least a disease risk assessment model based on static features, a change analysis model based on time-series features, and a comprehensive risk assessment model that integrates multi-source data.
[0011] The report generation and output module connects to the output of the AI dynamic analysis engine module and is linked to the personal health record database to read historical data. This data is used to generate a visual screening report and outputs risk level, abnormal area alerts, and health recommendations.
[0012] Furthermore, the multimodal image acquisition module includes a high-resolution color camera and a high-speed multispectral imager, and is equipped with an adaptive light source control unit, which dynamically adjusts the wavelength and intensity of the illumination source according to the real-time image quality to optimize the signal-to-noise ratio and microvascular imaging quality of the multispectral image sequence.
[0013] Furthermore, the data processing and feature extraction module also includes a dynamic feature extraction unit, which extracts the microvascular pulsation waveforms of each auricular acupoint region from the multispectral image sequence and performs time-frequency transformation on the pulsation waveforms to obtain the pulsation frequency, pulsation amplitude, waveform feature parameters, and pulsation phase difference between different regions as temporal features.
[0014] Furthermore, the multi-source data fusion module includes:
[0015] The weighted fusion unit is used to dynamically fuse ear image features, wearable device data and questionnaire information according to preset basic weights. The basic weights can be adjusted according to data quality and historical consistency.
[0016] The data consistency scoring unit is used to calculate the correlation coefficient between multi-source data and generate a consistency score. When the score is lower than the preset threshold, the fusion weight of the corresponding data is automatically reduced and a low confidence warning is marked in the report.
[0017] The cross-validation unit is used to detect whether there are significant contradictions between ear image features and wearable device data. When a contradiction occurs, it triggers a manual review prompt.
[0018] Furthermore, the personal health record database establishes an individualized health baseline based on the data collected from the user's first collection, and records the difference vector between the feature vector and the baseline for each subsequent collection.
[0019] The database also updates the model weights based on hospital diagnosis results reported by users, enabling personalized iteration.
[0020] Furthermore, the AI dynamic analysis engine module also includes a disease progression prediction model. This model uses a time-series neural network to predict the evolution trend of health risks within a specific future time period using the user's historical multimodal data sequence, and can output risk probability change curves and suggested intervention nodes.
[0021] Furthermore, it also includes a calibration mechanism module, which includes at least a skin color calibration unit, an age stratification unit, and a region adaptation unit;
[0022] The skin color calibration unit corrects skin color deviations in ear color analysis by acquiring multispectral images of the user's palm or forearm skin.
[0023] The age-stratified unit establishes independent model parameters for different age groups;
[0024] The region adaptation unit loads the corresponding region-specific model baseline values based on the user's geographical location.
[0025] Furthermore, the visualization report generated by the report generation and output module includes highlighted marks of abnormal areas on the standard auricular acupuncture chart, historical change curves of key indicators, consistency prompts for multi-source data, and future risk prediction curves, and pushes early warning notifications through the user interaction platform.
[0026] This invention discloses a health screening method based on dynamic temporal sequence of multimodal ear images, which uses any of the aforementioned health screening systems based on dynamic temporal sequence of multimodal ear images, and includes the following steps:
[0027] Acquire visible light and multispectral image sequences of the user's ear;
[0028] Image sequences are registered and auricular acupoint regions are segmented to extract static features and dynamic temporal features reflecting microvascular pulsation.
[0029] Obtain data from user-authorized external health monitoring devices and health questionnaire information, and perform consistency assessment and weighted fusion of multi-source data;
[0030] The fused feature vector is input into the AI dynamic analysis engine module, which, combined with the user's personal health baseline, outputs the current disease risk level, abnormal area prompts, and future risk predictions.
[0031] Generate visual reports and push them to users.
[0032] Furthermore, it also includes steps to automatically adjust the weights of each feature in the model based on hospital diagnosis results reported by users, and to update the individualized health baseline.
[0033] The beneficial effects of this invention are:
[0034] This invention enhances the accuracy and early identification capabilities of ear health screening by combining multimodal image acquisition with dynamic temporal feature extraction. The system is equipped with a high-resolution color camera and a high-speed multispectral imager, coupled with an adaptive light source control unit to dynamically optimize illumination conditions, effectively improving the signal-to-noise ratio of multispectral images and the quality of microvascular imaging, overcoming the limitations of traditional ear imaging that only analyzes static features. By extracting temporal features such as the frequency and amplitude of microvascular pulsation in the ear acupoint area, physiological functions such as peripheral blood circulation and vascular elasticity can be indirectly assessed, enabling early identification of subclinical conditions such as decreased vascular elasticity. Simultaneously, the calibration mechanism module eliminates interference from individual differences and environmental factors through multi-dimensional calibration based on skin color, age, and region, improving the accuracy of feature recognition for different populations and making the screening results more closely reflect the user's actual physiological state.
[0035] This invention, based on multi-source data fusion and the construction of a personalized health baseline, solves the problems of insufficient single-data dimensions and poor adaptability of population statistical baselines in traditional health assessments. The system integrates ear imaging features, wearable device physiological parameters, and health questionnaire information, forming a closed-loop data processing mechanism through weighted fusion, consistency scoring, and cross-validation. This effectively avoids assessment bias caused by distortion from single data points and improves the reliability of the results. The personalized health baseline, established based on the user's initial data collection and combined with a sliding update mechanism, can accurately distinguish between normal physiological fluctuations and abnormal pathological changes, reducing false positive and false negative results. Simultaneously, a feedback learning mechanism based on hospital diagnostic results continuously optimizes model weights, allowing the system's assessment accuracy to gradually improve with repeated use, achieving personalized health assessment iterations.
[0036] This invention, through an AI dynamic analysis engine and visualized report output, achieves a shift from passive screening to proactive prevention, enhancing the practicality and user experience of health screening. The AI engine integrates multi-model collaborative analysis, not only outputting the current disease risk level but also predicting the evolution of health risks over the next three to six months based on temporal neural networks, identifying intervention points, and providing users with forward-looking health management recommendations. The visualized report offers flexible configuration of displayed content, including historical indicator change curves, and can selectively display risk prediction curves to meet the cognitive preferences and usage scenarios of different users. Simultaneously, the system's short single-test time allows users to complete the test independently at home, lowering the operational threshold for health screening and balancing convenience and professionalism, providing an efficient and feasible solution for daily health management. Attached Figure Description
[0037] Figure 1 This is a schematic diagram of the overall module architecture of a health screening system based on dynamic temporal sequence of multimodal ear images according to an embodiment of this application.
[0038] Figure 2 This is a flowchart illustrating the data processing and feature extraction process in the embodiments of this application. Detailed Implementation
[0039] To enable those skilled in the art to better understand the present invention, the technical solutions in the specific embodiments of the present invention will be clearly and completely described below.
[0040] This invention discloses a health screening system based on dynamic temporal sequence of multimodal ear images, comprising:
[0041] The multimodal image acquisition module is used to acquire visible light image sequences and multispectral image sequences of the user's ear area over a continuous time period.
[0042] The data processing and feature extraction module, whose input end is connected to the output end of the multimodal image acquisition module, is used to register and segment the received image sequence, extract the static features of each auricular acupoint region and the temporal features reflecting the dynamic changes of subcutaneous microvessels;
[0043] The multi-source data fusion module has a first input end connected to the output end of the data processing and feature extraction module, a second input end for accessing data from user-authorized external health monitoring devices, and a third input end for accessing health questionnaire information, which is used to perform consistency assessment and weighted fusion of multi-source data.
[0044] The personal health record database is bidirectionally connected to the data processing and feature extraction module and the multi-source data fusion module. It is used to store multimodal image data collected in each iteration, extracted feature vectors, external device synchronization data and user feedback information, and to establish an individualized health baseline.
[0045] The AI dynamic analysis engine module has its input end connected to the output end of the multi-source data fusion module and bidirectionally connected to the personal health record database. It is used to read individualized health baselines and write analysis results, including at least a disease risk assessment model based on static features, a change analysis model based on time-series features, and a comprehensive risk assessment model that integrates multi-source data.
[0046] The report generation and output module connects to the output of the AI dynamic analysis engine module and is linked to the personal health record database to read historical data. This data is used to generate a visual screening report and outputs risk level, abnormal area alerts, and health recommendations.
[0047] The multimodal image acquisition module adopts an integrated design, including a high-resolution color camera and a high-speed multispectral imager with a frame rate of no less than 10 frames per second. The frame rate of the high-speed multispectral imager is set to 10 frames / second. Based on the physiological laws of human microvascular pulsation, the pulsation frequency is 0.8-1.5Hz in the resting state and up to 2.5Hz after exercise. 10 frames / second can completely capture the waveform characteristics of a single pulsation and avoid data redundancy caused by excessively high frame rates. The multispectral bands selected are 400nm-900nm, with a focus on the 540nm, 560nm, and 620nm bands. These bands are the characteristic absorption bands of hemoglobin to light. 540nm and 560nm are the absorption peaks of oxyhemoglobin, and 620nm is the absorption peak of deoxyhemoglobin. This can maximize the contrast of the imaging of subcutaneous microvessels in the ear and can acquire visible light image sequences and multispectral image sequences of multiple bands of the user's ear area in a continuous time period of 10 to 30 seconds. This module is equipped with an adaptive light source control unit, consisting of a ring-shaped adjustable LED light source array and a real-time image quality feedback circuit. During the acquisition process, it analyzes the signal-to-noise ratio and microvascular imaging clarity of the current multispectral image in real time, automatically adjusting the wavelength and intensity of the light source until an optimized image sequence that meets the analysis requirements is obtained. The input end of the data processing and feature extraction module is connected to the output end of the multimodal image acquisition module. After receiving the image sequence, it first uses a deep learning-based image registration algorithm to perform inter-frame registration to eliminate motion displacement deviations. To address the issue of inconsistent frame rates between the visible light image sequence (30 frames / second) and the multispectral image sequence (10 frames / second), the system first performs timestamp interpolation on the multispectral image sequence before performing temporal feature extraction. It then uses a cubic spline interpolation method to resample the multispectral image sequence to 30 frames / second, achieving time alignment with the visible light image sequence. The aligned image sequence uses the visible light image as the spatial reference frame, mapping the multispectral image to the same spatial coordinate system through feature point matching, ensuring that the subsequent static and dynamic feature extraction of each auricular acupoint region is based on the same spatiotemporal reference. After receiving the image sequence, the data processing and feature extraction module first performs image quality assessment, calculating four indicators: signal-to-noise ratio, sharpness, contrast, and microvascular imaging integrity. When any indicator falls below a preset threshold, a resampling command is sent back to the multimodal image acquisition module, and only qualified image sequences are subsequently registered and feature extracted. Then, based on a standard auricular hologram, the ear is segmented into multiple auricular acupoint regions corresponding to internal organs using an image segmentation network. For each region, the module extracts static features, including color, texture, morphology, and multispectral reflectance, while simultaneously extracting dynamic temporal features. This involves analyzing the pixel intensity changes in the near-infrared band of the multispectral image sequence to obtain the microvascular pulsation waveforms of each region, and performing time-frequency transformation on the waveforms to obtain the pulsation frequency, amplitude, waveform features, and phase difference between regions.Continuous acquisition and adaptive light source ensure stable image quality. The introduction of dynamic features enables the system to capture subcutaneous microvascular pulsation, a physiological information reflecting cardiovascular function and autonomic nervous regulation, providing richer functional dimensions for subsequent analysis.
[0048] The multi-source data fusion module has three input terminals. The first input terminal connects to the output terminal of the data processing and feature extraction module, receiving ear image features. The second input terminal accesses data from user-authorized external health monitoring devices via an application programming interface (API), including physiological parameters such as heart rate, blood oxygen saturation, sleep, and activity levels. The third input terminal accesses health questionnaire information filled out by the user through a mobile application, including text data such as symptom reports, lifestyle habits, and past medical history. Internally, the module includes a weighted fusion unit, a data consistency scoring unit, and a cross-validation unit. The weighted fusion unit fuses ear image features, wearable device features, and questionnaire text features processed by natural language processing. The clinically applicable base weights are set based on clinical experience: 60% for ear images, 30% for wearable devices, and 10% for the questionnaire. These weights can be dynamically adjusted based on the consistency score. Wearable devices can include smartwatches, blood pressure monitors, and blood glucose meters, connected via Bluetooth or Wi-Fi. The data consistency scoring unit calculates the correlation coefficient between the multi-source data, generating a consistency score of 0-1. When the score is below a preset threshold, the fusion weight of the corresponding data is automatically reduced, and a low-confidence warning is indicated in the report. The cross-validation unit incorporates a rule engine that triggers manual review when significant discrepancies arise between ear features and wearable device data. The personal health record database is bidirectionally connected to the data processing and feature extraction module and the multi-source data fusion module. Based on the user's initial data collection, an individualized health baseline is established, storing multimodal image data from each collection, feature vectors, external device synchronization data, and user feedback. It also records the difference between the feature vector collected each time and the baseline, and updates the model weights based on the hospital's diagnostic results provided by the user for personalized iteration. The advantages of this approach are: multi-source data fusion compensates for the limitations of a single ear image dimension; consistency scoring and dynamic weights enhance the reliability of the results; and the establishment of an individualized health baseline allows the system to perform longitudinal comparisons using the user's own historical data as a reference, effectively eliminating individual differences.
[0049] Personalized health baselines are established based on the user's initial data collection. If the initial data collection score is below 60 or there are obvious operational abnormalities, the system will prompt the user to recollect the data until valid data is obtained before a baseline can be established. Baseline updates use a sliding window mechanism, retaining the 12 most recent valid data collection records. The update weight is 20% for new data and 80% for historical data, and the normal fluctuation range is automatically adjusted according to seasonal changes, physiological cycles, and other normal fluctuations.
[0050] The hospital diagnosis results reported by the user trigger two update paths simultaneously: first, updating the feature weights of each model in the AI dynamic analysis engine, which affects the calculation logic of subsequent risk assessment; and second, updating the individualized health baseline in the personal health record database, which affects the calculation benchmark of the difference vector. These two processes are executed independently, with the model weight update using the diagnosis result as a supervisory signal, and the baseline update using the user's own historical data as a statistical basis.
[0051] The AI dynamic analysis engine module's input is connected to the output of the multi-source data fusion module and bidirectionally connected to the personal health record database to read individualized health baselines and write analysis results. This engine comprises several collaborative sub-models: a static feature-based disease risk assessment model using a deep neural network architecture, inputting current static features and outputting preliminary risk probabilities for each disease; a time-series feature-based change analysis model using a long short-term memory network to process the user's historical feature vector sequences, identifying abnormal change patterns and outputting the current deviation from the individualized baseline; a comprehensive risk assessment model jointly analyzes static risk outputs, time-series change features, and multi-source fusion features to generate risk levels, confidence levels, and major abnormal region identifiers; and a disease progression prediction model using a long short-term memory network to predict the evolution trend of health risks over the next three to six months, outputting risk probability change curves and suggested intervention points. The report generation and output module's input is connected to the AI dynamic analysis engine module's output and to a personal health record database to retrieve historical data. The generated visualized screening report includes highlighted areas of abnormality on a standard auricular acupuncture chart, historical change curves of key indicators, consistency indicators from multiple data sources, and future risk prediction curves. Reports and warning notifications are pushed to users via a mobile application. The beneficial effects of this approach are: multi-model collaboration comprehensively analyzes health status from different dimensions, improving the accuracy and reliability of risk assessment; disease progression prediction realizes a shift from passive screening to proactive prevention, providing users with forward-looking health management recommendations; and visualized reports make complex analysis results easier to understand, enhancing the system's usability and user experience.
[0052] As one implementation, the multimodal image acquisition module includes a high-resolution color camera and a high-speed multispectral imager, and is equipped with an adaptive light source control unit for dynamically adjusting the wavelength and intensity of the illumination source according to the real-time image quality, so as to optimize the signal-to-noise ratio and microvascular imaging quality of the multispectral image sequence.
[0053] The multimodal image acquisition module consists of a high-resolution color camera, a high-speed multispectral imager, a ring-shaped multi-band LED light source array, and an adaptive light source control unit. The high-resolution color camera uses a sensor with over 24 megapixels to acquire ear surface morphology and color information in the visible light band. The high-speed multispectral imager is equipped with 5 to 10 narrowband filters, covering a wavelength range of 400nm to 900nm, with a focus on the 540nm, 560nm, and 620nm bands sensitive to hemoglobin absorption. The ring-shaped multi-band LED light source array is divided into multiple independently controllable areas, each containing LEDs of different wavelengths. The light source intensity can be continuously adjusted from 0% to 100% via pulse width modulation. The adaptive light source control unit incorporates image quality assessment algorithms and light source driving circuitry, working synchronously with the camera and imager.
[0054] The adaptive light source control unit's workflow consists of two stages: image quality assessment and light source parameter adjustment. During image acquisition, the control unit reads camera preview frames in real time and calculates four indicators: average brightness, contrast, signal-to-noise ratio, and microvascular region sharpness. Average brightness is obtained by calculating the mean of the image's gray-level histogram; contrast is obtained by calculating the gray-level standard deviation; signal-to-noise ratio is determined by calculating the ratio of the signal area to the background area; and microvascular region sharpness is assessed using the Laplacian operator for edge sharpness evaluation. When any indicator falls below a preset threshold, the control unit initiates a light source adjustment program. Based on the current image quality deviation direction, it calculates the required combination of light source wavelengths and intensity ratios, and drives the corresponding LED light-emitting units through pulse width modulation signals. The adjustment process employs an iterative optimization method, re-acquiring the image and evaluating its quality after each adjustment until all four indicators reach the preset standards or the maximum number of adjustments is reached.
[0055] The four image quality metrics are calculated using the following formulas: Average brightness = arithmetic mean of the grayscale values of all pixels in the image; Contrast ratio = standard deviation of the grayscale values in the image; Signal-to-noise ratio = (mean grayscale value of the signal region - mean grayscale value of the background region) / standard deviation of the grayscale values in the background region; Microvascular clarity = average gradient value of the edge pixels of the image after convolution with the Laplacian operator. The light source iteration optimization termination condition is as follows: When all four metrics reach the preset thresholds (average brightness 120-180, contrast ratio ≥ 50, signal-to-noise ratio ≥ 20, microvascular clarity ≥ 1.5), or when the number of iterations reaches 5, the light source adjustment stops and formal acquisition begins. If the criteria are not met after 5 iterations, the user is prompted to check the shooting posture and re-acquire the image.
[0056] This adaptive light source control unit effectively improves the signal-to-noise ratio and microvascular imaging quality of multispectral image sequences. The adaptive light source control automatically optimizes lighting conditions based on the user's ear skin tone, ambient light variations, and shooting distance, avoiding increased image noise due to insufficient lighting or loss of detail due to excessive lighting. By dynamically adjusting the intensity of a specific wavelength of light, the contrast of hemoglobin's light absorption is enhanced, making subcutaneous microvascular structures more clearly discernible in multispectral images.
[0057] As one implementation, the data processing and feature extraction module also includes a dynamic feature extraction unit, which extracts the microvascular pulsation waveforms of each auricular acupoint region from the multispectral image sequence and performs time-frequency transformation on the pulsation waveforms to obtain the pulsation frequency, pulsation amplitude, waveform feature parameters, and pulsation phase difference between different auricular acupoint regions as temporal features.
[0058] The dynamic feature extraction unit extracts microvascular pulsation waveforms from continuously acquired multispectral image sequences. First, inter-frame registration is performed on the image sequence, and a feature point matching algorithm is used to ensure the spatial consistency of the same auricular acupoint region in each frame. Then, multiple sampling points are selected within each auricular acupoint region, and the reflection intensity values of each sampling point in the 540nm and 560nm bands are read. The ratio of the reflection intensity of the two bands is calculated to eliminate interference from skin surface reflection. The continuously acquired 30 to 60 frames are arranged chronologically to obtain a curve showing the change in reflection intensity of each sampling point over time, reflecting the pulsation signal of subcutaneous microvessels. The raw signal is preprocessed using a bandpass filter with a filtering frequency range of 0.5 Hz to 3 Hz to remove high-frequency noise and low-frequency drift. When the filtered microvascular pulsation frequency is below 0.5 Hz or above 3 Hz, the system determines that the microvascular pulsation in that area is abnormal and marks this feature as a health risk characteristic, incorporating it into the subsequent multi-source data fusion risk assessment system.
[0059] The preprocessed pulsation signal was subjected to time-frequency transformation to obtain temporal characteristic parameters. The time-frequency transformation employed continuous wavelet transform, using Morlet wavelets as the basis function and a scale parameter ranging from 1 to 32. The pulsation frequency was extracted from the time-frequency transformation results, defined as the frequency value corresponding to the peak value of the energy spectrum. The pulsation amplitude was extracted, defined as the average difference between the peaks and troughs of the time-domain signal. Waveform characteristic parameters were extracted, including the ratio of systolic to diastolic area, the ratio of ascending to descending time, and the proportion of dominant frequency energy. The pulsation phase difference between different auricular acupoint regions was extracted by calculating the time difference of the pulsation peaks in each region and dividing it by the pulsation period to obtain the standardized phase value. All characteristic parameters were stored as feature vectors for subsequent analysis.
[0060] Traditional ear imaging analysis focuses only on static color and texture features, failing to reflect the dynamic functional state of the cardiovascular system. By extracting microvascular pulsation waveforms and their time-frequency characteristics, peripheral blood circulation, vascular elasticity, and autonomic nervous system regulation can be indirectly assessed. Abnormal pulsation frequency may indicate heart rate variability problems, decreased pulsation amplitude may indicate insufficient peripheral perfusion, and phase differences often indicate vascular conduction dysfunction.
[0061] As one implementation method, the multi-source data fusion module includes:
[0062] The weighted fusion unit is used to dynamically fuse ear image features, wearable device data and questionnaire information according to preset basic weights. The basic weights can be adjusted according to data quality and historical consistency.
[0063] The data consistency scoring unit is used to calculate the correlation coefficient between multi-source data and generate a consistency score. When the score is lower than the preset threshold, the fusion weight of the corresponding data is automatically reduced and a low confidence warning is marked in the report.
[0064] The cross-validation unit is used to detect whether there are significant contradictions between ear image features and wearable device data. When a contradiction occurs, it triggers a manual review prompt.
[0065] The basic weight allocation in this scheme is determined based on the "Clinical Guidelines for Auricular Holographic Diagnosis and Treatment Technology" (2023 edition) and the results of multi-center clinical trials. The trials included 5,000 healthy and subclinical individuals of different age groups. The results showed that ear imaging features contributed the most to health risk assessment, wearable device physiological parameters were important auxiliary factors, and health questionnaires were a clinical supplement. The consistency score threshold of 0.4 was set based on the clinical statistical judgment standard of Pearson correlation coefficient. When the correlation coefficient < 0.4, there is no significant linear correlation between multi-source data, which is judged as low consistency.
[0066] When manual review is triggered, the system pauses the automatic generation of reports for that batch and pushes the relevant data to a qualified physician or system operator. After review, if the data is valid, the system resumes report generation and marks it as reviewed; if the data is invalid, the system removes the batch of data and prompts the user to re-collect the data. The review results are also stored in the user's personal health record database for optimizing subsequent consistency assessment rules.
[0067] The multi-source data fusion module consists of a weighted fusion unit, a data consistency scoring unit, and a cross-validation unit. It integrates three data sources: ear imaging features, wearable device data, and health questionnaire information. Ear imaging features include temporal feature vectors such as microvascular pulsation frequency, amplitude, and phase difference. Wearable device data includes physiological parameters such as heart rate, blood pressure, and blood oxygen saturation. Health questionnaire information includes structured data such as symptom scores, past medical history, and lifestyle habits. The weighted fusion unit assigns initial weights to each data type. For example, in one approach, the weight of ear imaging features suitable for home and portable use is set to 0.5, the weight of wearable device data is set to 0.35, and the weight of questionnaire information is set to 0.15, with a total weight of 1.0. Each unit receives input data through a standard data interface and outputs the fused comprehensive feature vector for use by the AI analysis engine.
[0068] The weighted fusion unit dynamically adjusts weights based on data quality and historical consistency. Data quality scoring is calculated based on three indicators: missing value rate, outlier ratio, and collection time validity. Each indicator has a maximum score of 100 points, and the weighted average yields the quality score. Historical consistency is calculated by determining the Euclidean distance between the current data and the user's three previous data collections; a smaller distance indicates higher consistency. The data consistency scoring unit calculates the Pearson correlation coefficient between the three data sources. When the correlation coefficient is below 0.4, it is considered low consistency, automatically reducing the weight of the corresponding data source by 50% and marking a low-confidence warning in the final report. The cross-validation unit sets up 10 rule bases. Each rule defines a reasonable correspondence between ear image features and wearable device data. For example, if the difference between the ear microvascular pulsation frequency and heart rate exceeds 20 beats per minute, it is considered a significant contradiction, triggering a manual review prompt and pausing automatic report generation.
[0069] By introducing consistency scoring and cross-validation mechanisms, data distortion caused by sensor malfunctions, user errors, or abnormal physiological states can be effectively detected. When the credibility of a certain type of data decreases, the system automatically reduces its weight instead of discarding it completely, ensuring the continuity of the evaluation while avoiding interference from erroneous data.
[0070] In this embodiment, the human reviewer is a qualified physician or system operator. The review primarily includes three aspects: data validity verification, equipment malfunction investigation, and confirmation of user physiological abnormalities. After the review process, the system will execute different feedback strategies based on the review results. If the review result determines the data is invalid, such as due to equipment acquisition failure or data transmission error, the system will automatically remove the invalid data and re-analyze and calculate based on the remaining valid data. If the review result determines that the user has potential physiological abnormalities, the system will highlight the abnormal characteristics and simultaneously push professional medical advice to the user. By clearly defining the reviewer, content, and feedback mechanism, a complete closed loop is constructed from problem triggering and human intervention to result feedback, effectively improving the reliability and rigor of the entire health screening system.
[0071] As one implementation method, the personal health record database establishes an individualized health baseline based on the data collected from the user's first collection, and records the difference vector between the feature vector and the baseline for each subsequent collection.
[0072] The database also updates the model weights based on hospital diagnosis results reported by users, enabling personalized iteration.
[0073] Before establishing an individualized health baseline, the system verifies the validity of the initially collected data: a data consistency score is calculated. If the score is below a preset threshold (e.g., 60 points) or there are obvious operational anomalies such as blurred images or failed region segmentation, the user is prompted to recollect data until valid data is obtained before the baseline can be established. This mechanism ensures the reliability of the baseline and avoids subsequent assessment biases due to data collection quality issues.
[0074] The personal health record database uses a hybrid architecture of relational database and vector database to store user health data. When a user completes a health assessment after passing the initial assessment, the system extracts 28 indicators across 3 categories: ear image feature vectors, wearable device physiological parameters, and questionnaire scores. These 28 indicators are specifically defined as follows: Ear image features (13 items): static features (color L / a / b values, texture grayscale co-occurrence matrix, morphological area / perimeter, etc., 5 items), dynamic temporal features (pulsation frequency, amplitude, systolic / diastolic area ratio, ascending / descending limb time ratio, regional phase difference, etc., 8 items); Wearable device physiological parameters (10 items): heart rate, systolic / diastolic blood pressure, blood oxygen saturation, body temperature, respiratory rate, deep sleep duration, light sleep duration, daily activity level, and blood glucose (fasting / postprandial); Health questionnaire scores (5 items): symptom self-assessment score, lifestyle habit score, past medical history score, family medical history score, and psychological state score. Each questionnaire score uses a 0-10 point quantification standard, is converted into feature vectors through natural language processing, and the mean and standard deviation of each indicator are calculated as an individualized health baseline. Baseline data includes five fields: feature name, baseline value, normal fluctuation range, collection timestamp, and confidence score. During each subsequent collection, the system automatically calculates the difference vector between the current feature vector and the baseline. The difference value is calculated using a standardized formula: current value minus baseline value, divided by the baseline standard deviation. The difference vector is stored in a separate data table, containing user ID, collection batch, difference value, trend direction, and clinical significance annotation.
[0075] Relational databases are primarily used to store structured data such as basic user information, timestamps of each data collection, data collection device information, clinical diagnosis results, and health screening report texts. Vector databases, on the other hand, are specifically used to store unstructured high-dimensional data such as feature vectors extracted from ear images, difference vectors after multi-source data fusion, and microvascular pulsation time-series feature vectors. This database-splitting approach ensures both efficient retrieval of basic user information and meets the need for rapid comparison of high-dimensional feature vectors.
[0076] Personalized iterative model weighting is achieved through a feedback learning mechanism. When a user enters a hospital diagnosis into the system, the system stores the diagnosed disease type, diagnosis time, and diagnostic basis in a structured manner. The database calls the model update algorithm to compare the consistency between the system's prediction and the hospital's diagnosis. For correctly predicted samples, the original model weights remain unchanged. For samples with prediction errors, a gradient descent algorithm is used to adjust the feature weights; the adjustment magnitude is determined by the magnitude of the prediction error—the larger the error, the larger the adjustment. Weight updates employ a sliding window mechanism, using only the user's most recent 10 data collections for calculation, avoiding excessive influence of earlier data on the current model. After each weight update, the system recalculates the user's health risk assessment result and pushes an updated report.
[0077] Traditional health assessment systems rely on population statistical baselines, which fail to recognize the inherent differences in individual physiological characteristics, leading to false positive or false negative results for some users. By establishing individualized baselines, the system can distinguish between normal physiological fluctuations and abnormal pathological changes, reducing the false positive rate. Personalized iteration of model weights enables the system to learn each user's unique health patterns, resulting in continuously improving assessment accuracy with increased usage.
[0078] As one implementation method, the AI dynamic analysis engine module also includes a disease progression prediction model. This model uses a time-series neural network to predict the evolution trend of health risks within a specific future time period using the user's historical multimodal data sequence, and can output risk probability change curves and suggested intervention nodes.
[0079] The disease progression prediction model is built on a temporal neural network architecture. The input data consists of a sequence of historical multimodal data from the user, including ear image feature vectors, wearable device physiological parameters, and health questionnaire scores. The model uses a Long Short-Term Memory (LSTM) network as its core algorithm. The network structure includes an input layer, three hidden layers, and an output layer. The input layer receives feature data collected each time within the past 12 months, arranged chronologically to form a temporal input matrix. Each hidden layer contains 128 neurons, and the ReLU activation function is used. The output layer outputs the health risk probability values for the next 1 month, 3 months, and 6 months.
[0080] The model training employs supervised learning, using historical user data as the training set. In this embodiment, the training set contains no fewer than 5000 samples, a number set according to the statistical sample size requirements for machine learning model training. For multimodal and multi-population stratified feature dimensions, each stratum has no fewer than 100 samples to meet the model's generalization training requirements. The sample inclusion criteria are: age 18-80 years, no ear skin diseases, able to cooperate in ear image acquisition, and at least one hospital health checkup / diagnosis result as a label. Samples with distorted feature extraction due to local ear lesions are excluded. The samples cover user groups of different ages, skin colors, and geographical regions, and each sample contains at least three consecutive multimodal acquisition data points and corresponding health status labels to ensure the model's generalization ability. The training data includes input feature sequences and corresponding actual health result labels, derived from the user's subsequent hospital diagnosis records or follow-up results. The cross-entropy function is used as the loss function, the Adam algorithm is used as the optimizer, and the learning rate is set to 0.001. An early stopping strategy is employed during training: training stops when the validation set loss no longer decreases for five consecutive epochs. The model outputs a risk probability change curve, with time on the horizontal axis and disease risk probability on the vertical axis. It also outputs suggested intervention points, determining the optimal intervention time by calculating the slope of the risk probability increase; intervention points are marked when the slope exceeds a preset threshold.
[0081] By mining historical data through temporal neural networks to identify patterns in health deterioration trends, it is possible to identify them in advance. Risk probability change curves provide users with a clear understanding of their health trajectory, and suggested intervention points offer users specific timeframes for action.
[0082] For new users or users with fewer than 12 historical data collection records, the system employs a cold start strategy: using statistical model parameters from populations of the same age and region as the initial prediction model, while incorporating the multimodal characteristics of the user's current single data collection, to output short-term risk prediction results. As the number of data collections increases, the model gradually migrates to personalized time-series prediction, ensuring that the system can provide effective health risk assessments throughout the entire lifespan.
[0083] As one implementation, it also includes a calibration mechanism module, which includes at least a skin color calibration unit, an age stratification unit, and a region adaptation unit;
[0084] The skin color calibration unit corrects skin color deviations in ear color analysis by acquiring multispectral images of the user's palm or forearm skin.
[0085] The age-stratified unit establishes independent model parameters for different age groups;
[0086] The region adaptation unit loads the corresponding region-specific model baseline values based on the user's geographical location.
[0087] The calibration mechanism module consists of a skin color calibration unit, an age stratification unit, and a geographic adaptation unit, designed to eliminate the influence of individual differences and environmental factors on health assessment results. The skin color calibration unit acquires multispectral images of the palm or forearm skin upon first use, with acquisition bands including 450nm, 540nm, 560nm, 620nm, and 660nm. A skin color correction coefficient matrix is established by calculating the optical reflectance characteristics of the reference area skin. When applied to ear image analysis, this correction coefficient eliminates the interference of melanin content differences on blood vessel color recognition. The correction algorithm uses a linear regression method to establish a mapping relationship between the reflectance of the reference area and the reflectance of the ear area, resulting in a corrected ear color value that more closely approximates the true state of blood vessels.
[0088] The age stratification unit divides users into six age groups: 18-30, 31-40, 41-50, 51-60, 61-70, and 71 and above. Each age group has its own independent model parameter library, containing the normal physiological range, disease risk thresholds, and feature weight coefficients for that age group. The model parameters are trained based on large-scale population cohort study data, with a sample size of no less than 5000 cases for each age group. Users enter their birthdate during registration, and the system automatically matches the model parameters for the corresponding age group. The geographic adaptation unit loads relevant baseline values from the geographic model library based on the user's registered geographic location or mobile phone GPS location information. The geographic model library covers seven major geographical regions of China, and each region includes correction parameters for influencing factors such as climate, diet, and lifestyle habits.
[0089] After skin color calibration, the accuracy of vascular feature recognition for users with dark skin improved from 65% to 88%, resolving the technical challenge caused by melanin interference. Age stratification makes the assessment criteria for each age group more in line with physiological characteristics, improving the accuracy of cardiovascular risk prediction for users over 40 years old by approximately 25%. Geographic adaptation takes into account the impact of environmental factors on health indicators, reducing the deviation of blood pressure assessment in winter by approximately 30% for users in northern regions.
[0090] The calibration mechanism module has its input terminals connected to the output terminals of the multimodal image acquisition module and the data processing and feature extraction module, respectively, and its output terminal connected to the input terminal of the multi-source data fusion module. It is used to perform multi-dimensional calibration of the feature vector based on skin color, age, and region after feature extraction and before data fusion.
[0091] As one implementation method, the visualization report generated by the report generation and output module includes highlighted marks of abnormal areas on the standard auricular acupuncture chart, historical change curves of key indicators, consistency prompts of multi-source data, and future risk prediction curves, and pushes early warning notifications through the user interaction platform.
[0092] The report generation and output module uses a combination of front-end rendering and back-end data to generate visualized health reports. The standard auricular acupoint chart is constructed based on the standard anatomical atlas of the human ear, containing the coordinates of 78 standard auricular acupoints. Based on the coordinates of abnormal areas output by the AI dynamic analysis engine, the system highlights abnormal acupoints in red, suspicious areas in yellow, and normal areas in green on the auricular acupoint chart. The highlighted areas use a semi-transparent overlay technology with a transparency set to 60% to ensure the acupoint names and numbers are clearly visible. Historical change curves of key indicators are displayed as line graphs, with the horizontal axis representing the collection date and the vertical axis representing the indicator value, showing up to 24 data collections from the past 12 months. The curves support zooming and hovering to view specific values.
[0093] The multi-source data consistency alert module features a status bar at the top of the report, displaying consistency scores for the three data sources using icons and text. A green "pass" indicator is displayed when the consistency score is above 80, a yellow "caution" indicator is displayed between 60 and 80, and a red "warning" indicator with detailed explanations is displayed when the score is below 60. The future risk prediction curve uses an area chart to show the trend of health risk probability changes over the next 1, 3, and 6 months, with risk threshold lines marked by dashed lines. The user interaction platform sends alert notifications through three channels: mobile application push notifications, SMS, and email. Alert levels are divided into three levels: Level 1 alerts are pushed immediately, Level 2 alerts are pushed daily in summary, and Level 3 alerts are pushed weekly in summary.
[0094] Visualized auricular acupoint charts allow users to intuitively locate abnormal areas, improving comprehension by approximately 50%. Historical change curves help users track health trends and enhance long-term management awareness. Multi-source data consistency indicators increase report transparency, allowing users to understand the reliability of assessment results. Risk prediction curves provide forward-looking health information, prompting users to take early intervention measures.
[0095] This invention discloses a health screening method based on dynamic temporal sequence of multimodal ear images, which uses any of the aforementioned health screening systems based on dynamic temporal sequence of multimodal ear images, and includes the following steps:
[0096] Acquire visible light and multispectral image sequences of the user's ear;
[0097] Image sequences are registered and auricular acupoint regions are segmented to extract static features and dynamic temporal features reflecting microvascular pulsation.
[0098] Obtain data from user-authorized external health monitoring devices and health questionnaire information, and perform consistency assessment and weighted fusion of multi-source data;
[0099] The fused feature vector is input into the AI dynamic analysis engine module, which, combined with the user's personal health baseline, outputs the current disease risk level, abnormal area prompts, and future risk predictions.
[0100] Generate visual reports and push them to users.
[0101] The health screening method first acquires images of the user's ear using a dedicated ear imaging device. Visible light image sequences are acquired at a frequency of 30 frames per second, for a total of 90 frames over 3 seconds. Multispectral image sequences are acquired at 450nm, 540nm, 560nm, 620nm, and 660nm, with 10 frames acquired for each wavelength. Image registration uses a feature point matching algorithm to extract anatomical landmarks such as the ear contour, tragus, and antihelix as references, with registration accuracy controlled within 0.5 mm. Ear acupoint region segmentation is based on a deep learning semantic segmentation network, trained on 5000 labeled ear images, accurately identifying 78 standard ear acupoint regions. Static features include 15 indicators of ear acupoint color, texture, and morphology. Dynamic temporal features are extracted using optical volumetric plethysmography to obtain microvascular pulsation signals, including 8 indicators of pulsation amplitude, frequency, and waveform characteristics.
[0102] In the multi-source data fusion phase, the system acquires data from user-authorized external health monitoring devices via an application programming interface (API), including 12 physiological parameters such as heart rate, blood pressure, blood oxygen saturation, and blood glucose. This external health monitoring device data includes, but is not limited to, heart rate, blood pressure, blood oxygen saturation, body temperature, respiratory rate, sleep quality, and activity level. Data is synchronized in real-time via Bluetooth, Wi-Fi, or API interfaces and aligned with the ear image acquisition timestamp to ensure data synchronization. The health questionnaire covers 35 questions covering symptom self-assessment, lifestyle habits, and family medical history. The questionnaire is collected via an interactive form on a mobile device, using a combination of structured multiple-choice questions and open-ended text. The text content undergoes symptom keyword extraction and sentiment analysis using natural language processing algorithms, transforming it into quantifiable feature vectors before fusion. Multi-source data consistency assessment uses correlation analysis to calculate the correlation coefficient between ear image features and external device data. A data anomaly alert is triggered when the correlation coefficient is below 0.4. The fused feature vector is input into the AI dynamic analysis engine, which calculates the difference vector by combining it with the user's personal health baseline. The output shows the current disease risk level, which is divided into three levels: low, medium, and high. Abnormal areas are marked with specific auricular acupoint locations. Future risk predictions are displayed with risk probability curves for 1 month, 3 months, and 6 months.
[0103] This method allows for preliminary screening via ear imaging, with a single test taking no more than 5 minutes, which users can perform at home. Multimodal data fusion improves assessment accuracy. Clinical validation based on 1000 high-risk individuals for cardiovascular disease shows a sensitivity of 85% and a specificity of 82% for cardiovascular disease screening. Dynamic temporal features capture changes in microvascular function, enabling early identification of subclinical conditions such as decreased vascular elasticity, detecting risk approximately 3 months earlier than traditional methods. Personalized baselines eliminate interference from population differences, making assessment results more accurate and reliable for users of different skin colors, ages, and regions.
[0104] In this embodiment, the warning levels are divided into Level 1, Level 2, and Level 3. Level 1 is the highest level, triggered when the disease risk probability is greater than or equal to 80% and the multi-source data consistency score is greater than or equal to 80%. When these conditions are met, the system immediately determines that the user has an extremely high health risk and sends a warning notification via push notification, reminding the user to seek medical attention as soon as possible. Level 2 is an intermediate level warning, triggered when the disease risk probability is between 60% and 80%, or the multi-source data consistency score is between 60% and 80%. At this point, the system determines that the user has a moderate health risk and adjusts the push notification frequency to once a day, continuously monitoring data changes. Level 3 is a basic level warning, triggered when the disease risk probability is less than 60% but the multi-source data consistency score is still greater than or equal to 80%. In this case, the user's health risk has not yet reached an emergency level, and the system pushes health advice weekly for long-term health management. Through the above-mentioned quantified trigger conditions, a precise match between the warning level and the risk level is achieved, ensuring the practical value of the health screening report.
[0105] As one implementation method, it also includes the steps of automatically adjusting the weights of each feature in the model based on the hospital diagnosis results reported by the user, and updating the individualized health baseline.
[0106] The model's adaptive optimization function achieves continuous learning through a closed-loop user feedback system. After a user is diagnosed at the hospital, they upload an image of the diagnostic report via a mobile application or manually enter the diagnosis result, including the disease name, diagnosis date, and test result values. The system uses optical character recognition (OCR) technology to automatically extract key information from the report, achieving an accuracy rate of over 95%. After the user uploads the report, the system sends a confirmation request. Once the user confirms the authenticity of the diagnosis, the data is marked as a valid feedback sample and stored in the feedback database. The system sets a feedback time window, only accepting diagnostic results within 3 months after screening as valid feedback to ensure the correlation between feedback data and screening results.
[0107] Feature weight adjustment employs an online learning algorithm, comparing user diagnosis results as true labels with model predictions. When predictions and diagnoses differ, the system calculates the contribution of each feature to the prediction error and adjusts feature weights using gradient descent. Features with higher contributions receive larger weight adjustments, while those with lower contributions receive smaller adjustments. Each weight adjustment is limited to 5% of the current weight to prevent drastic model fluctuations. Personalized health baseline updates use a moving average method, incorporating the latest screening data into the baseline calculation while removing the earliest data, maintaining a stable baseline data volume of 12 collection records. The baseline update weight is set to 20% for new data and 80% for historical data to ensure a smooth baseline transition.
[0108] When a user performs their first health data collection, if there are fewer than 12 historical data collection records, the system will temporarily use the baseline values of a population of the same age group and matching geographical characteristics as a temporary health baseline. This temporary baseline is used to provide a basic health status reference and comparison for users collecting data for the first time. As users continue to collect data, the system will continuously supplement the baseline dataset with new and valid data collection records. When the cumulative number of data collection records reaches 12, the system will automatically switch from the temporary baseline to a personalized baseline based entirely on the user's own data. This approach effectively solves the baseline establishment problem caused by insufficient data for new users, ensuring that the system can provide stable and reliable health assessment services throughout the entire lifecycle.
[0109] Through a feedback loop, the model can continuously learn and improve, with an overall accuracy increase of approximately 2% for every 1000 additional valid feedback cases. The individualized baseline is dynamically updated as the user's health condition changes, avoiding assessment bias caused by baseline aging. After 6 months of use by a long-term user, the system's disease prediction accuracy for that user can increase from the initial 75% to 89%. The feedback mechanism also enhances user engagement and trust, with a user feedback rate reaching 35%, higher than the industry average of 12%.
[0110] It should be understood that those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
Claims
1. A health screening system based on dynamic timing of multi-modal ear images, characterized in that, include: The multimodal image acquisition module is used to acquire visible light image sequences and multispectral image sequences of the user's ear area over a continuous time period. The data processing and feature extraction module, whose input end is connected to the output end of the multimodal image acquisition module, is used to register and segment the received image sequence, extract the static features of each auricular acupoint region and the temporal features reflecting the dynamic changes of subcutaneous microvessels; The multi-source data fusion module has a first input end connected to the output end of the data processing and feature extraction module, a second input end for accessing data from user-authorized external health monitoring devices, and a third input end for accessing health questionnaire information, which is used to perform consistency assessment and weighted fusion of multi-source data. The personal health record database is bidirectionally connected to the data processing and feature extraction module and the multi-source data fusion module. It is used to store multimodal image data collected in each iteration, extracted feature vectors, external device synchronization data and user feedback information, and to establish an individualized health baseline. The AI dynamic analysis engine module has its input end connected to the output end of the multi-source data fusion module and bidirectionally connected to the personal health record database. It is used to read individualized health baselines and write analysis results, including at least a disease risk assessment model based on static features, a change analysis model based on time-series features, and a comprehensive risk assessment model that integrates multi-source data. The report generation and output module connects to the output of the AI dynamic analysis engine module and is linked to the personal health record database to read historical data. This data is used to generate a visual screening report and outputs risk level, abnormal area alerts, and health recommendations.
2. The health screening system based on multimodal ear imaging dynamic temporal sequence according to claim 1, characterized in that: The multimodal image acquisition module includes a high-resolution color camera and a high-speed multispectral imager, and is equipped with an adaptive light source control unit to dynamically adjust the wavelength and intensity of the illumination source according to the real-time image quality, so as to optimize the signal-to-noise ratio and microvascular imaging quality of the multispectral image sequence.
3. The health screening system based on multimodal ear imaging dynamic temporal sequence according to claim 1, characterized in that: The data processing and feature extraction module also includes a dynamic feature extraction unit, which extracts the microvascular pulsation waveforms of each auricular acupoint region from the multispectral image sequence and performs time-frequency transformation on the pulsation waveforms to obtain the pulsation frequency, pulsation amplitude, waveform feature parameters, and pulsation phase difference between different regions as temporal features.
4. A health screening system based on dynamic temporal sequence of multimodal ear images according to claim 1, characterized in that: The multi-source data fusion module includes: The weighted fusion unit is used to dynamically fuse ear image features, wearable device data and questionnaire information according to preset basic weights. The basic weights can be adjusted according to data quality and historical consistency. The data consistency scoring unit is used to calculate the correlation coefficient between multi-source data and generate a consistency score. When the score is lower than the preset threshold, the fusion weight of the corresponding data is automatically reduced and a low confidence warning is marked in the report. The cross-validation unit is used to detect whether there are significant contradictions between ear image features and wearable device data. When a contradiction occurs, it triggers a manual review prompt.
5. A health screening system based on dynamic temporal sequence of multimodal ear images according to claim 1, characterized in that: The personal health record database establishes an individualized health baseline based on the data collected from the user's first collection, and records the difference vector between the feature vector and the baseline for each subsequent collection. The database also updates the model weights based on hospital diagnosis results reported by users, enabling personalized iteration.
6. A health screening system based on dynamic temporal sequence of multimodal ear images according to claim 1, characterized in that: The AI dynamic analysis engine module also includes a disease progression prediction model. This model uses a time-series neural network to predict the evolution trend of health risks within a specific future time period using the user's historical multimodal data sequence, and can output risk probability change curves and suggested intervention nodes.
7. A health screening system based on dynamic temporal sequence of multimodal ear images according to claim 1, characterized in that: It also includes a calibration mechanism module, which includes at least a skin color calibration unit, an age stratification unit, and a region adaptation unit. The skin color calibration unit corrects skin color deviations in ear color analysis by acquiring multispectral images of the user's palm or forearm skin. The age-stratified unit establishes independent model parameters for different age groups; The region adaptation unit loads the corresponding region-specific model baseline values based on the user's geographical location.
8. A health screening system based on dynamic temporal sequence of multimodal ear images according to claim 1, characterized in that: The visualization report generated by the report generation and output module includes highlighted marks of abnormal areas on the standard auricular acupoint chart, historical change curves of key indicators, consistency prompts of multi-source data, and future risk prediction curves, and pushes early warning notifications through the user interaction platform.
9. A method of health screening based on dynamic timing of multi-modal ear images, characterized in that, Using any of the multimodal ear image dynamic temporal health screening systems described in claims 1-8, the following steps are included: Acquire visible light and multispectral image sequences of the user's ear; Image sequences are registered and auricular acupoint regions are segmented to extract static features and dynamic temporal features reflecting microvascular pulsation. Obtain data from user-authorized external health monitoring devices and health questionnaire information, and perform consistency assessment and weighted fusion of multi-source data; The fused feature vector is input into the AI dynamic analysis engine module, which, combined with the user's personal health baseline, outputs the current disease risk level, abnormal area prompts, and future risk predictions. Generate visual reports and push them to users.
10. A health screening method based on dynamic temporal sequence of multimodal ear images according to claim 9, characterized in that: It also includes steps to automatically adjust the weights of each feature in the model based on hospital diagnosis results reported by users, and to update the individualized health baseline.