Individualized health assessment method and device based on patient multi-modal information

By collecting and fusing multimodal medical data, and using deep learning and geometric mean theory to construct a personalized health benchmark, this technology solves the problem of lack of comprehensive assessment and risk factor detection in existing technologies, and achieves quantitative assessment of health status and improved treatment outcomes.

CN121506466APending Publication Date: 2026-02-10CHENGDU UNIV OF TRADITIONAL CHINESE MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511402357.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing individualized health assessment methods for patients mainly focus on the prediction of disease incidence, risk assessment, and efficacy evaluation of a specific disease. They lack comprehensive assessment of all diseases, fail to identify disease risk factors in a timely manner, and lack comparative analysis before and after treatment.

Method used

By collecting and integrating patients' multimodal medical data, including symptom data, tongue image data, facial image data, audio data, and pulse image data, deep learning networks are used for feature extraction and digital processing. Combined with geometric mean theory, an individualized health scale is constructed to achieve quantitative assessment of health status. Furthermore, disease risk factors are discovered through quantitative difference and parameter hole analysis.

Benefits of technology

It enables comprehensive assessment and trend analysis of patients' health status, helping doctors to identify disease risk factors in a timely manner, improve medication quality, and shorten the treatment cycle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121506466A_ABST
    Figure CN121506466A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of health assessment, and particularly discloses an individualized health assessment method and device based on patient multi-modal information. According to the method, the multi-modal medical data, including disease question and answer text information, tongue picture and face picture information and voice information, of the patient is collected and fused, the current individualized health scale of the patient is constructed by using the same inch and geometric mean theory, and quantitative evaluation of the comprehensive health state is achieved. Through construction of health scales of patients in different treatment periods, doctors are helped to visually understand the trend and development of health states of the patients, and disease treatment is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical information technology, and in particular to a method and device for individualized health assessment based on patient multimodal information. Background Technology

[0002] Existing individualized health assessments for patients are mostly focused on the prediction, risk assessment, early warning, and efficacy evaluation of a particular disease. They are mostly based on relevant medical data to complete the risk assessment or efficacy evaluation of a certain disease. They lack a comprehensive assessment of the health status of patients applicable to all diseases, fail to identify disease risk factors in a timely manner, and lack comparative evaluation and analysis of patients before and after treatment. Summary of the Invention

[0003] To overcome the aforementioned shortcomings in existing individualized health assessments for patients, this invention provides an individualized health assessment method and device based on multimodal patient information.

[0004] This invention provides a personalized health assessment device based on patient multimodal information, the device comprising a memory and a processor; The memory is used to store computer programs; the processor is used to call and execute the computer programs, so that the device performs a personalized health assessment method based on patient multimodal information, the personalized health assessment method including: Obtain multimodal parameters of patients across multiple consultations; By selecting multimodal parameters in each visit in pairs, and combining the selected multimodal parameters and their corresponding reciprocals, the quantitative differences of each combination are calculated. Based on the combination corresponding to the maximum quantitative difference, the geometric mean of the corresponding number of visits is calculated to obtain the health scale; The patient's health was assessed based on the health standards used in each consultation. The multimodal parameters include the patient's symptom data, tongue image data, facial image data, audio data, and pulse image data.

[0005] According to one specific implementation, in the above-mentioned assessment device, the symptom data and pulse data are obtained by encoding the patient's symptom text and pulse diagnosis text according to a unique code and assigning values ​​according to the severity of the symptoms.

[0006] According to one specific implementation, in the aforementioned evaluation device, the tongue image data and facial image data are digital tongue diagnosis and digital facial diagnosis obtained by annotating and extracting features from the patient's tongue image images and facial image images respectively using image analysis technology, specifically including: Label the tongue color, coating color, tongue body, and tongue coating in the tongue image; label the face color, lip color, and facial luster in the face image. The labeled tongue images and the labeled facial images were standardized and augmented. The processed tongue image is input into a pre-trained deep learning network to extract tongue color features, coating color features, tongue body features, and tongue coating features to obtain digital tongue diagnosis. The processed facial images are input into a pre-trained deep learning network to extract facial color features, lip color features, and facial luster features, resulting in a digital facial diagnosis.

[0007] According to one specific implementation, in the above-mentioned assessment device, the audio data is digitized audio obtained by denoising and feature extraction of patient audio; wherein, the patient audio includes heart sounds, respiratory sounds, speech, and tissue vibrations.

[0008] According to one specific implementation, in the above-mentioned evaluation device, the formula for calculating the quantitative difference is:

[0009] Where x1 and x 22 These are the two multimodal parameters or their corresponding reciprocals in each combination.

[0010] According to one specific implementation, the above-mentioned assessment device calculates the geometric mean of the corresponding number of diagnoses, specifically including: Calculate the geometric mean of the symptom data in the corresponding consultation sessions to obtain the geometric mean of each symptom; The geometric mean of each symptom is combined with various parameters to calculate the geometric mean of the corresponding number of visits. The formula is as follows: a=e, Among them, SGM l For the first l The geometric mean of the number of diagnoses, where N is the number of multimodal parameters, x i Let be the i-th multimodal parameter.

[0011] According to one specific embodiment, in the above-mentioned evaluation equipment, the evaluation method further includes: Normalize the same multimodal parameter for a patient across multiple visits; The proportion of a single multimodal parameter in each diagnosis session is calculated based on the normalized multimodal parameters. If the calculation results based on the proportion do not meet the requirements of consistency and continuity, an early warning will be issued; the early warning is used to indicate that the patient's multimodal parameters are not accurate.

[0012] According to one specific embodiment, in the above-mentioned evaluation equipment, the evaluation method further includes: Under the condition that the calculation results based on the proportion meet the requirements of consistency and continuity, the normalized multimodal parameters are divided by the geometric mean of the corresponding number of diagnoses to obtain the easy flow model. Based on the aforementioned easy-flow model, parameter hole calculations are performed to obtain key cross-modal parameters; The patient's health score is calculated based on the number of key crossmodal parameters and the number of multimodal parameters.

[0013] According to one specific implementation, in the above evaluation method, the calculation of the parameter hole includes: The patient's multiple visits are divided into pre-treatment and post-treatment stages according to preset rules; The quantitative difference between each multimodal parameter and 1 in the pre-treatment easy flow model is calculated to obtain the first parameter hole; The quantitative difference between each multimodal parameter and 1 in the easy flow model during the later stage of treatment is calculated to obtain the second parameter hole; If the first parameter hole and the second parameter hole are greater than a preset threshold, the corresponding parameter will be used as the key cross-modal parameter.

[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention collects and integrates multimodal medical data from patients, including disease-related question-and-answer text information, tongue and facial images, and voice information. Utilizing body measurements and geometric mean theory, it constructs a personalized health scale for each patient, enabling a quantitative assessment of their overall health status. By constructing health scales for patients at different treatment stages, it helps doctors intuitively understand the trend of patients' health status, thus aiding in disease treatment. This invention utilizes the longitudinal and lateral normalization of multidimensional patient data to obtain an individualized model, and uses the principle of quantitative difference to obtain parameter holes (i.e. disease risk factors), helping doctors to identify disease risk factors in a timely manner for causal treatment, thereby improving the quality of medication use and shortening the treatment cycle for patients. Attached Figure Description

[0015] Figure 1 This is a flowchart illustrating the personalized health assessment method based on patient multimodal information provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of comparative evaluation and analysis provided for an embodiment of the present invention. Detailed Implementation

[0016] The present invention will now be described in further detail with reference to specific embodiments. However, this should not be construed as limiting the scope of the present invention to the following embodiments; all technologies implemented based on the content of the present invention fall within the scope of the present invention.

[0017] This invention provides a personalized health assessment device based on patient multimodal information. The device includes a memory and a processor. The memory stores a computer program, and the processor calls and executes the computer program to enable the device to perform a personalized health assessment method based on patient multimodal information.

[0018] Please refer to Figure 1 This illustration shows a flowchart of a personalized health assessment method based on patient multimodal information provided in an embodiment of the present invention. The method includes: Step 1: Based on the patient's multimodal parameters across multiple visits, calculate the quantitative differences between the various multimodal parameters in each visit.

[0019] The multimodal parameters include the patient's symptom data, tongue image data, facial image data, audio data, and pulse image data.

[0020] According to one specific implementation, the symptom data and pulse data are obtained by encoding the patient's symptom text and pulse diagnosis text according to a unique code, and assigning values ​​according to the severity of the symptoms.

[0021] Specifically, to ensure the accuracy of subsequent calculations, 0 in the unique hot code is replaced with 1, and 1 is replaced with 2. The severity of symptoms is assigned according to (severe: 3, moderate: 2, mild: 1) to complete the digitization and standardization of symptoms and pulse diagnosis.

[0022] According to a specific implementation, in the above evaluation method, the tongue image data and facial image data are digital tongue diagnosis and digital facial diagnosis obtained by annotating and extracting features from the patient's tongue image images and facial image images respectively using image analysis technology, specifically including: Label the tongue color, coating color, tongue body, and tongue coating in the tongue image; label the face color, lip color, and facial luster in the face image. The labeled tongue images and the labeled facial images were standardized and augmented. The processed tongue image is input into a pre-trained deep learning network to extract tongue color features, coating color features, tongue body features, and tongue coating features to obtain digital tongue diagnosis. The processed facial images are input into a pre-trained deep learning network to extract facial color features, lip color features, and facial luster features, resulting in a digital facial diagnosis.

[0023] Specifically, the labelimg tool was used to annotate the tongue image information in four dimensions.

[0024] The four dimensions are as follows: Tongue color: pale white, pale red, red, crimson, bluish-purple Moss color: white, yellow, grayish-black Tongue body: old, young, fat, thin, dots, prickles, cracks, teeth marks Tongue coating: thin, thick, moist, dry, greasy, rotten, peeling, partial, complete, true, false Tongue color labeling: For tongue color classification, a rectangular box is used to label the tongue area, and the corresponding tongue color category (pale white, light red, red, crimson, bluish-purple) is recorded in the labeling file.

[0025] Tongue coating color labeling: Use a rectangle to label the tongue coating area and record the coating color category (white, yellow, gray-black).

[0026] Tongue annotation: Based on the different characteristics of the tongue, the polygon annotation method is used to accurately annotate the position and shape of the tongue, such as old, young, fat, thin, dots, thorns, cracks, and teeth marks.

[0027] Tongue coating feature annotation: For tongue coating features such as thin, thick, moist, dry, greasy, rotten, peeling, off-center, complete, true, and false, semantic segmentation annotation method is used to annotate the tongue coating area at the pixel level and record the tongue coating feature category corresponding to each pixel.

[0028] To ensure the accuracy of subsequent image analysis, the tongue images underwent a series of processing steps, including standardization, enhancement, and color space conversion.

[0029] Standardization processing: In order to reduce the influence of external factors (such as lighting, angle, etc.) on the analysis, the original image needs to be cropped, rotated, scaled and other operations to ensure that the tongue occupies the center position of the image, and the bicubic interpolation method (cv2.INTER_CUBIC) is used to adjust the image to a uniform size, such as 384×384 pixels.

[0030] Augmentation: Data augmentation techniques (such as random flipping, color jittering, etc.) are used to increase sample diversity and help the model generalize better.

[0031] Color space conversion: Sometimes RGB images are converted to other color spaces (such as HSV) in order to extract color features more effectively.

[0032] Furthermore, ResNet50 deep learning technology was used to automatically extract tongue image features. The first four convolutional stages were retained, the fully connected layers were removed, and the output feature map size was 12×12×2048 (corresponding to 384×384 input).

[0033] Tongue color and tongue coating feature extraction: The feature map (384→96×96×256) is output in the conv2 stage of the backbone network, and the dimension is reduced to 3 channels (corresponding to RGB color components) through 1×1 convolution. Global average pooling is used to obtain the color mean feature. Tongue feature learning: In the conv3 stage, candidate bounding boxes for tongue edges are generated through RPN (Region Proposal Network). The aspect ratio of the bounding boxes is calculated (e.g., fat tongue > 0.85, thin tongue < 0.65) and the edge curvature (e.g., absolute value of curvature of the teeth mark region > 0.5). Cracks are judged by combining the mean distance of the contour points (e.g., distance of contour points in the crack region > 0.1 × tongue radius). Tongue coating feature learning: Multi-scale convolution is used in the conv5 stage to extract tongue coating texture and calculate texture entropy (e.g., greasy coating entropy value < 3.0, rotten coating > 4.0) and energy value (e.g., dry coating energy value < 0.2, moist coating > 0.4).

[0034] The above tongue color characteristics (1 attribute), tongue coating characteristics (1 attribute), tongue body characteristics (3 attributes), and tongue coating characteristics (2 attributes), totaling 4 characteristics and 7 attributes, are combined to form a digital set of tongue diagnosis.

[0035] Furthermore, the labelimg tool was used to annotate the tongue image information in three dimensions.

[0036] The three dimensions are as follows: Complexion, lip color, and facial luster; among which, complexion can be categorized as normal, red, white, black, yellow, or bluish. Lip color: pale white, light red, red, dark red, purple Facial luster: shiny, less shiny, dull.

[0037] To ensure the accuracy of subsequent image analysis, the tongue images underwent a series of processing steps, including standardization, enhancement, and color space conversion.

[0038] Standardization processing: In order to reduce the influence of external factors (such as lighting, angle, etc.) on the analysis, the original image needs to be cropped, rotated, scaled and other operations to ensure that the tongue occupies the center position of the image, and the bicubic interpolation method (cv2.INTER_CUBIC) is used to adjust the image to a uniform size, such as 384×384 pixels.

[0039] Augmentation: Data augmentation techniques (such as random flipping, color jittering, etc.) are used to increase sample diversity and help the model generalize better.

[0040] Color space conversion: Sometimes RGB images are converted to other color spaces (such as HSV) in order to extract color features more effectively.

[0041] Furthermore, ResNet50 deep learning technology was used to automatically extract tongue image features. The first four convolutional stages were retained, the fully connected layers were removed, and the output feature map size was 12×12×2048 (corresponding to 384×384 input).

[0042] Facial color feature extraction: The feature map output in the 4th stage of the backbone network (448→56×56×512) is reduced to 3 channels (corresponding to RGB skin color components) through 1×1 convolution. Region weighted average pooling (1.2 weight for the center region of the face and 0.8 weight for the edge region) is used to obtain the color mean features (e.g., the mean value of the red component of red face color is >0.62, and the mean value of the blue component of cyan face color is >0.45).

[0043] Lip color feature extraction: Based on the lip mask obtained from segmentation, the lip region of the 5th stage feature map (28×28×1024) of the backbone network is cropped, and the dimensionality is reduced to 3 channels through 1×1 convolution. Global average pooling is used to obtain lip color features (such as the mean value of red component of light white lips <0.35, and the red component of dark red lips >0.55 and the blue component >0.3).

[0044] Facial gloss feature extraction: Analysis was performed on the L channel (brightness) of the LAB color space. Highlight regions were extracted by Gaussian difference filtering, and the proportion of highlight area was calculated (glossy > 15%, matte < 5%). Combined with the texture features of the third stage of the backbone network (112×112×256), gloss uniformity was captured by multi-scale convolution (3×3 / 7×7) (variance of low gloss regions > 0.05).

[0045] The above three features and three attributes—facial color feature (1 attribute), lip color feature (1 attribute), and facial luster feature (1 attribute)—form a digital set for facial diagnosis.

[0046] According to one specific implementation, in the above evaluation method, the audio data is digitized audio obtained by denoising and feature extraction of patient audio; wherein, the patient audio includes heart sounds, respiratory sounds, speech, and tissue vibrations.

[0047] Specifically, techniques such as Fourier transform are used to perform noise reduction and feature extraction on the patient's audio, resulting in a digital result for patient voice recognition.

[0048] First, preprocessing is performed, including the following steps: 1) Noise reduction Static noise suppression: The spectral subtraction method is used. By collecting 1-2 seconds of pure noise samples (such as ambient noise before the collection), the noise power spectrum is estimated, and the noise spectrum is subtracted from the original signal spectrum in the frequency domain.

[0049] Formula (1) in, As an over-reduction factor, it is generally taken as 1.2-1.5.

[0050] Dynamic noise filtering: For low-frequency signals such as heart sounds / breath sounds, use a Butterworth low-pass filter (cutoff frequency 2000Hz); for speech signals, use a band-pass filter (300-3400Hz) to preserve the main frequency bands; Wavelet denoising: For non-stationary noise (such as sudden interference), a 5-level decomposition using the db6 wavelet basis is employed, and soft thresholding is applied to the high-frequency coefficients (thresholding). , (where N is the noise standard deviation and N is the signal length), reconstruct the signal to preserve transient characteristics.

[0051] 2) Signal segmentation and alignment Automatic segmentation: Heart sounds are segmented by cardiac cycle, respiratory sounds are segmented by respiratory cycle, and speech is segmented by syllable / word; Time alignment: For multiple signals of the same type (such as multiple heart sounds from the same patient), the waveforms are aligned using a dynamic time warping algorithm to eliminate time shifts caused by individual differences.

[0052] 3) Normalization processing Z-Score standardization is adopted. (For the signal mean and standard deviation), ensuring that the feature scales of different samples are consistent.

[0053] Furthermore, audio feature extraction is performed. Based on the signal type (heart sounds, breath sounds, speech, tissue vibration), Fourier transform and other techniques are used to extract targeted features, covering time domain, frequency domain, time-frequency domain and nonlinear features.

[0054] 1) Extraction of heart sound related features Heart sound signals contain hallmark components such as S1 (start of systole) and S2 (start of diastole). Feature extraction needs to focus on changes in period, amplitude, and spectrum. Temporal characteristics: Cyclic characteristics: duration of cardiac cycle, S1-S2 interval (systole), S2-S1 interval (diastole); Amplitude characteristics: S1 peak amplitude, S2 peak amplitude, S1 / S2 amplitude ratio (normally about 1.2-1.5), signal energy (integral sum of squares); Morphological characteristics: S1 rise time (duration from base value to peak value), S2 fall slope (slope within 10ms after peak value).

[0055] Frequency domain characteristics: Power spectrum analysis: The power distribution in the 0-2000Hz frequency band was calculated by FFT. The main frequency of S1 is concentrated in 50-250Hz, and the main frequency of S2 is concentrated in 100-300Hz.

[0056] Time-frequency domain characteristics: Short-time Fourier transform: window length 256ms, overlap rate 50%, extracting spectrogram features (such as energy concentration regions) in time periods S1 and S2.

[0057] 2) Extraction of breath sound related features Breath sounds include inspiratory phase, expiratory phase, and abnormal sounds (such as rales and wheezing). It is important to distinguish between the breathing pattern and the abnormal components. Temporal characteristics: Cyclic characteristics: respiratory rate (number of breaths per minute), inspiratory duration, expiratory duration, inspiratory-expiratory ratio (normal is about 1:1.5-2); Amplitude characteristics: peak inspiratory amplitude, peak expiratory amplitude, and breath sound intensity; Abnormal sound detection: Identification is performed by short-term energy mutations (rallies) or periodic pulses (wheezing), and the proportion of abnormal sounds is calculated.

[0058] Frequency domain characteristics: Frequency energy ratio: The energy ratio of high frequency (500-2000Hz) to low frequency (50-500Hz) during the inspiratory phase (normal <1.0), with a significant increase in the high frequency energy of wheezing; Spectral entropy ( : Reflects spectral complexity, abnormal respiratory spectral entropy > 0.8 (normal < 0.6).

[0059] Time-frequency domain characteristics: Wavelet packet decomposition: A 6-level decomposition based on the db4 wavelet basis was used to calculate the energy proportion of each frequency band. The rallies were prominent in the 250-500Hz range, and the wheezing was prominent in the 500-2000Hz range.

[0060] 3) Extraction of speech-related features Speech signals are related to the functions of the vocal organs (larynx, lungs, and oral cavity), and their features focus on acoustic parameters and prosodic characteristics.

[0061] Acoustic characteristics: Fundamental frequency (F0): Extracted by autocorrelation algorithm, reflecting the frequency of vocal cord vibration (85-180Hz for males, 165-255Hz for females). When abnormal, the standard deviation of F0 is >50Hz. Formants (F1-F4): Extracted through LPC analysis, reflecting the shape of the vocal tract; Energy and duration: syllable energy, vowel duration (normal 100-300ms), speech rate (syllables / second).

[0062] Spectral characteristics: The spectral envelope of speech is captured using 13-dimensional MFCC and its first and second-order differences (39 dimensions in total).

[0063] 4) Extraction of tissue vibration-related features Tissue vibration signals (such as vascular murmurs and joint friction sounds) are weak and require targeted enhancement and feature extraction. Temporal characteristics: Vibration period (e.g., the regularity of the systolic / diastolic phase of a vascular murmur), peak interval, and amplitude variation coefficient; Frequency domain characteristics: Dominant frequency (e.g., the dominant frequency of joint friction noise is 100-500Hz), harmonic components (whether there are integer multiples of frequency). Nonlinear characteristics: Sample entropy (SampEn): measures signal complexity. Normal tissue vibration SampEn < 1.5, abnormal tissue vibration SampEn > 2.0; Fractal dimension (calculated by box counting): reflects the irregularity of the vibration waveform. The fractal dimension increases in pathological conditions.

[0064] The above features, including heart sound-related features (12 attributes), breath sound-related features (11 attributes), speech-related features (44 attributes), and tissue vibration-related features (7 attributes), totaling 4 features and 74 attributes, are used to form a digital set for patient audio analysis.

[0065] Based on the above steps, the patient's symptoms, pulse diagnosis, tongue diagnosis, facial diagnosis, and audio information are digitized, realizing the digitization and fusion of multimodal patient information. Combined with digitized laboratory indicators, this constitutes integrated TCM and Western medicine data from multiple patient visits (symptoms, pulse diagnosis, tongue diagnosis, facial diagnosis, audio, and laboratory indicators).

[0066] Step 2: Select multimodal parameters from each visit in pairs, and combine them with their corresponding reciprocals to calculate the quantitative difference of the corresponding combinations.

[0067] Understandably, based on the selected multimodal parameters, four combinations can be generated using these parameters and their corresponding reciprocals: multimodal parameter 1 and multimodal parameter 2, the reciprocal of multimodal parameter 1 and multimodal parameter 2, the reciprocals of multimodal parameter 1 and multimodal parameter 2, and the reciprocal of multimodal parameter 1 and multimodal parameter 2. The quantitative difference for each combination is then calculated. Specifically, the formula for calculating the quantitative difference is: (Formula 2) Here, x1 and x2 are the two multimodal parameters or their corresponding reciprocals in each combination.

[0068] Specifically, the pairwise selections mentioned above can be random or sequential, ultimately yielding quantitative differences between various multimodal parameters.

[0069] Step 3: Calculate the geometric mean of the corresponding number of visits based on the combination corresponding to the maximum quantitative difference, and obtain the health scale.

[0070] It is understandable that the quantitative differences calculated in each combination in step 2, the combination corresponding to the maximum quantitative difference includes one of four combinations, that is, taking the reciprocal of each multimodal parameter or not taking the reciprocal, and continuing the subsequent calculation according to this form.

[0071] Specifically, using the body measurement theory, the SGM of each symptom is generated according to formula (3) for each patient's symptoms and their severity at each consultation. Then, the various parameters (symptom SGM, tongue diagnosis digitization result, face diagnosis digitization result, audio digitization result, and multiple test indicators) of each patient at each consultation are calculated to obtain the structural geometric mean SGM of each patient at each consultation.

[0072] Formula (3) Among them, SGM l For the first l The geometric mean of the number of diagnoses, where N is the number of multimodal parameters, x i It is the i-th multimodal parameter or its corresponding reciprocal.

[0073] Take the reciprocal of each of the N parameters from the M patient visits, calculate and compare the QDMM after each calculation. When the QDMM reaches its maximum value, retain the form of that set of parameters (either by taking the reciprocal or not). The structural geometric mean (SGM) of each visit when the QDMM reaches its maximum value is the sum of the ten periods (GT). The SGM that satisfies this condition is the GT.

[0074] Step 4: Conduct a health assessment of the patient based on the health standards used in each visit.

[0075] Understandably, the obtained health scale can characterize a patient's health status. An upward trend in the health scale across different visits indicates that the patient is becoming healthier; conversely, a downward trend suggests a need to monitor the patient's health and implement targeted treatment based on the corresponding multimodal parameters.

[0076] Furthermore, according to a specific implementation, the above evaluation method also includes means for handling problems with individual multimodal parameters in each consultation, specifically including: Normalize the same multimodal parameter for a patient across multiple visits; The proportion of a single multimodal parameter in each diagnosis session is calculated based on the normalized multimodal parameters. If the calculation results based on the proportion do not meet the requirements of consistency and continuity, an early warning will be issued; the early warning is used to indicate that the patient's multimodal parameters are not accurate.

[0077] Specifically, an individualized model is obtained by using longitudinal and lateral normalization of multidimensional patient data, and parameter holes (i.e. disease risk factors) are obtained by using the principle of quantitative difference, which helps to identify disease risk factors in a timely manner, as well as multimodal parameters that lead to a decline in the patient's health status.

[0078] 1. Longitudinal normalization: Calculate the functional geometric mean (FGM) of each parameter for M visits to the patient, and divide each parameter value by the functional geometric mean of that parameter.

[0079] 2. Calculate the proportion: Calculate the sum of all parameter values ​​for each measurement after longitudinal normalization, and then calculate the proportion of each individual parameter in the total parameter sum for each measurement.

[0080] 3. Calculate the structural geometric mean of the proportions: Calculate the structural geometric mean of all proportions in each consultation.

[0081] 4. Determine conservation: If the geometric mean of the M proportions satisfies the conservation of consistency (4) and continuity (5), the geometric mean of the structures that satisfies (4) and (5) is called the golden center (GC). If it does not satisfy (4) and (5), the data of that patient may be inaccurate and cannot be used in subsequent calculations.

[0082] The consistency of QD, where QD is less than α for the geometric mean (indicated by an overline) at each time step and for all time steps, is called the conservation of consistency. Formula (4) If we sort the M time intervals, and QD is less than a in any two consecutive time intervals, then we have a conserved continuity. Formula (5) According to one specific implementation, the above evaluation method further includes: Under the condition that the calculation results based on the proportion meet the requirements of consistency and continuity, the normalized multimodal parameters are divided by the geometric mean of the corresponding number of diagnoses to obtain the easy flow model. Based on the aforementioned easy-flow model, parameter hole calculations are performed to obtain key cross-modal parameters; The patient's health score is calculated based on the number of key crossmodal parameters and the number of multimodal parameters.

[0083] Specifically, this embodiment provides a method for calculating a health score:

[0084] Where HS is the patient's health score at a certain consultation, m is the number of key cross-modal parameters, and n is the number of multimodal parameters at that consultation.

[0085] According to one specific implementation, in the above evaluation method, the calculation of the parameter hole includes: The quantitative difference between each parameter and 1 in the pre-treatment flow model is calculated to obtain the first parameter hole; The quantitative difference between each parameter and 1 in the easy flow model during the later stage of treatment is calculated to obtain the second parameter hole; If the first parameter hole and the second parameter hole are greater than a preset threshold, the corresponding parameter will be used as the key cross-modal parameter.

[0086] Specifically, after vertical normalization, each parameter for each consultation is divided by a factor of the N parameters for that consultation. After both vertical and horizontal normalization, each parameter for each consultation is normalized to the same level, and the resulting data stream is called the patient's easy-flow model.

[0087] For the data after two normalizations (both longitudinal and transverse), use formula (2) to calculate the QD of each parameter with 1 in the early and late stages of treatment. Compare the calculated QD with β. If it is greater than β, the parameter is a parameter hole, indicating that the patient needs to pay close attention to this parameter. If it is less than β, the parameter is not a parameter hole, and the patient can continue treatment as before.

[0088] Furthermore, based on parameter-based causal treatment, a comparative evaluation of the patients' overall health status before and after treatment was completed through comparative analysis of multidimensional health scales and quantitative differences in the pre- and post-treatment periods. Please refer to [reference needed]. Figure 2 This illustrates a comparative evaluation analysis diagram provided by an embodiment of the present invention.

[0089] Through the above steps, calculate the GT of the corresponding clinic in the later stage of treatment (compare the difference between the GT in the early stage of treatment and the later stage of treatment; the larger the GT, the healthier the patient is, that is, the GT after treatment should be larger than before). Use formula (2) to calculate the QD of the GT in the early stage of treatment and the later stage of treatment. If it is greater than β, it proves that the treatment has a significant effect.

[0090] In one possible implementation, multiple patient visits are divided into pre-treatment and post-treatment phases. This division can be done by allocating the first third of the visits to the pre-treatment phase and the last two-thirds to the post-treatment phase. If a phase cannot be divided, the data can be rounded up according to a preset rule. For example, if there are four visits, the pre-treatment phase would be the first visit, and the post-treatment phase would be the last three visits.

[0091] Furthermore, the above α and β are the values ​​of the Weber threshold according to Noether's theorem, where α is 0.268259 and β is 0.8648.

[0092] Based on the above technical solution, this invention collects and integrates multimodal medical data from patients, including disease-related question-and-answer text information, tongue and facial images, and voice information. Utilizing body measurements and geometric mean theory, it constructs a personalized health scale for each patient, enabling a quantitative assessment of their overall health status. By constructing health scales for patients at different treatment stages, doctors can intuitively understand the trend of patients' health status, which is beneficial for disease treatment.

[0093] Furthermore, this invention utilizes the longitudinal and lateral normalization of multidimensional patient data to obtain an individualized model, and uses the principle of quantitative difference to obtain parameter holes (i.e. disease risk factors), helping doctors to identify disease risk factors in a timely manner for causal treatment, thereby improving the quality of medication use and shortening the treatment cycle for patients.

[0094] In addition, in embodiments of the present invention, the processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0095] The various methods, steps, and logic diagrams disclosed in the embodiments of this invention can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The processor reads information from the storage medium and, in conjunction with its hardware, completes the steps of the above methods.

[0096] The storage medium can be memory, such as volatile memory or non-volatile memory, or may include both volatile and non-volatile memory.

[0097] Among them, non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory.

[0098] Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), sync link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM).

[0099] Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0100] Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner for ease of understanding.

[0101] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A personalized health assessment device based on patient multimodal information, characterized in that, The device includes a memory and a processor; The memory is used to store computer programs; the processor is used to call and execute the computer programs, so that the device performs a personalized health assessment method based on patient multimodal information, the personalized health assessment method including: Obtain multimodal parameters of patients across multiple consultations; By selecting multimodal parameters in each visit in pairs, and combining the selected multimodal parameters and their corresponding reciprocals, the quantitative differences of each combination are calculated. Based on the combination corresponding to the maximum quantitative difference, the geometric mean of the corresponding number of visits is calculated to obtain the health scale; The patient's health was assessed based on the health standards used in each consultation. The multimodal parameters include the patient's symptom data, tongue image data, facial image data, audio data, and pulse image data.

2. The personalized health assessment device based on patient multimodal information according to claim 1, characterized in that, The symptom data and pulse data are obtained by the device encoding the patient's symptom text and pulse diagnosis text according to a unique code and assigning values ​​according to the severity of the symptoms.

3. The personalized health assessment device based on patient multimodal information according to claim 1, characterized in that, The tongue image data and facial image data are digital tongue diagnosis and digital facial diagnosis obtained by the device through image analysis technology, which involves annotation and feature extraction of the patient's tongue image and facial image, respectively. Specifically, they include: Label the tongue color, coating color, tongue body, and tongue coating in the tongue image; label the face color, lip color, and facial luster in the face image. The labeled tongue images and the labeled facial images were standardized and augmented. The processed tongue image is input into a pre-trained deep learning network to extract tongue color features, coating color features, tongue body features, and tongue coating features to obtain digital tongue diagnosis. The processed facial images are input into a pre-trained deep learning network to extract facial color features, lip color features, and facial luster features, resulting in a digital facial diagnosis.

4. The personalized health assessment device based on patient multimodal information according to claim 1, characterized in that, The audio data is digitized audio obtained by the device through noise reduction and feature extraction of the patient's audio; wherein, the patient's audio includes heart sounds, respiratory sounds, speech, and tissue vibrations.

5. The personalized health assessment device based on patient multimodal information according to claim 1, characterized in that, The formula for calculating the quantitative difference is: , Here, x1 and x2 are the two multimodal parameters or their corresponding reciprocals in each combination.

6. The personalized health assessment device based on patient multimodal information according to claim 1, characterized in that, Calculate the geometric mean of the corresponding number of visits, specifically including: Calculate the geometric mean of the symptom data in the corresponding consultation sessions to obtain the geometric mean of each symptom; The geometric mean of each symptom is combined with various parameters to calculate the geometric mean of the corresponding number of visits. The formula is as follows: ,a=e, Among them, SGM l For the first l The geometric mean of the number of diagnoses, where N is the number of multimodal parameters, x i Let be the i-th multimodal parameter.

7. The personalized health assessment device based on patient multimodal information according to claim 1, characterized in that, The individualized health assessment method also includes: Normalize the same multimodal parameter for a patient across multiple visits; The proportion of a single multimodal parameter in each diagnosis session is calculated based on the normalized multimodal parameters. If the calculation results based on the proportion do not meet the requirements of consistency and continuity, an early warning will be issued; the early warning is used to indicate that the patient's multimodal parameters are not accurate.

8. The personalized health assessment device based on patient multimodal information according to claim 1, characterized in that, The individualized health assessment method also includes: Under the condition that the calculation results based on the proportion meet the requirements of consistency and continuity, the normalized multimodal parameters are divided by the geometric mean of the corresponding number of diagnoses to obtain the easy flow model. Based on the aforementioned easy-flow model, parameter hole calculations are performed to obtain key cross-modal parameters; The patient's health score is calculated based on the number of key crossmodal parameters and the number of multimodal parameters.

9. A personalized health assessment device based on patient multimodal information according to claim 8, characterized in that, The calculation of the parameter hole includes: The patient's multiple visits are divided into pre-treatment and post-treatment stages according to preset rules; The quantitative difference between each multimodal parameter and 1 in the pre-treatment flow model is calculated to obtain the first parameter hole; The quantitative difference between each multimodal parameter and 1 in the easy flow model during the later stage of treatment is calculated to obtain the second parameter hole; If the first parameter hole and the second parameter hole are greater than a preset threshold, the corresponding parameter will be used as the key cross-modal parameter.