Traditional Chinese medicine cardiovascular risk assessment system based on pulse condition and tongue condition
By simultaneously acquiring and deeply correlated multimodal data, and combining non-invasive sensing of microenvironment biochemical indicators, a cardiovascular risk assessment system with pathophysiological interpretability was constructed. This system solves the problem of insufficient tongue image analysis in existing technologies and enables dynamic monitoring and early warning of cardiovascular risks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-19
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies for tongue image analysis lack the ability to detect non-invasive and simultaneous biochemical indicators of the tongue surface microenvironment. The fusion depth of pulse and tongue image features is insufficient, making it impossible to achieve dynamic and continuous monitoring and early warning. This results in insufficient biological basis for cardiovascular risk assessment and weak interpretability of assessment results.
The system employs a multimodal data synchronous acquisition module, a non-invasive sensing module for microenvironment biochemical indicators, a multimodal feature deep correlation analysis module, a pathophysiological mechanism mapping and risk assessment module, and a dynamic monitoring and early warning module to achieve multimodal deep fusion and dynamic monitoring of pulse and tongue images.
It enables non-invasive and simultaneous detection of cardiovascular risk biochemical markers in saliva on the tongue surface, establishes a deep correlation model between pulse, tongue appearance and biochemical indicators, provides a clear pathophysiological explanation, realizes dynamic monitoring and early warning of cardiovascular risk, and improves the objectivity and clinical interpretability of assessment results.
Smart Images

Figure CN121839136A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of medical data processing, and particularly relates to a traditional Chinese medicine cardiovascular risk assessment system based on pulse and tongue manifestations. BACKGROUND
[0002] Artificial intelligence technology is increasingly widely used in the field of medical health, and particularly shows great potential in disease risk assessment and auxiliary diagnosis.
[0003] Traditional Chinese medicine diagnosis, as a traditional medicine based on the overall view and syndrome differentiation, its objectification and quantification research is the key direction to promote its modernization development. Pulse and tongue manifestations, as the core information source of traditional Chinese medicine diagnosis, contain rich characteristics reflecting the physiological and pathological state of the human body.
[0004] Based on pulse and tongue information, the risk assessment of cardiovascular diseases is an important research direction of intelligent diagnosis of traditional Chinese medicine. This technical direction aims to collect and analyze the characteristics of pulse waveform, rhythm, and strength, as well as the visual characteristics of tongue color, shape, and fur quality, to construct a mathematical model to quantitatively assess the risk level of individuals suffering from cardiovascular diseases.
[0005] The existing technology usually uses independent pulse instrument and tongue collection equipment to obtain data, and uses feature extraction and machine learning model for risk prediction. However, the existing technology has significant limitations: tongue analysis mainly relies on surface visual information, lacks non-invasive and synchronous detection capability of biochemical indicators (such as inflammatory factors and oxidative stress markers) in the tongue microenvironment (such as saliva) closely related to cardiovascular risk, resulting in insufficient biological basis and limited depth of risk assessment. At the same time, the fusion of pulse and tongue characteristics is mostly limited to the data level, and the correlation model between them and the deep pathological and physiological mechanism (such as vascular endothelial function and autonomic nervous regulation) has not been effectively established, making the interpretability and specificity of the evaluation results not strong. In addition, the existing system is usually based on static one-time detection, and it is difficult to realize continuous monitoring and early warning of the dynamic changes of cardiovascular risk.
[0006] Therefore, how to realize the multi-modal deep fusion of pulse, tongue macroscopic characteristics and microenvironment biochemical indicators, and construct a cardiovascular risk assessment model with clear pathological and physiological interpretation, has become a technical problem to be solved in this field. SUMMARY
[0007] The present application provides a traditional Chinese medicine cardiovascular risk assessment system based on pulse and tongue manifestations to solve the problems of lack of non-invasive and synchronous detection capability of biochemical indicators in the tongue microenvironment, insufficient depth of pulse and tongue characteristic fusion, and inability to realize dynamic and continuous monitoring and early warning in the prior art.
[0008] The technical scheme of the present application is a traditional Chinese medicine cardiovascular risk assessment system based on pulse and tongue manifestations, which comprises a multi-modal data synchronous acquisition module, a microenvironment biochemical index non-invasive sensing module, a multi-modal feature deep correlation analysis module, a pathological and physiological mechanism mapping and risk assessment module, and a dynamic monitoring and early warning module.
[0009] The multi-modal data synchronous acquisition module is used to synchronously acquire the pulse original signal and the tongue original image of a user under a unified time reference. The multi-modal data synchronous acquisition module is integrated with a high-precision pressure sensor array and a multi-spectral imaging unit. The pressure sensor array acquires pulse wave pressure signals of the Cun, Guan and Chi sections at a specific sampling frequency. The multi-spectral imaging unit acquires a digital image sequence of the tongue body region under specific wavelength combinations of illumination conditions, and the digital image sequence contains information of the visible light band and at least two near-infrared bands.
[0010] The microenvironment biochemical index non-invasive sensing module is integrated with the tongue image acquisition end of the multi-modal data synchronous acquisition module, and is used to non-invasively and in situ detect the concentration of specific biochemical indicators related to cardiovascular risk in the salivary film on the tongue surface. The microenvironment biochemical index non-invasive sensing module comprises a microfluidic sensing chip array and a surface plasmon resonance optical detection unit. The microfluidic sensing chip array passively adsorbs and enriches the tongue surface saliva through capillary action, and the surface is fixed with biological probe molecules that specifically bind to the target biochemical indicators. The surface plasmon resonance optical detection unit then monitors the change in the refractive index of the chip surface caused by the binding of the biological probe molecules to the target biochemical indicators in real time, and converts the change in the refractive index of the chip surface into an electrochemical concentration signal of the corresponding biochemical indicators. The target biochemical indicators at least include C-reactive protein, interleukin 6 and malondialdehyde.
[0011] The multi-modal feature deep correlation analysis module is used to receive and process the original data from the multi-modal data synchronous acquisition module and the microenvironment biochemical index non-invasive sensing module, extract and correlate multi-dimensional features. The multi-modal feature deep correlation analysis module first preprocesses the pulse wave pressure signal, including baseline drift correction and power frequency interference removal, and then extracts time domain features, frequency domain features and nonlinear dynamics features. The time domain features include the main wave amplitude, the double wave amplitude, the time interval between the main wave and the double wave, and the pulse rate variability index. The frequency domain features are calculated by fast Fourier transform, including the ratio of low frequency power to high frequency power. The nonlinear dynamics features include sample entropy and Lyapunov exponent.
[0012] Simultaneously, the multimodal feature deep correlation analysis module segments and extracts features from the original tongue image. The segmentation process employs a semantic segmentation algorithm based on a deep convolutional neural network to accurately separate the tongue region from the background and teeth regions. Color features, texture features, and morphological features are extracted from the segmented tongue image. Color features include the average chromaticity and saturation values of the tongue body in a specific color space, as well as the quantified values of the thickness and uniformity of the tongue coating. Texture features are calculated using a gray-level co-occurrence matrix, including contrast and energy values. Morphological features include the tongue's plumpness index and the depth and number of teeth marks on the edges. Further, the multimodal feature deep correlation analysis module performs time alignment and normalization on the extracted pulse feature set, tongue visual feature set, and microenvironment biochemical indicator concentration set to construct a multidimensional feature vector. Subsequently, the multimodal feature deep correlation analysis module performs correlation analysis, employing an algorithm based on mutual information and Granger causality tests to calculate the statistical dependencies and potential causal temporal relationships between pulse features, tongue visual features, and biochemical indicator concentrations, and generates a correlation matrix representing the deep correlation patterns among the three.
[0013] The pathophysiological mechanism mapping and risk assessment module is used to construct and run a cardiovascular risk assessment model with a clear pathophysiological interpretation based on the association matrix. The cardiovascular risk assessment model is a multi-level computational graph network, with its input layer being a multi-dimensional feature vector. The first hidden layer of the model is a feature-mechanism mapping layer, which contains multiple parallel sub-networks. Each sub-network corresponds to a preset cardiovascular pathophysiological mechanism dimension, including vascular endothelial function status, autonomic nervous system regulatory balance, systemic inflammation level, and oxidative stress level. Each sub-network selectively weights and nonlinearly transforms the input multi-dimensional feature vector according to pre-learned weights in the association matrix, outputting a scalar value representing the user's quantitative assessment score in the current mechanism dimension. The second hidden layer of the model is a mechanism-risk integration layer. This risk integration layer receives the assessment scores from all mechanism dimensions and learns the contribution weights of each mechanism to the final cardiovascular risk through a fully connected network, integrating multiple mechanism scores into a comprehensive cardiovascular risk index. The model's output layer maps the cardiovascular risk index to discrete risk levels, with at least three levels: low risk, medium risk, and high risk.
[0014] The training process for cardiovascular risk assessment employs a supervised learning method with pathophysiological constraints. The training data includes multidimensional feature vector samples and corresponding cardiovascular event risk labels confirmed by the gold standard method. At the same time, a regularization term is added to the loss function to encourage the feature-mechanism mapping relationship learned by the model to be consistent with known medical knowledge.
[0015] The dynamic monitoring and early warning module is used to perform continuous or periodic repetitive assessments of users and generate early warnings based on time-series data. This module drives the system to automatically execute the entire process from data collection to risk assessment according to a preset monitoring cycle. The system records the multidimensional feature vectors, scores for each pathophysiological mechanism dimension, cardiovascular risk index, and risk level obtained from each user's assessment, forming the user's longitudinal health time-series profile. The dynamic monitoring and early warning module incorporates a time-series prediction model based on long short-term memory networks. This model uses the user's historical risk index sequence and key feature change sequence as input to predict the evolution trend of the cardiovascular risk index within a specific future time window. When the predicted risk index exceeds the user's personal baseline value by a certain percentage, or when the currently assessed risk level jumps, the dynamic monitoring and early warning module generates and sends a graded early warning signal. The early warning signal is transmitted to a designated client or medical monitoring platform through the system interface.
[0016] As one embodiment of the present invention, the adsorption process of the microfluidic sensor chip array in the microenvironment biochemical index non-invasive sensing module is controlled by an active negative pressure micropump to ensure that a sufficient amount of saliva sample is obtained for detection within the set collection time window, while avoiding discomfort to the user's tongue.
[0017] As one embodiment of the present invention, the deep convolutional neural network used for tongue segmentation in the multimodal feature deep correlation analysis module adopts an encoder-decoder structure. The encoder part uses a residual network pre-trained on a large medical image dataset to extract features, and the decoder part recovers spatial details through progressive upsampling and skip connections, and finally outputs a probability map of each pixel belonging to the tongue region.
[0018] In one embodiment of the present invention, each sub-network of the feature-mechanism mapping layer in the pathophysiological mechanism mapping and risk assessment module is provided with a feature importance interpretation unit. During the forward propagation of the model, the feature importance interpretation unit calculates the gradient of each input feature with respect to the output score of the sub-network, and normalizes the absolute value of the gradient as an interpretability quantification index of the feature's contribution to the current pathophysiological mechanism dimension. This interpretability quantification index is output along with the evaluation result.
[0019] As one embodiment of the present invention, the time series prediction model based on a long short-term memory network in the dynamic monitoring and early warning module employs a multi-task learning framework during its training process. The primary task is to predict future risk indices, while the auxiliary task is to predict future changes in the concentrations of key biochemical indicators. By sharing the underlying parameters of the network, the model's accuracy and physiological rationality in predicting risk trends are enhanced by utilizing the physiological significance of changes in biochemical indicators.
[0020] In one embodiment of the present invention, the overall system architecture adopts an edge-cloud collaborative computing model. The multimodal data synchronous acquisition module and the non-invasive sensing module for microenvironmental biochemical indicators are deployed on edge computing devices, responsible for the acquisition, preliminary preprocessing, and encryption of raw data. The core computing tasks of the multimodal feature deep correlation analysis module, the pathophysiological mechanism mapping and risk assessment module, and the dynamic monitoring and early warning module are deployed on a cloud server. Edge devices upload encrypted feature data or preprocessing results to the cloud for computation via a secure communication protocol and receive the returned assessment results and early warning information.
[0021] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0022] 1. This invention, by integrating a non-invasive sensing module for microenvironmental biochemical indicators, achieves for the first time in-situ, simultaneous, and non-invasive detection of specific cardiovascular risk biochemical markers in saliva on the tongue surface during tongue image acquisition. This technological breakthrough deepens traditional Chinese medicine tongue diagnosis from a single visual morphological analysis to a comprehensive assessment including molecular-level biochemical information, providing solid modern biological evidence for cardiovascular risk assessment and significantly enhancing the objectivity and scientific depth of the assessment results.
[0023] 2. This invention constructs a multimodal feature deep correlation analysis module and an assessment model with clear pathophysiological interpretation. It deeply correlates visual features of pulse and tongue appearance with biochemical indicators, mapping them to specific pathophysiological mechanisms such as vascular endothelial function and autonomic nervous system regulation. This design goes beyond simple data fusion, establishing an interpretable correlation model between macroscopic signs and microscopic mechanisms. This ensures that the final risk assessment result is not only numerical or graded, but also includes a quantitative interpretation of potential physiological and pathological states, greatly enhancing the system's clinical interpretability and decision support value.
[0024] 3. This invention upgrades the system from a single, static assessment tool to a dynamic monitoring system that continuously tracks cardiovascular health status by designing a dynamic monitoring and early warning module. This system can accumulate longitudinal user data and proactively assess risk evolution trends using time series prediction models, achieving a leap from passive assessment to proactive early warning. This helps in the early detection of signs of risk deterioration, providing continuous and dynamic data support for personalized health management and the selection of clinical intervention timing, aligning with the long-term and continuous needs of chronic disease risk management. Attached Figure Description
[0025] Figure 1 This is a schematic diagram of the overall technical solution architecture of the present invention;
[0026] Figure 2 This is a schematic diagram of the core principle framework of multimodal feature deep correlation analysis in this invention;
[0027] Figure 3 This is a schematic diagram of the hierarchical computational framework of the pathophysiological mechanism mapping and risk assessment model in this invention;
[0028] Figure 4 This is a logical flow and data flow framework diagram of the dynamic monitoring and early warning module in this invention;
[0029] Figure 5 This is a schematic diagram of system interaction and data flow in the edge-cloud collaborative computing mode of this invention. Detailed Implementation
[0030] Example 1: The overall technical architecture of the TCM cardiovascular risk assessment system based on pulse and tongue diagnosis proposed in this invention is shown in the attached figure. Figure 1 As shown, the system consists of five functional units: a multimodal data synchronous acquisition module, a non-invasive sensing module for microenvironment biochemical indicators, a multimodal feature deep correlation analysis module, a pathophysiological mechanism mapping and risk assessment module, and a dynamic monitoring and early warning module. The units are seamlessly connected through a strict time synchronization mechanism and standardized data interfaces, forming a closed-loop, interpretable, and predictable TCM cardiovascular health assessment system.
[0031] First, the multimodal data synchronous acquisition module, as the system's front-end sensing unit, is responsible for synchronously acquiring the user's raw pulse signal and raw tongue image under a unified time reference. This module integrates a high-precision pressure sensor array and a multispectral imaging unit. The high-precision pressure sensor array is configured with a three-channel structure, corresponding to the three pulse positions (cun, guan, chi) in Traditional Chinese Medicine. Each channel contains no fewer than 16 miniature piezoresistive sensing units, continuously acquiring pulse wave pressure signals at a sampling frequency of 200 Hz to ensure complete capture of key morphological details such as the main wave, dicrotic wave, and tidal wave.
[0032] The multispectral imaging unit consists of a programmable multi-band light source and a high-resolution CMOS image sensor. The light source system can precisely switch between visible light bands (450 nm to 650 nm) and at least two near-infrared bands (850 nm and 940 nm). During each tongue image acquisition, different wavelength combinations of illumination sources are sequentially lit, and the image sensor is simultaneously triggered to acquire the corresponding sequence of digital images of the tongue region. The entire acquisition process is uniformly scheduled by the embedded main control chip, ensuring that the timestamp error between the pulse signal and the tongue image does not exceed 1 millisecond, thus providing a precise time alignment basis for subsequent multimodal data fusion.
[0033] The non-invasive microenvironment biochemical index sensing module is tightly integrated into the tongue image acquisition end of the multimodal data synchronous acquisition module. Its core function is to achieve in-situ, non-invasive, and quantitative detection of specific cardiovascular risk-related biochemical indicators in the saliva membrane of the tongue surface without interfering with the visual acquisition of the tongue image. This non-invasive microenvironment biochemical index sensing module includes a microfluidic sensor chip array and a surface plasmon resonance optical detection unit. The microfluidic sensor chip array is made of biocompatible polymer material, and its surface is constructed with multiple independent microchannel networks through micro-nano fabrication technology. Each microchannel has a micron-sized sampling hole at its end, which directly contacts the user's tongue surface. When acquisition is initiated, an active negative pressure micropump is activated, forming a controllable negative pressure gradient within the microchannel. Through capillary action, it actively adsorbs and enriches the naturally secreted saliva sample from the tongue surface. The entire adsorption process lasts for 30 seconds to ensure that sufficient sample is obtained for subsequent testing. At the same time, the negative pressure value is strictly controlled within -5 kPa to avoid any physical stimulation or discomfort to the user's tongue mucosa.
[0034] Each detection area of the microfluidic chip is pre-immobilized with highly specific biological probe molecules. These probe molecules are designed to target three biochemical indicators—C-reactive protein, interleukin-6, and malondialdehyde—and possess picomolar affinity. When a saliva sample flows through the detection area, the target biochemical indicator specifically binds to the corresponding probe, causing a change in the local dielectric constant of the chip. The surface plasmon resonance optical detection unit, consisting of a broadband light source, a prism coupler, a high-sensitivity photodetector array, and signal processing circuitry, monitors the refractive index change on the chip surface in real time caused by molecular binding events and converts this optical signal into an electrochemical output signal linearly related to the concentration of the biochemical indicator. This signal is then converted from analog to digital and output in digital form, in nanograms per milliliter (Ng / mL), with a typical detection range of 0.1 to 100 Ng / mL and a detection limit better than 0.05 Ng / mL.
[0035] The multimodal feature deep correlation analysis module receives the raw data streams from the two modules mentioned above and performs a series of complex signal processing and feature extraction operations. Please refer to the appendix. Figure 2 The processing flow of this multimodal feature deep correlation analysis module is divided into four stages: pulse feature extraction, tongue feature extraction, multimodal feature alignment, and correlation analysis. In the pulse feature extraction stage, the original pulse wave pressure signal is first processed by a baseline drift correction algorithm, which uses a fifth-order polynomial fitting to remove slowly drifting components; then, a 50 Hz notch filter and a bandpass filter (0.5 Hz to 40 Hz) are used to filter out power frequency interference and high-frequency noise.
[0036] The preprocessed signal was used to extract three types of features: temporal features, including the amplitude of the dominant wave (defined as the vertical distance between the pulse wave peak and the baseline, in millimeters of mercury), the amplitude of the dicrotic wave (distance between the dicrotic wave peak and the baseline), the time interval between the dominant wave and the dicrotic wave (in milliseconds), and pulse rate variability indices (including SDNN and RMSSD) calculated based on a continuous 5-minute pulse interval sequence; frequency domain features were obtained by performing a fast Fourier transform on 512 pulse wave segments, focusing on extracting the ratio of low-frequency power (0.04 Hz to 0.15 Hz) to high-frequency power (0.15 Hz to 0.4 Hz) (LF / HF); and nonlinear dynamic features were obtained by calculating the sample entropy (embedding dimension of 2, similarity tolerance of 0.2 standard deviations) and the maximum Lyapunov exponent (using a small data volume method, reconstruction dimension of 5). In the tongue image feature extraction stage, the multispectral image sequence first underwent white balance correction and illumination normalization to eliminate the influence of ambient light.
[0037] Subsequently, a semantic segmentation algorithm based on a deep convolutional neural network (DCNN) was used to accurately extract the tongue region. This DCNN employs an encoder-decoder structure. The encoder uses a ResNet-50 backbone network pre-trained on ImageNet and a large TCM tongue image database to extract high-dimensional semantic features layer by layer. The decoder performs four upsampling operations through transposed convolution and introduces skip connections to fuse spatial detail information from each stage of the encoder. The final output is a probability map of the same size as the input image, where each pixel value represents the confidence level of belonging to the tongue region. After thresholding, a binary tongue mask is obtained. Based on this binary tongue mask, color features are extracted from the visible light image: in the CIELAB color space, the color features of the tongue body region (excluding the area covered by tongue coating) are calculated. Mean and standard deviation, and convert to colorimetric values. The hue angle (hab) is used; the tongue coating area is separated into thick and thin coatings through morphological opening and closing operations, the area ratio of thick coating is calculated as the quantitative value of tongue coating thickness, and the distribution uniformity is evaluated by calculating the gray standard deviation of the tongue coating area.
[0038] Texture features are obtained by calculating the gray-level co-occurrence matrix in the tongue region (direction at 0 degrees, 45 degrees, 90 degrees, and 135 degrees, distance of 1 pixel), with a focus on extracting the mean of contrast and energy values. For morphological features, the tongue's plumpness index is defined as the ratio of the tongue's area to the area of its circumscribed rectangle. Edge dents are identified using Canny edge detection and Hough transform, counting the number of dents with a depth greater than 2 pixels and their average depth. After completing single-modal feature extraction, the module performs multi-modal feature alignment: since the pulse sampling frequency is 200 Hz and the tongue image and biochemical indicators are single snapshots, the system uses the tongue image acquisition time as the time anchor, extracting pulse signal windows for 15 seconds before and after, and calculating the moving average of all pulse features within this window, thereby generating a pulse feature vector strictly aligned with the tongue image / biochemical indicator timestamps. Subsequently, all features (including pulse features, tongue visual features, C-reactive protein concentration, interleukin-6 concentration, and malondialdehyde concentration) were normalized to the 0 to 1 range, forming a multidimensional feature vector of dimension N (N is usually 25 to 35).
[0039] Finally, the association analysis phase employs a dual-algorithm approach to verify the deep relationships between features: first, the mutual information value between any two features is calculated to measure their statistical dependence strength; second, for feature pairs with high mutual information, a Granger causality test is performed (with a lag order of 3) to determine whether a significant causal time-series relationship exists. All calculation results are organized into an N×N association matrix, where the value of matrix element (i,j) represents the causal influence strength of the i-th feature on the j-th feature. This matrix serves as a key input to the subsequent risk assessment model.
[0040] The pathophysiological mechanism mapping and risk assessment module is the core decision engine of this system, and its computational framework is shown in the attached figure. Figure 3 As shown, this pathophysiological mechanism mapping and risk assessment module operates a multi-level computational graph network, aiming to map multi-dimensional feature vectors to cardiovascular pathophysiological mechanism dimensions with clear medical significance, and ultimately output an actionable risk level. The model input layer receives the aforementioned normalized multi-dimensional feature vectors. The first hidden layer is a feature-mechanism mapping layer, containing four parallel sub-networks, corresponding to four preset pathophysiological mechanism dimensions: vascular endothelial functional state, autonomic nervous system regulatory balance, systemic inflammation level, and oxidative stress level.
[0041] Each sub-network is a three-layer fully connected neural network, with its weights determined through supervised learning during model training. Training data is derived from a large-scale clinical cohort, where each sample contains a multi-dimensional feature vector and a corresponding gold standard label—for example, the vascular endothelial function status label is quantified from the results of blood flow-mediated vasodilation function testing. In the loss function, in addition to the standard mean squared error term, a regularization term based on a medical knowledge graph is added, forcing features strongly correlated with known mechanisms (such as the pulse LF / HF ratio and autonomic nervous system regulation) to receive higher weights in their corresponding sub-networks. Each sub-network outputs a scalar value between 0 and 1, representing the user's score on the degree of abnormality in that mechanism dimension (0 for completely normal, 1 for severely abnormal). The second hidden layer is a mechanism-risk integration layer, which receives four mechanism scores as input and learns the combined contribution weights of each mechanism to the final cardiovascular risk through two fully connected layers. The training label for this network is the risk probability of major adverse cardiovascular events (such as myocardial infarction and stroke) occurring within the next 12 months, determined by clinical follow-up data. The integration layer outputs a cardiovascular risk index from 0 to 100. The output layer maps this index to discrete risk levels: 0 to 30 is low risk, 31 to 70 is medium risk, and 71 to 100 is high risk.
[0042] In addition, each sub-network embeds a feature importance interpretation unit. During each forward propagation, it calculates the gradient of the input feature with respect to the sub-network output score, takes the absolute value of the gradient and normalizes it to generate a feature contribution vector. This feature contribution vector is output along with the evaluation result, providing users with an interpretable basis for "why it was judged to a certain risk level".
[0043] The dynamic monitoring and early warning module enables the system to perform long-term health management. Please refer to the appendix. Figure 4 The dynamic monitoring and early warning module automatically triggers a complete assessment process according to a preset monitoring cycle (e.g., weekly or daily). After each assessment, the system stores the multidimensional feature vector, scores of four mechanisms, cardiovascular risk index, and risk level in the user's personalized longitudinal health time series archive. This longitudinal health time series archive is indexed by timestamps, forming a structured database. The module incorporates a time series prediction model based on a long short-term memory network. This model takes the user's most recent K (K=10) historical risk index sequences and key feature (e.g., C-reactive protein concentration, pulse rate variability SDNN) change sequences as input to predict the evolution trajectory of the cardiovascular risk index within the next T days (T=30). The time series prediction model is trained using a multi-task learning framework: the primary task is to predict future risk indices, and the auxiliary task is to predict future changes in C-reactive protein and malondialdehyde concentrations.
[0044] The two tasks share the underlying parameters of the LSTM encoder but have their own independent decoders. This design utilizes the physiological and pathological significance of changes in biochemical indicators to constrain the prediction direction of the main task, making it more consistent with biological laws. The system generates an early warning signal when either of the following conditions is met: 1) the predicted future risk index exceeds the user's personal historical baseline value (defined as the moving average of the risk index over the past 30 days) by 20%; 2) the current assessed risk level jumps upward compared to the previous assessment (e.g., from medium risk to high risk). The early warning signal is divided into two levels: Level 1 warning (risk index increase of 20% to 50% or level jump of one level) is pushed to the mobile application; Level 2 warning (risk index increase of more than 50% or direct entry into high risk) is transmitted in real time to the monitoring platform of the contracted medical institution through a secure communication protocol, along with a complete assessment report and feature contribution analysis.
[0045] The entire system is deployed using an edge-cloud collaborative computing model, as shown in the attached diagram. Figure 5 As shown, the multimodal data synchronous acquisition module and the non-invasive sensing module for microenvironmental biochemical indicators are integrated into a portable edge computing device (such as a dedicated handheld terminal). This portable edge computing device has a built-in ARM Cortex-A72 processor and a security encryption chip, responsible for the acquisition of raw data, preliminary preprocessing (such as pulse filtering and image compression), and AES-256 encryption. The encrypted feature data or lightweight preprocessing results are uploaded to the cloud server via 4G / 5G or Wi-Fi networks. The cloud server cluster deploys core computing services for the multimodal feature deep correlation analysis module, the pathophysiological mechanism mapping and risk assessment module, and the dynamic monitoring and early warning module, using containerization technology to achieve high-concurrency processing. The assessment results and early warning information are returned to the edge device through the same encrypted channel and presented to the user after local decryption. This architecture not only ensures the privacy and security of the user's biological data but also makes full use of the powerful computing resources of the cloud, achieving a balance between performance and security.
[0046] Example 2: Based on Example 1, this example optimizes the microfluidic chip structure of the non-invasive sensing module for microenvironmental biochemical indicators to improve the collection efficiency and detection stability of saliva samples. Specifically, the microchannel network of the microfluidic sensing chip array adopts a fractal tree topology design, with a main channel diameter of 200 micrometers, branching stepwise to the end sampling aperture with a diameter of 50 micrometers. This fractal tree topology design mimics the human capillary network, enabling saliva to be evenly distributed along the fractal path to all detection areas under the drive of an active negative pressure micropump, avoiding detection failure caused by local blockage in traditional parallel channel designs.
[0047] Meanwhile, in addition to fixing the biological probe, each detection area is also coated with an antifouling coating (such as a polyethylene glycol derivative) to effectively suppress non-specific protein adsorption and reduce background noise to less than 5% of the original signal. The surface plasmon resonance optical detection unit uses a wavelength-tunable laser as its light source, which can scan in the range of 780 nm to 850 nm. By finding the wavelength corresponding to the minimum resonance angle, it achieves higher precision measurement of refractive index changes and controls the relative error of biochemical indicator concentration detection to within 3%.
[0048] Furthermore, this embodiment improves the tongue segmentation network in the multimodal feature deep association analysis module. The encoder no longer uses the general-purpose ResNet-50, but instead employs a lightweight attention enhancement network specifically designed for tongue images. This lightweight attention enhancement network embeds channel attention modules and spatial attention modules into the residual blocks. The former calculates the importance weights of each channel through global average pooling, while the latter generates a spatial position weight map through convolutional layers, thus focusing on key tongue regions (such as the tip and sides) in the early stages of the network and suppressing interference from teeth and oral cavity background. The decoder introduces a dilated convolutional pyramid to capture multi-scale contextual information during upsampling, effectively solving the edge blurring problem caused by changes in tongue pose. Testing showed that this improved network achieved a segmentation intersection-union ratio (IUU) of 92.5% on a private dataset containing 10,000 labeled tongue images, an improvement of 2.3 percentage points compared to the original scheme.
[0049] In the pathophysiological mechanism mapping and risk assessment module, this embodiment expands the number of sub-networks in the feature-mechanism mapping layer. In addition to the original four mechanism dimensions, two new dimensions, "lipid metabolism disorder" and "hypercoagulability," are added. The former is assessed by integrating features such as the thickness of the tongue coating, malondialdehyde concentration, and the amplitude of the main pulse wave; the latter is correlated with the purplish-darkness of the tongue, C-reactive protein concentration, and the absence of the dicrotic pulse wave. Each new sub-network is also equipped with a feature importance interpretation unit, and corresponding gold standard labels (such as low-density lipoprotein cholesterol levels and D-dimer concentration) are introduced during training. The mechanism-risk integration layer is correspondingly adjusted to a six-input fully connected network, enabling cardiovascular risk assessment to cover a more comprehensive pathophysiological spectrum.
[0050] The dynamic monitoring and early warning module has also been upgraded. The input to the time series prediction model not only includes the user's own historical data but also incorporates group health benchmark data as a contextual reference. Specifically, the cloud server maintains a database of group risk index distributions stratified by age, gender, and underlying diseases. During prediction, the model considers both individual user trends and the average evolution pattern of the same group, dynamically adjusting their weights through a gating mechanism. When user data is sparse (e.g., for new users), it relies more on group patterns; when user data is abundant, it prioritizes individual trends. This design significantly improves the reliability of early warnings for new users. An additional warning trigger condition has been added: when a user's current risk index exceeds the 90th percentile of the same group, a Level 1 warning is triggered, indicating a potential high risk, even without a tier jump.
[0051] The system as a whole still adopts an edge-cloud collaborative architecture, but edge devices have added local caching and breakpoint resume functionality. During network outages, devices can cache up to 7 days of raw data and automatically resume transmission after the network is restored, ensuring the integrity of longitudinal data. The cloud service introduces a federated learning mechanism, which, while protecting the data privacy of various medical institutions, combines data from multiple parties to continuously optimize the risk assessment model, achieving dynamic evolution and improved generalization capabilities of the model.
[0052] The two embodiments above together constitute the complete technical solution of the present invention. Embodiment 1 provides the basic implementation framework of the system, while Embodiment 2 demonstrates in-depth optimization and functional expansion of key modules. Both strictly follow the core concept of the present invention, namely, to construct a dynamic assessment and early warning system for cardiovascular risk with pathophysiological interpretability through multimodal deep fusion of pulse, tongue appearance, and biochemical indicators of the tongue microenvironment.
Claims
1. A Traditional Chinese Medicine cardiovascular risk assessment system based on pulse and tongue diagnosis, characterized in that, include: The multimodal data synchronous acquisition module is used to synchronously acquire the user's original pulse signal and original tongue image under a unified time reference. A non-invasive sensing module for microenvironment biochemical indicators is integrated into the tongue image acquisition end of the multimodal data synchronous acquisition module. It is used to non-invasively and in situ detect the concentration of specific biochemical indicators related to cardiovascular risk in the salivary membrane of the tongue. The multimodal feature deep correlation analysis module is used to receive and process raw data from the multimodal data synchronous acquisition module and the microenvironment biochemical index non-invasive sensing module, and extract and correlate multi-dimensional features. The pathophysiological mechanism mapping and risk assessment module is used to construct and run a cardiovascular risk assessment model with a clear pathophysiological explanation based on the correlation matrix. The dynamic monitoring and early warning module is used to conduct continuous or periodic repeated assessments of users and generate early warnings based on time series data.
2. The TCM cardiovascular risk assessment system based on pulse and tongue diagnosis according to claim 1, characterized in that, The multimodal data synchronous acquisition module integrates a high-precision pressure sensor array and a multispectral imaging unit; The high-precision pressure sensor array collects pulse wave pressure signals from the Cun, Guan, and Chi points at a specific sampling frequency. The multispectral imaging unit acquires digital image sequences of the tongue region under illumination conditions with specific wavelength combinations; The digital image sequence contains information in the visible light band and at least two near-infrared bands; The non-invasive sensing module for microenvironment biochemical indicators includes a microfluidic sensing chip array and a surface plasmon resonance optical detection unit. The microfluidic sensor chip array actively adsorbs and enriches saliva from the tongue surface through capillary action, and its surface is immobilized with biological probe molecules that specifically bind to target biochemical indicators. The surface plasmon resonance optical detection unit monitors in real time the change in refractive index of the chip surface caused by the binding of biological probe molecules with the target biochemical index, and converts the change in refractive index of the chip surface into the electrochemical concentration signal of the corresponding biochemical index. The target biochemical indicators include at least C-reactive protein, interleukin-6, and malondialdehyde.
3. The TCM cardiovascular risk assessment system based on pulse and tongue diagnosis according to claim 1, characterized in that, The multimodal feature deep correlation analysis module first preprocesses the pulse wave pressure signal, including baseline drift correction and power frequency interference filtering, and then extracts time domain features, frequency domain features and nonlinear dynamic features. The time-domain features include the amplitude of the main wave, the amplitude of the diabetic wave, the time interval between the main wave and the diabetic wave, and the pulse rate variability index. The frequency domain features are calculated using fast Fourier transform and include the ratio of low-frequency power to high-frequency power. The nonlinear dynamic features include sample entropy and Lyapunov exponent. The multimodal feature deep correlation analysis module segments and extracts features from the original tongue image. The segmentation process uses a semantic segmentation algorithm based on deep convolutional neural networks to accurately separate the tongue region from the background and teeth region, and extracts color features, texture features and morphological features from the segmented tongue image. The color features include the average chromaticity and saturation values of the tongue body in a specific color space, as well as the quantified values of the thickness and uniformity of the tongue coating. The texture features are calculated using a gray-level co-occurrence matrix, including contrast and energy values. The morphological features include the tongue's plumpness index and the depth and number of edge teeth marks. The multimodal feature deep correlation analysis module performs time alignment and normalization processing on the extracted pulse feature set, tongue visual feature set, and microenvironment biochemical index concentration set to form a multidimensional feature vector. The multimodal feature deep correlation analysis module performs correlation analysis, using an algorithm based on mutual information and Granger causality test to calculate the statistical dependence and potential causal temporal relationship between pulse features, tongue visual features and biochemical index concentrations, and generates a correlation matrix characterizing the deep correlation pattern among the three.
4. The TCM cardiovascular risk assessment system based on pulse and tongue diagnosis according to claim 1, characterized in that, The model is a multi-level computational graph network, and its input layer is the multi-dimensional feature vector; The first hidden layer of the model is a feature-mechanism mapping layer, which contains multiple parallel sub-networks. Each sub-network corresponds to a preset cardiovascular pathophysiological mechanism dimension, including vascular endothelial function status, autonomic nervous regulation balance, systemic inflammation level, and oxidative stress level. Each sub-network selectively weights and nonlinearly transforms the input multidimensional feature vector according to the weights learned in the correlation matrix, and outputs a scalar value representing the user's quantitative evaluation score in the current mechanism dimension. The second hidden layer of the model is a mechanism-risk integration layer. This risk integration layer receives the evaluation scores of all mechanism dimensions and learns the contribution weight of each mechanism to the final cardiovascular risk through a fully connected network, integrating the scores of multiple mechanisms into a comprehensive cardiovascular risk index. The output layer of the model maps the cardiovascular risk index to discrete risk levels, which are at least divided into three levels: low risk, medium risk, and high risk.
5. The TCM cardiovascular risk assessment system based on pulse and tongue diagnosis according to claim 1, characterized in that, The dynamic monitoring and early warning module drives the system to automatically execute the entire process from data collection to risk assessment according to the preset monitoring cycle, and records the multidimensional feature vectors, scores of each pathophysiological mechanism dimension, cardiovascular risk index and risk level obtained by the user in each assessment, forming the user's longitudinal health time series profile. The dynamic monitoring and early warning module has a built-in time series prediction model based on long short-term memory network. This time series prediction model takes the user's historical risk index sequence and key feature change sequence as input to predict the evolution trend of cardiovascular risk index within a specific time window in the future. When the predicted risk index exceeds the user's personal baseline value by a certain percentage, or when the currently assessed risk level jumps. The dynamic monitoring and early warning module generates and sends tiered early warning signals.
6. The TCM cardiovascular risk assessment system based on pulse and tongue diagnosis according to claim 1, characterized in that, The adsorption process of the microfluidic sensor chip array in the microenvironment biochemical index non-invasive sensing module is controlled by an active negative pressure micropump to ensure that a sufficient amount of saliva sample is obtained for detection within the set collection time window. The deep convolutional neural network used for tongue segmentation in the multimodal feature deep correlation analysis module adopts an encoder-decoder structure. The encoder part uses a residual network pre-trained on a large medical image dataset to extract features, and the decoder part recovers spatial details through progressive upsampling and skip connections, and finally outputs a probability map of each pixel belonging to the tongue region.
7. The TCM cardiovascular risk assessment system based on pulse and tongue diagnosis according to claim 1, characterized in that, Within each sub-network of the feature-mechanism mapping layer in the pathophysiological mechanism mapping and risk assessment module, a feature importance interpretation unit is set up; During the forward propagation of the model, the feature importance interpretation unit calculates the gradient of each input feature with respect to the output score of the sub-network, and normalizes the absolute value of the gradient as a quantitative indicator of the interpretability of the feature's contribution to the current pathophysiological mechanism dimension. The time series prediction model based on long short-term memory network in the dynamic monitoring and early warning module adopts a multi-task learning framework in its training process. The main task is to predict the future risk index, and the auxiliary task is to predict the changes in the concentration of key biochemical indicators in the future. By sharing the underlying parameters of the network, the model enhances the accuracy and physiological rationality of risk trend prediction by utilizing the physiological significance of changes in biochemical indicators.
8. The TCM cardiovascular risk assessment system based on pulse and tongue diagnosis according to claim 1, characterized in that, The overall system architecture adopts an edge-cloud collaborative computing model. The multimodal data synchronous acquisition module and the microenvironment biochemical index non-invasive sensing module are deployed on the edge computing device and are responsible for the acquisition, preliminary preprocessing and encryption of raw data. The core computing tasks of the multimodal feature deep correlation analysis module, the pathophysiological mechanism mapping and risk assessment module, and the dynamic monitoring and early warning module are deployed on a cloud server. Edge devices upload encrypted feature data or preprocessing results to the cloud for calculation through a secure communication protocol, and receive the returned evaluation results and early warning information.
9. The TCM cardiovascular risk assessment system based on pulse and tongue diagnosis according to claim 2, characterized in that, The negative pressure value generated by the active negative pressure micro-pump is controlled within -5 kPa.
10. The TCM cardiovascular risk assessment system based on pulse and tongue diagnosis according to claim 3, characterized in that, In the deep convolutional neural network used for tongue segmentation, the residual network used in the encoder part is ResNet-50.