Periodontal disease risk prediction system based on cardiovascular and oral microbiota data fusion

By integrating cardiovascular and oral microbiota data into a multimodal prediction system, the challenges of data heterogeneity and model deployment in periodontal disease risk prediction have been solved. This system enables accurate, rapid, and non-invasive periodontal disease risk prediction and is suitable for clinical, home, and large-scale population screening.

CN122117408APending Publication Date: 2026-05-29BAODING SECOND CENT HOSPITAL

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BAODING SECOND CENT HOSPITAL
Filing Date
2026-03-19
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing periodontal disease risk prediction technologies suffer from problems such as single data sources, poor heterogeneous data fusion, inability to deploy prediction models in a lightweight manner, inaccurate risk classification, and cumbersome testing procedures, failing to meet the needs for accurate, rapid, and non-invasive screening in clinical and home settings.

Method used

A periodontal disease risk prediction system based on the fusion of cardiovascular and oral microbiota data is adopted. Through multimodal data acquisition, heterogeneous data preprocessing, cross-modal feature fusion, lightweight bi-branch model prediction and dynamic risk classification, it can achieve accurate, rapid and non-invasive prediction of periodontal disease risk.

Benefits of technology

It breaks through the limitations of a single data source, significantly improves the accuracy of periodontal disease risk prediction, realizes the deployment of a lightweight model, adapts dynamic threshold calibration to individual characteristics, simplifies the detection process, and improves ease of operation and prediction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122117408A_ABST
    Figure CN122117408A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of biological detection and disease risk prediction, in particular to a periodontal disease risk prediction system based on cardiovascular and oral flora data fusion, comprising a multi-modal data acquisition module, a heterogeneous data preprocessing module, a cross-modal feature fusion module, a double-branch risk prediction model module and a risk grading output module. The multi-modal data acquisition module non-invasively acquires cardiovascular physiological data and oral mucosa swab flora sequencing data of the subject; the preprocessing module denoises and normalizes the time series physiological data, and completes species annotation and abundance correction of the flora data; the cross-modal feature fusion module realizes adaptive alignment and deep fusion of the two types of heterogeneous data through a cross-modal attention mechanism; the double-branch prediction model outputs a risk probability after being trained by a comorbidity correlation data set; and the risk grading module divides the individual risk level in combination with a dynamic threshold. The whole-body cardiovascular and oral flora data are deeply fused to improve the prediction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biological detection and disease risk prediction technology, specifically a periodontal disease risk prediction system based on the fusion of cardiovascular and oral microbiota data. Background Technology

[0002] Periodontal disease is a chronic infectious disease caused by an imbalance in the oral microbiota. It is the leading cause of tooth loss in adults. Furthermore, periodontal disease has a clear bidirectional link with cardiovascular disease. Periodontal pathogens can trigger systemic inflammatory responses through the bloodstream, exacerbating cardiovascular diseases such as atherosclerosis, hypertension, and dyslipidemia. Conversely, the systemic inflammatory state and altered immune function in patients with cardiovascular disease can further disrupt the oral microecological balance, accelerating the progression of periodontal disease. Clinical studies have confirmed that early screening and intervention for periodontal disease can significantly reduce the risk of tooth loss and improve the prognosis of patients with cardiovascular disease. Therefore, developing accurate and non-invasive periodontal disease risk prediction tools has significant clinical value and public health implications.

[0003] Existing periodontal disease risk prediction technologies are mainly divided into two categories. One category is the traditional assessment method based on oral clinical examination indicators, which relies on clinical parameters such as periodontal probing depth, attachment loss, and gingival bleeding index. This requires professional physician operation, cannot achieve early screening at home, and the test results are greatly affected by the physician's subjective operation. The other category is the molecular detection method based on oral flora or salivary biomarkers, which only focuses on the local oral microecology or biomolecular characteristics and ignores the regulatory role of the systemic physiological state on the occurrence and development of periodontal disease. The generalization ability and accuracy of the prediction model are limited.

[0004] In existing technologies, some disease prediction systems attempt to integrate multi-source data, but no periodontal disease risk prediction scheme has yet emerged that deeply integrates cardiovascular physiological data and oral microbiota data. Existing multi-source data fusion schemes often employ simple feature stitching, failing to address the heterogeneity between cardiovascular time-series physiological data and oral microbiota physiological data, resulting in low feature utilization efficiency. Furthermore, existing prediction models are mostly large-scale deep learning models with a large number of parameters, making them unsuitable for deployment on portable clinical devices and home terminals, and thus unable to meet real-time screening needs. In addition, existing risk prediction systems often use fixed thresholds for risk grading, without dynamically calibrating based on individual baseline characteristics such as age, gender, and lifestyle habits, leading to significant deviations in prediction results across different populations.

[0005] Furthermore, existing oral microbiota testing and cardiovascular physiology testing follow different testing procedures, resulting in incompatible data acquisition equipment and inconsistent data formats. Data must be manually processed before analysis, making the process cumbersome, time-consuming, and unable to achieve rapid, integrated testing. In summary, existing periodontal disease risk prediction technologies suffer from drawbacks such as a single data source, low data fusion efficiency, difficulties in model deployment, inaccurate risk grading, and cumbersome testing procedures, failing to meet the needs for accurate, rapid, and non-invasive periodontal disease risk screening in both clinical and home settings. Summary of the Invention

[0006] The purpose of this invention is to solve the problems of single data source, poor heterogeneous data fusion effect, inability to deploy lightweight prediction models, and low accuracy of risk classification in existing periodontal disease risk prediction technologies. It provides a periodontal disease risk prediction system based on the fusion of cardiovascular and oral microbiota data. By non-invasively collecting cardiovascular physiological data and oral microbiota data, and through heterogeneous data preprocessing, cross-modal deep feature fusion, lightweight bi-branch model prediction, and dynamic risk classification, it achieves accurate, rapid, and non-invasive prediction of periodontal disease risk.

[0007] The technical solution adopted by this invention to solve its technical problem is: a periodontal disease risk prediction system based on the fusion of cardiovascular and oral microbiota data, comprising a multimodal data acquisition module, a heterogeneous data preprocessing module, a cross-modal feature fusion module, a bi-branch risk prediction model module, and a risk grading output module connected in sequence; the multimodal data acquisition module is used to non-invasively acquire the subject's cardiovascular physiological data and oral microbiota sequencing data. The cardiovascular physiological data includes resting blood pressure, heart rate variability characteristics, serum lipid indicators, and electrocardiogram time-domain and frequency-domain characteristics. The oral microbiota sequencing data is 16S smears from oral mucosal swabs. The rRNA high-throughput sequencing data; the heterogeneous data preprocessing module is used to denoise, fill missing values ​​and normalize cardiovascular time-series physiological data, and to perform sequence quality control, species annotation, abundance standardization and diversity index calculation on oral microbiota data; the cross-modal feature fusion module uses a cross-modal attention mechanism to adaptively align and deeply fuse the preprocessed cardiovascular feature vector and oral microbiota feature vector to generate a fused feature vector; the dual-branch risk prediction model module includes a cardiovascular feature branch, an oral microbiota feature branch and a fusion discrimination branch. After the model is trained on a periodontal disease-cardiovascular comorbidity labeled dataset, it outputs the periodontal disease risk probability based on the fused feature vector; the risk grading output module has a built-in dynamic threshold calibration unit, which dynamically adjusts the risk threshold in combination with the subject's age, gender and oral hygiene habits covariates, and converts the risk probability into three levels of periodontal disease risk results (low, medium and high) and outputs them.

[0008] Specifically, the multimodal data acquisition module is compatible with home-use portable physiological testing devices and clinical microbiology testing devices. It automatically collects and transmits data through a standardized interface without the need for manual format conversion.

[0009] Specifically, the heterogeneous data preprocessing module processes the microbial community data by removing chimeric sequences, dividing into operable taxa, annotating to the genus level, and calculating the Shannon diversity index and the Simpson diversity index.

[0010] Specifically, the cross-modal attention mechanism is a structure that combines self-attention and cross-attention, automatically weighting key dimensions related to periodontal disease risk in cardiovascular features and oral microbiota features.

[0011] Specifically, the dual-branch risk prediction model is a lightweight neural network. The cardiovascular branch uses a temporal convolutional network to process physiological time-series data, the microbiome branch uses a multilayer perceptron to process omics data, and the fusion discrimination branch uses a fully connected layer to output risk probabilities.

[0012] Specifically, the dynamic threshold calibration unit takes the subject's baseline characteristics as input and generates an appropriate risk grading threshold through a linear regression model. The threshold is updated in real time with the subject's baseline characteristics.

[0013] Specifically, the risk grading output module is also associated with the intervention suggestion unit, which outputs corresponding oral care and cardiovascular health management suggestions according to different risk levels.

[0014] Specifically, the lightweight neural network has a model parameter compression rate of no less than 70% and can be deployed on edge computing terminals and mobile terminals.

[0015] Specifically, the periodontal disease-cardiovascular comorbidity labeled dataset contains paired cardiovascular data and oral microbiota data of clinically diagnosed periodontal disease patients and healthy subjects, and the labeled information includes the severity grading of periodontitis.

[0016] The beneficial effects of this invention are: 1. By integrating systemic cardiovascular physiological data with local oral microbiota data, it overcomes the limitations of prediction based on a single data source and significantly improves the accuracy of periodontal disease risk prediction by combining systemic and local characteristics; 2. Employing a cross-modal attention mechanism enables deep fusion of heterogeneous data, resulting in more accurate key feature identification and higher feature utilization efficiency; 3. The lightweight dual-branch model can be deployed on multiple types of terminals, enabling rapid prediction in both clinical and home scenarios; 4. Dynamic threshold calibration adapts to different individual characteristics, resulting in more accurate risk classification; 5. Non-invasive integrated data collection, simple operation, high subject compliance, suitable for large-scale population screening. Attached Figure Description

[0017] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0018] Figure 1 This is an architecture diagram of the periodontal disease risk prediction system based on the fusion of cardiovascular and oral microbiota data provided by the present invention. Detailed Implementation

[0019] To make the technical means, creative features, objectives and effects of this invention easier to understand, the invention will be further described below in conjunction with specific embodiments.

[0020] like Figure 1 As shown, the periodontal disease risk prediction system based on the fusion of cardiovascular and oral microbiota data described in this invention includes a multimodal data acquisition module, a heterogeneous data preprocessing module, a cross-modal feature fusion module, a dual-branch risk prediction model module, and a risk grading output module. The modules work together to complete the fully automated processing from data acquisition to risk result output.

[0021] The multimodal data acquisition module employs a non-invasive testing method, compatible with both home-use portable physiological testing devices and clinical microbiology testing devices. It automatically acquires cardiovascular physiological data and oral microbiota sequencing data from subjects without manual intervention. The data acquisition process is simple, non-invasive, and results in high subject compliance. The cardiovascular physiological data covers core indicators such as resting systolic blood pressure, diastolic blood pressure, time and frequency domain characteristics of heart rate variability, total cholesterol, triglycerides, high-density lipoprotein, low-density lipoprotein, PR interval, QRS duration, and ST segment deviation on electrocardiogram, comprehensively reflecting the subject's overall cardiovascular physiological status. The oral microbiota sequencing data consists of high-throughput 16S rRNA sequencing data from oral mucosal swab samples. Sample collection is non-invasive, enabling rapid acquisition of microbial genome data.

[0022] The heterogeneous data preprocessing module is designed with a dedicated processing flow to address the heterogeneous characteristics of the two types of data. For cardiovascular time-series physiological data, wavelet denoising, K-nearest neighbor missing value imputation, and min-max normalization are performed to eliminate noise and dimensional differences. For oral microbiota data, sequence quality control, chimera removal, operable taxonomic unit division, species annotation to the genus level are performed, microbiota abundance and diversity indices are calculated, and low-abundance noisy species are removed to ensure the validity and comparability of microbiota data.

[0023] The cross-modal feature fusion module adopts a cross-modal attention mechanism that combines self-attention and cross-attention. It automatically identifies and weights the feature dimensions that are highly correlated with the risk of periodontal disease in the two types of data, and realizes adaptive alignment and deep fusion of cardiovascular feature vectors and oral microbiota feature vectors to generate high-dimensional fused feature vectors. This solves the problems of information redundancy and loss of key features caused by traditional feature splicing, and greatly improves the efficiency of feature utilization.

[0024] The dual-branch risk prediction model module features a lightweight neural network structure, divided into a cardiovascular feature branch, an oral microbiota feature branch, and a fusion discriminant branch. The cardiovascular branch uses a temporal convolutional network to process temporal physiological data, adapting to the temporal characteristics of cardiovascular data. The microbiota branch uses a multilayer perceptron to process microbiota biometric data, adapting to the high-dimensional and sparse characteristics of microbiota data. The fusion discriminant branch integrates the output features of the two branches to output the probability of periodontal disease risk. The model is trained on a large-scale periodontal disease-cardiovascular comorbidity labeled dataset and achieves lightweight design through parameter compression, reducing the number of parameters by more than 70%. It can be deployed on edge computing terminals, mobile devices, and clinical testing equipment to meet real-time prediction needs.

[0025] The risk grading output module incorporates a dynamic threshold calibration unit and an intervention suggestion unit. The dynamic threshold calibration unit takes covariates such as the subject's age, gender, oral hygiene habits, and smoking history as inputs, and generates personalized risk thresholds through a pre-trained linear regression model, replacing traditional fixed thresholds and improving the accuracy of grading different populations. The intervention suggestion unit automatically matches corresponding oral care plans, periodontal intervention measures, and cardiovascular health management suggestions based on low, medium, and high risk levels, achieving integrated prediction and intervention.

[0026] Example 1: This example provides a basic operating scheme for a periodontal disease risk prediction system based on the fusion of cardiovascular and oral microbiota data. The system modules operate according to the following process: First, the multimodal data acquisition module was activated. After the subject sat quietly for 5 minutes, resting systolic and diastolic blood pressure were collected using a portable blood pressure monitor, heart rate variability and electrocardiogram data were collected using an ECG wristband, and serum lipid levels were collected using a dry biochemical analyzer. All of these devices were connected to the system via a standardized Bluetooth interface and automatically transmitted data to the preprocessing module. The entire acquisition process was non-invasive and took approximately 10 minutes. Simultaneously, oral mucosal swab samples were collected from the subject and sent to a high-throughput sequencing platform for 16S rRNA sequencing. The sequencing data was transmitted to the preprocessing module via a network interface. The sequencing process was a standard Illumina sequencing process with a sequence length of 250 bp, requiring no special optimization.

[0027] After receiving the data, the heterogeneous data preprocessing module processes the cardiovascular physiological data: wavelet transform is used to remove power frequency noise and electromyographic noise from the electrocardiogram signal; missing blood lipid indicators are filled using the K-nearest neighbor algorithm with K set to 5; all physiological data are normalized to the 0-1 interval, generating a 32-dimensional cardiovascular feature vector; the oral microbiota data is processed: QIIME2 software is used for sequence quality control, removing sequences with quality scores below 20 and chimeric sequences; the DADA2 algorithm is used to divide the data into operable taxa; sequences are annotated to the genus level; low-abundance species with relative abundance below 0.01% are removed; Shannon diversity index and Simpson diversity index are calculated; and the microbiota abundance data is standardized, generating a 64-dimensional oral microbiota feature vector. The preprocessing process is automated and requires no manual operation, taking approximately 15 minutes.

[0028] The cross-modal feature fusion module receives a 32-dimensional cardiovascular feature vector and a 64-dimensional oral microbiota feature vector. It initiates a cross-modal attention mechanism, first extracting the internal correlation features of the two types of features through a self-attention layer, and then calculating the cross-modal correlation weights of the cardiovascular features and microbiota features through a cross-attention layer. It automatically weights the abundance features of periodontal pathogens such as Rhodopseudomonas erythrosporum, Porphyromonas gingivalis, and Actinobacillus actinomycetii, as well as cardiovascular risk features such as heart rate variability and low-density lipoprotein. The weighted features are then spliced ​​and their dimensions compressed to generate a 48-dimensional fusion feature vector, thus completing the cross-modal deep fusion.

[0029] The dual-branch risk prediction model module loads a pre-trained lightweight neural network model. The cardiovascular branch is a 3-layer temporal convolutional network with a kernel size of 3 and a stride of 1; the microbiome branch is a 4-layer multilayer perceptron with ReLU activation function in the hidden layers; and the fusion discriminant branch is a 2-layer fully connected layer with Sigmoid activation function in the output layer, outputting a periodontal disease risk probability between 0 and 1. The model training set contains paired data from 1200 clinically diagnosed periodontal disease patients and 800 healthy subjects. The annotation information includes chronic periodontitis, aggressive periodontitis, and healthy controls. During training, an adaptive moment estimation optimizer is used with a learning rate of 0.001 and a batch size of 32. After training, the model parameters are compressed by 75%, making it suitable for mobile devices.

[0030] After the risk grading output module obtains the risk probability, the dynamic threshold calibration unit reads covariates such as the subject's age, gender, and number of brushing sessions per day, inputs them into a linear regression model, and generates personalized risk thresholds: 0-0.3 for low risk, 0.31-0.6 for medium risk, and 0.61-1 for high risk. If the subject is a male over 60 years old and brushes his teeth once a day, the thresholds are adjusted to 0-0.25 for low risk, 0.26-0.55 for medium risk, and 0.56-1 for high risk. The system outputs the risk level and corresponding intervention recommendations: daily oral hygiene maintenance is recommended for low risk, oral examinations every 3 months are recommended for medium risk, and immediate clinical periodontal treatment and cardiovascular health monitoring is recommended for high risk.

[0031] In this embodiment, the system takes approximately 30 minutes to complete a single prediction, with a prediction accuracy of 92.3%. This is a significant improvement compared to the 81.5% accuracy of predictions based solely on oral microbiota data and the 76.8% accuracy of predictions based solely on cardiovascular data.

[0032] Compared with Example 1-1, using the same subject data as in Example 1, only oral microbiota data was used as the sole input. The cardiovascular data acquisition and processing module was removed, and a traditional multilayer perceptron model was used for prediction without cross-modal fusion. The risk grading used a fixed threshold (0-0.3 for low risk, 0.31-0.6 for medium risk, and 0.61-1 for high risk). The prediction accuracy was 81.5%, and the missed diagnosis rate for high-risk patients reached 18.2%.

[0033] Compared with Examples 1-2, the same subject data as in Example 1 were used, with only cardiovascular physiological data as the sole input. The oral microbiota data collection and processing module was removed, and a traditional logistic regression model was used for prediction. The prediction accuracy was 76.8%, and the misdiagnosis rate of healthy subjects was 22.5%.

[0034] Compared with Examples 1-3, using the same subject data as in Example 1, cardiovascular features and microbiome features were simply concatenated. Without using cross-modal attention mechanism, traditional fully connected neural network prediction was used. The prediction accuracy was 85.7%, the feature redundancy was high, and the model running time increased by 40%.

[0035] Example 2: This example provides a system operation plan for periodontal disease risk prediction in clinical outpatient settings. It optimizes data acquisition and model accuracy based on Example 1. The specific process is as follows: The multimodal data acquisition module is compatible with clinical testing equipment, connecting to fully automated electrocardiogram analyzers, fully automated biochemical analyzers, and high-throughput microbial sequencers. It collects 24-hour ambulatory blood pressure, ambulatory electrocardiogram, fasting lipid profile, and high-depth oral microbiota 16S rRNA sequencing data from subjects. The data acquisition accuracy is higher than that of home devices, the cardiovascular physiological data dimensions are expanded to 48 dimensions, the sequencing depth of oral microbiota data is increased to 50,000 sequences / sample, and species annotation is done at the species level, further improving data accuracy.

[0036] The heterogeneous data preprocessing module optimizes the processing flow, employing sliding window denoising for dynamic cardiovascular data and multiple imputation to fill missing values, ensuring the continuity of time-series data. For oral microbiota data, a functional annotation process is added, using PICRUSt2 software to predict microbiota metabolic functions, screening for lipopolysaccharide synthesis and amino acid metabolism functional modules related to periodontal inflammation, and combining microbiota functional features with abundance features to generate an 80-dimensional oral microbiota feature vector, increasing the feature information of the microbiota functional dimension.

[0037] The cross-modal feature fusion module upgrades the cross-modal attention mechanism by introducing a multi-head attention structure with 4 attention heads. Features are extracted from four dimensions: bacterial abundance, bacterial function, cardiovascular time series, and cardiovascular biochemistry. Multi-scale cross-modal fusion is then performed to generate a 64-dimensional fusion feature vector, further exploring the intrinsic relationship between the two types of data and improving feature representation capabilities.

[0038] The dual-branch risk prediction model module optimizes the lightweight network structure. The cardiovascular branch uses gated recurrent units instead of temporal convolutional networks to adapt to long-term dynamic cardiovascular data. The microbiome branch uses batch normalization layers to optimize the multilayer perceptron, accelerating model convergence. The model training set is expanded to include severe cases such as invasive periodontitis and periodontal abscess, increasing the total sample size to 2,500 cases. The model training accuracy is improved to 95.1%, while the parameter compression rate remains at 72%. It can be deployed on clinical outpatient computer terminals and portable testing devices.

[0039] The risk grading output module optimizes the dynamic threshold calibration unit by adding covariates such as history of diabetes, smoking history, and periodontal disease history, and constructs a multivariate logistic regression threshold calibration model to generate accurate thresholds for subjects of different ages and underlying diseases in outpatient settings. The intervention suggestion unit adds clinical treatment pathway recommendations, directly linking high-risk subjects to joint treatment plans from periodontology and cardiology departments, recommending local drug interventions for medium-risk subjects, and providing oral hygiene guidance and cardiovascular health follow-up plans for low-risk subjects.

[0040] In this embodiment, the system was applied to 100 subjects in a clinical outpatient setting, including 65 patients clinically diagnosed with periodontal disease and 35 healthy subjects. The system predicted a true positive rate of 93.8% and a true negative rate of 97.1%, with both the misdiagnosis rate and the missed diagnosis rate being less than 5%. Compared with traditional clinical probing and testing, the efficiency is improved by 60%, and no professional physician is required to operate it, thus reducing clinical labor costs.

[0041] Compared with Example 2-1, using the same clinical data as Example 2, but without adding the functional features of the microbial community, only the abundance features of the microbial community were used for prediction. The prediction accuracy dropped to 90.2%, and the identification rate of patients with severe periodontitis decreased by 12%.

[0042] Compared with Example 2-2, the same clinical data as in Example 2 were used, but an uncompressed large-scale deep learning model was used with a parameter compression rate of 0. The model running time increased by 3 times, making it impossible to run on portable clinical devices. It could only run on high-performance servers, resulting in poor practicality.

[0043] Example 3: This example provides a solution for a periodontal disease risk prediction system for large-scale population screening. Based on Example 1, it simplifies the process, improves detection speed, and adapts to community health checkups and home screening scenarios. The specific process is as follows: The multimodal data acquisition module adopts a minimally invasive acquisition method, using only a portable ECG and blood pressure dual-function device to collect core cardiovascular data (systolic blood pressure, diastolic blood pressure, heart rate, and heart rate variability). It uses a disposable oral swab self-sampling kit, where subjects collect oral mucosal samples themselves. The samples are then sent to a rapid sequencing platform by the community. Using rapid 16S rRNA sequencing technology, the sequencing time is reduced to 4 hours. The data is transmitted to the system via the cloud, enabling non-invasive sampling at home / community without the need for clinical equipment support.

[0044] The heterogeneous data preprocessing module employs a lightweight processing algorithm. For cardiovascular data, it uses mean denoising and simple normalization. For oral microbiota data, it uses a simplified sequence quality control process, retaining core periodontal pathogens (such as Rhodopseudomonas, Porphyromonas gingivalis, and Peptostreptococcus) and removing irrelevant microorganisms. This generates a 32-dimensional oral microbiota feature vector, reducing preprocessing time to 5 minutes and lowering the system's computing power requirements.

[0045] The cross-modal feature fusion module adopts a lightweight cross-modal attention structure, reduces attention layer parameters, retains the core feature weighting function, and quickly completes the fusion of cardiovascular and microbiome features to generate a 32-dimensional fusion feature vector. The fusion time is less than 1 minute, which is suitable for the rapid processing needs of large-scale screening.

[0046] The dual-branch risk prediction model module uses an ultra-lightweight neural network. The dual-branch structure is simplified to a 2-layer temporal convolutional network and a 2-layer multilayer perceptron. The model parameter compression rate reaches 80%. It can run on cloud servers and mobile terminals such as home tablets and mobile phones. The prediction time for a single sample is no more than 10 seconds, which meets the needs of batch processing of large-scale populations.

[0047] The risk grading output module uses a simplified dynamic threshold calibration, with age and gender as the core covariates, to generate generalized personalized thresholds. After the risk level is output, it is simultaneously pushed to the community health management center to provide free oral examination appointment services for high-risk groups, home care packages for medium-risk groups, and health science knowledge for low-risk groups.

[0048] In this embodiment, the system was applied to screen 500 middle-aged and elderly people in the community. The entire process of sampling, testing and prediction was completed automatically without the need for professional personnel. The screening efficiency reached 50 people per hour, and the prediction accuracy rate reached 89.7%. Compared with traditional community oral examinations, the efficiency was improved by 10 times. 32 patients with undiagnosed early periodontal disease were successfully screened, realizing the early detection and early intervention of periodontal disease.

[0049] Compared with Example 3-1, using the same screening data as Example 3, using fixed threshold grading, and without dynamic calibration, the prediction accuracy dropped to 82.3%, and the high-risk missed diagnosis rate in the middle-aged and elderly population reached 15%.

[0050] Compared with Example 3-2, using the same screening data as Example 3, but removing the data fusion step, cardiovascular risk and oral microbiota risk were output separately without joint prediction. The final risk assessment consistency was only 78.5%, and the results had low reliability.

[0051] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of protection claimed by the present invention. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A periodontal disease risk prediction system based on the fusion of cardiovascular and oral microbiota data, characterized in that, It includes a multimodal data acquisition module, a heterogeneous data preprocessing module, a cross-modal feature fusion module, a dual-branch risk prediction model module, and a risk classification output module, which are connected in sequence via communication. The multimodal data acquisition module is used to non-invasively acquire cardiovascular physiological data and oral microbiota sequencing data of the subjects. The cardiovascular physiological data includes resting blood pressure, heart rate variability characteristics, serum lipid indicators, and electrocardiogram time and frequency domain characteristics. The oral microbiota sequencing data is 16S rRNA high-throughput sequencing data from oral mucosal swabs. The heterogeneous data preprocessing module is used to perform noise reduction, missing value filling and normalization on cardiovascular time-series physiological data, and to perform sequence quality control, species annotation, abundance standardization and diversity index calculation on oral microbiota data. The cross-modal feature fusion module employs a cross-modal attention mechanism to adaptively align and deeply fuse the preprocessed cardiovascular feature vector and oral microbiota feature vector, generating a fused feature vector. The dual-branch risk prediction model module includes a cardiovascular feature branch, an oral microbiota feature branch, and a fusion discrimination branch. After being trained on a periodontal disease-cardiovascular comorbidity labeled dataset, the model outputs the probability of periodontal disease risk based on the fusion feature vector. The risk grading output module has a built-in dynamic threshold calibration unit that dynamically adjusts the risk threshold based on the subject's age, gender, and oral hygiene habits, converting the risk probability into three levels of periodontal disease risk and outputting them.

2. The periodontal disease risk prediction system based on the fusion of cardiovascular and oral microbiota data according to claim 1, characterized in that: The multimodal data acquisition module is compatible with home-use portable physiological testing devices and clinical microbiology testing devices. It automatically collects and transmits data through a standardized interface without the need for manual format conversion.

3. The periodontal disease risk prediction system based on the fusion of cardiovascular and oral microbiota data according to claim 1, characterized in that: The heterogeneous data preprocessing module processes the microbial community data by removing chimeric sequences, dividing into operable taxa, annotating to the genus level, and calculating the Shannon diversity index and the Simpson diversity index.

4. The periodontal disease risk prediction system based on the fusion of cardiovascular and oral microbiota data according to claim 1, characterized in that: The cross-modal attention mechanism is a structure that combines self-attention and cross-attention, automatically weighting key dimensions related to periodontal disease risk in cardiovascular features and oral microbiota features.

5. The periodontal disease risk prediction system based on the fusion of cardiovascular and oral microbiota data according to claim 1, characterized in that: The dual-branch risk prediction model is a lightweight neural network. The cardiovascular branch uses a temporal convolutional network to process physiological time-series data, the microbiome branch uses a multilayer perceptron to process omics data, and the fusion discrimination branch uses a fully connected layer to output risk probabilities.

6. The periodontal disease risk prediction system based on the fusion of cardiovascular and oral microbiota data according to claim 1, characterized in that: The dynamic threshold calibration unit takes the subject's baseline characteristics as input and generates an appropriate risk grading threshold through a linear regression model. The threshold is updated in real time according to the subject's baseline characteristics.

7. The periodontal disease risk prediction system based on the fusion of cardiovascular and oral microbiota data according to claim 1, characterized in that: The risk grading output module is also associated with the intervention suggestion unit, which outputs corresponding oral care and cardiovascular health management suggestions based on different risk levels.

8. The periodontal disease risk prediction system based on the fusion of cardiovascular and oral microbiota data according to claim 5, characterized in that: The lightweight neural network has a model parameter compression rate of no less than 70% and can be deployed on edge computing terminals and mobile terminals.

9. The periodontal disease risk prediction system based on the fusion of cardiovascular and oral microbiota data according to claim 1, characterized in that: The periodontal disease-cardiovascular comorbidity labeled dataset contains paired cardiovascular data and oral microbiota data of clinically diagnosed periodontal disease patients and healthy subjects, and the labeling information includes the severity grading of periodontitis.