Elderly caries onset risk prediction model construction method
By collecting multi-dimensional oral data and constructing a risk prediction model for caries in the elderly, the accuracy and stability of caries prediction in the middle-aged and elderly people are solved, and the support of personalized health management and early intervention is achieved.
Patent Information
- Application Number
- CN202510349488.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-08
AI Technical Summary
The existing methods for predicting caries in the elderly rely on a single diagnostic data and cannot fully reflect the patient's health status. It ignores the comprehensive analysis and time-varying of multi-dimensional oral data, resulting in low accuracy and stability of prediction results.
Multi-dimensional oral data were collected, including three-dimensional morphological characteristics of the occlusal surface, time sequence data of saliva microbial metabolic maps, and bioelectrochemical data of dental restoration materials. Through sliding time window mechanism and time-varying feature fusion technology, a risk prediction model for caries in the elderly was constructed, and a cross-cycle verification strategy was adopted.
It improves the accuracy and stability of caries risk prediction in elderly people, provides scientific basis for personalized health management and treatment plans, and supports early intervention and decision-making by generating a time-dimensional risk rating map.
Smart Images

Figure CN120280143A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of healthcare, and particularly to a method for constructing a prediction model for the incidence risk of geriatric caries. Background Art
[0002] Geriatric caries refers to a dental caries disease caused by changes in oral health status, especially occlusal surface wear, microbial metabolic disorders, and the use of dental restoration materials. It is particularly common in the elderly population. With the intensification of population aging, the incidence of geriatric caries has been increasing year by year, and its impact on the elderly population has become increasingly serious, directly affecting the quality of life and health status of the elderly. Effective prediction of geriatric caries can not only provide a basis for early intervention but also help improve the overall oral health level of patients. Therefore, early and accurate prediction of the occurrence of geriatric caries has important public health significance.
[0003] Most existing methods for predicting geriatric caries rely on single diagnostic data, such as clinical examinations and oral X-rays. Although these methods can provide certain prediction bases, they often cannot comprehensively reflect the health status of patients. In the prior art, many models ignore the comprehensive analysis of multi-dimensional oral data, and most methods are difficult to process dynamic data that changes over time, resulting in great limitations in time series data processing and long-term risk prediction. In addition, existing prediction models often ignore the mutual relationship and time-variability between features, and lack an effective evaluation of the importance of different features at different time periods, making the accuracy and stability of the prediction results not high.
[0004] The purpose of the present invention is to provide a method for constructing a prediction model for the incidence risk of geriatric caries based on multi-dimensional oral data, which not only improves the accuracy and stability of risk prediction but also provides a scientific basis for the formulation of personalized health management and treatment plans, thereby realizing accurate prediction and effective intervention of geriatric caries. Summary of the Invention
[0005] The present invention provides a method for constructing a prediction model for the incidence risk of geriatric caries.
[0006] A method for constructing a prediction model for the incidence risk of geriatric caries includes the following steps:
[0007] S1, Oral data collection: Collect multi-dimensional oral data of the target population, including three-dimensional topographic features of the occlusal surface, time-series data of the salivary microbial metabolic profile, and bioelectrochemical data of dental restoration materials;
[0008] S2, Data preprocessing: Perform dynamic preprocessing on the collected multi-dimensional oral data, including reconstructing the occlusal stress distribution, aligning the microbial metabolic pathway features, standardizing the ion precipitation rate of the restoration materials, and generating a preprocessed data set with timestamp marks;
[0009] S3, Risk feature matrix extraction: Based on the sliding time window mechanism, extract the risk feature matrix from the preprocessed dataset, including the cumulative amount of mechanical wear, the entropy change rate of the microbial community metabolism, and the evolution degree of the micro-leakage at the repair interface;
[0010] S4, Feature fusion: Generate the feature fusion weight coefficients according to the time-varying correlation degree of the risk feature matrix, and fuse the risk feature matrix into a composite risk feature vector;
[0011] S5, Model construction: Based on the composite risk feature vector, construct a prediction model for the onset risk of geriatric caries;
[0012] S6, Cross-cycle verification and risk prediction: Adopt the cross-cycle verification strategy to verify the prediction model for the onset risk of geriatric caries, and output a risk level map with time dimension prediction ability.
[0013] Optionally, the oral data collection in S1 includes:
[0014] S11, Three-dimensional topography feature collection of the occlusal surface: Use a three-dimensional scanner to perform three-dimensional scanning on the occlusal surface of the target population, obtain the morphological data of the oral occlusal surface, and generate the three-dimensional topography features of the occlusal surface;
[0015] S12, Collection of time-series data of the salivary microbial metabolism map: Use a saliva sample collector to collect saliva samples from the target population, analyze the microbial community and its metabolites in the saliva using mass spectrometry analysis technology, generate the salivary microbial metabolism map, and record its time-series changes;
[0016] S13, Collection of bioelectrochemical data of dental restoration materials: Perform electrochemical tests on the surface of the oral restoration materials of the target population through an electrochemical workstation, measure the electrochemical characteristics of the materials, including conductivity, current response, and ion precipitation rate, and collect the bioelectrochemical reaction data during the use of the restoration materials.
[0017] Optionally, the data preprocessing in S2 includes:
[0018] S21, Reconstruction of the occlusal stress distribution: Perform mechanical analysis on the three-dimensional topography features of the occlusal surface of the target population, reconstruct the occlusal stress distribution map, calculate the stress distribution data under different occlusal states, and smooth the stress distribution data;
[0019] S22, Alignment of microbial metabolic pathway features: Perform multi-level alignment on the time-series data of the salivary microbial metabolism map, and eliminate time delay and metabolic cycle differences;
[0020] S23, Standardization of the ion precipitation rate of the restoration material: Standardize the bioelectrochemical data of the dental restoration material, and use the ion precipitation rate data collected by the electrochemical workstation to remove the differences between materials through a normalization algorithm;
[0021] S24, Generate a preprocessed dataset with timestamp markings: Organize all preprocessed multi-dimensional oral data according to the time series, and add timestamp markings to the data points of each multi-dimensional oral data to generate a preprocessed dataset with timestamp markings.
[0022] Optionally, the reconstruction of the occlusal stress distribution in S21 includes:
[0023] S211, Three-dimensional topography data modeling: Use a three-dimensional scanning device to scan the occlusal surface of the target population to obtain three-dimensional point cloud data, and convert the three-dimensional point cloud data into a three-dimensional geometric model through data processing software;
[0024] S212, Mechanical analysis and stress distribution calculation: According to the three-dimensional geometric model of the target population, apply the finite element analysis (FEA) method to perform mechanical simulation on the occlusal surface and calculate the stress distribution under different occlusal states;
[0025] S213, Reconstruction of the occlusal stress distribution: Combine the mechanical analysis results under different occlusal states, use the numerical integration method to reconstruct the stress distribution, calculate the local stress distribution data of different parts under different occlusal states, and synthesize them into an occlusal stress distribution map;
[0026] S214, Smoothing processing: Smooth the stress distribution map data reconstructed by the moving average method.
[0027] Optionally, the alignment of the microbial metabolic pathway characteristics in S22 includes:
[0028] S221, Preliminary time alignment: For the time series data of the salivary microbial metabolic map collected at different time points, perform preliminary time alignment through linear interpolation;
[0029] S222, Time delay elimination: Use the dynamic time warping (DTW) algorithm to align the time series data of the salivary microbial metabolic map of different samples;
[0030] S223, Elimination of metabolic cycle differences: Use the fast Fourier transform method to convert the time domain data into frequency domain data and remove the differences in periodic changes of different samples.
[0031] Optionally, the extraction of the risk feature matrix in S3 includes:
[0032] S31. Define a sliding time window: Define the size W and sliding step S of the sliding time window. Assume that the preprocessed data set includes N time points;
[0033] S32. Extract mechanical wear cumulative amount: Within each time window, calculate the mechanical wear cumulative amount P(t) by multiplying the stress and contact area within each time window;
[0034] S33. Extract microbial community metabolic entropy change rate: Within each sliding window, calculate the microbial community metabolic entropy change rate ΔH(t);
[0035] S34. Extract the evolution degree of the repair interface microleakage: Within each sliding window, calculate the evolution degree L(t) of the repair interface microleakage by calculating the ion precipitation rate and the electrochemical characteristics of the material;
[0036] S35. Construct a risk feature matrix: Within each time window, construct a risk feature matrix F(t) by extracting the mechanical wear cumulative amount, microbial community metabolic entropy change rate, and evolution degree of the repair interface microleakage.
[0037] Optionally, the feature fusion in S4 includes:
[0038] S41. Calculate the time-varying correlation degree: By calculating the time-varying correlation degree between different risk feature matrices, measure the importance of each feature at different time points. The time-varying correlation degree is calculated using the Pearson correlation coefficient;
[0039] S42. Generate feature fusion weight coefficients: Generate the weight coefficients w i (t) of feature fusion, and use the time-varying correlation degree as the weight of feature fusion;
[0040] S43. Fuse risk feature vectors: Weightedly fuse all risk feature matrices through the corresponding weight coefficients to obtain a composite risk feature vector F composite (t).
[0041] Optionally, the elderly caries incidence risk prediction model in S5 adopts a Gaussian process regression (GPR) model. The Gaussian process regression (GPR) model includes:
[0042] S51. Define the multi-dimensional kernel function of the composite risk feature vector: For the composite risk feature vector, introduce a weighted radial basis function (RBF) kernel to handle the non-linear relationship between features;
[0043] S52. Optimize the kernel function parameters: Optimize the parameters of the multi-dimensional kernel function by maximizing the log marginal likelihood function;
[0044] S53. Introduce a noise term: Add a noise term to the covariance matrix
[0045] S54, Risk prediction: Given a training dataset X, corresponding target values y, and a new test point X * make a prediction to obtain a predicted mean μ * and a predicted covariance ∑ * ;
[0046] S55, Result output: Output the predicted value of the incidence risk of geriatric caries (predicted mean μ * ) and the confidence level (predicted covariance ∑ * ), and calculate the corresponding confidence interval
[0047] Optionally, the cross - period validation and risk prediction in S6 include:
[0048] S61, Define time period and data splitting: Define a time period (half - year or 1 year), and split the pre - processed dataset into multiple time segments to verify the prediction ability of the geriatric caries incidence risk prediction model in future time segments;
[0049] S62, Evaluate the cross - period prediction ability of the model: Evaluate the prediction ability of the geriatric caries incidence risk prediction model on the pre - processed dataset for each time period, and calculate its evaluation metrics in different periods, including mean squared error (MSE) and root mean squared error (RMSE);
[0050] S63, Generate a risk level map: Generate a risk level map with a time dimension based on the results of cross - period validation to reflect the prediction results in different time periods.
[0051] Optionally, the generation of the risk level map in S63 includes:
[0052] S631, Risk level classification: Based on the predicted value of the incidence risk of geriatric caries, i.e., the predicted mean μ * , classify the risk level. When μ * ≤0.25, the risk level is low risk; when 0.25 < μ * ≤0.5, the risk level is medium risk; when 0.5 < μ * ≤0.75, the risk level is high risk; when μ * >0.75, the risk level is extremely high risk;
[0053] S632, Generate a risk level map: Based on the classified risk levels, draw a risk level map and visualize different risk levels by using color coding.
[0054] Advantages of the present invention:
[0055] The present invention comprehensively reflects all aspects of an individual's oral health status by collecting and integrating multi-dimensional oral data, including three-dimensional topographical features of the occlusal surface, time-series data of the salivary microbial metabolic profile, and bioelectrochemical data of dental restorative materials. This diversified data collection method can not only accurately depict an individual's oral health status but also reveal biological processes at the microscopic level, providing a more comprehensive evaluation basis for establishing a prediction model for the onset risk of geriatric caries, providing a more reliable basis for disease prediction, improving the effect of personalized treatment and health management, and further promoting the prevention and treatment of geriatric caries.
[0056] The present invention, by adopting a sliding time window mechanism and a time-varying feature fusion technology, combined with a cross-cycle verification strategy, enables the model to accurately capture long-term health change trends when processing dynamically changing oral health data. Through these technologies, the model can adapt to the law of data change over time, thereby improving the accuracy and stability of prediction. In addition, cross-cycle verification ensures the generalization ability of the model within different time periods, enabling it to predict risk changes in future time periods, significantly enhancing the reliability and adaptability of the prediction of geriatric caries risk.
[0057] The present invention presents complex prediction results to doctors and health managers in an intuitive way by generating a risk level map with a time dimension. Different risk levels are distinguished by color coding, and the map clearly shows low-risk and high-risk areas, effectively supporting the decision-making process. This risk level map can not only provide real-time risk assessment but also reflect risk fluctuations over time, providing strong support for the formulation of personalized health intervention and treatment plans. Through this visualization method, medical staff can implement early intervention more precisely, improve the treatment effect and the quality of life of patients, and promote more scientific and personalized health management. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only for the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0059] Figure 1 It is a schematic flow chart of the construction method of the embodiment of the present invention;
[0060] Figure 2 It is a schematic diagram for extracting the risk feature matrix of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; moreover, the accompanying drawings are only for more specific description of the embodiments, and are not intended to specifically limit the present invention.
[0062] It should be noted that in the specification, the mention of "an embodiment", "embodiments", "exemplary embodiments", "some embodiments", etc. indicates that the described embodiments may include specific features, structures or characteristics, but not necessarily every embodiment includes such specific features, structures or characteristics. Additionally, when combining embodiments to describe specific features, structures or characteristics, implementing such features, structures or characteristics in combination with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the relevant art.
[0063] Generally, terms can be understood at least in part from their use in context. For example, at least in part depending on the context, the term "one or more" used herein can be used to describe any feature, structure or characteristic in a singular sense, or can be used to describe a combination of features, structures or characteristics in a plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey a set of exclusive factors, but rather, at least in part depending on the context, can allow for the existence of other factors that are not necessarily explicitly described.
[0064] As Figure 1 - Figure 2 shown, a method for constructing a prediction model for the risk of geriatric caries includes the following steps:
[0065] S1, Oral data collection: Collect multi-dimensional oral data of the target population, including three-dimensional topographic features of the occlusal surface, time-series data of the saliva microbial metabolic profile, and bioelectrochemical data of dental restorative materials;
[0066] S2, Data preprocessing: Dynamically preprocess the collected multi-dimensional oral data, including reconstructing the occlusal stress distribution, aligning the microbial metabolic pathway features, standardizing the ion precipitation rate of the restorative material, and generating a preprocessed data set with timestamp markings;
[0067] S3, Risk feature matrix extraction: Extract a risk feature matrix from the preprocessed data set based on the sliding time window mechanism, including the cumulative amount of mechanical wear, the entropy change rate of the microbial community metabolism, and the evolution degree of the microleakage at the restoration interface;
[0068] S4, Feature fusion: Generate a feature fusion weight coefficient according to the time-varying correlation degree of the risk feature matrix, and fuse the risk feature matrix into a composite risk feature vector;
[0069] S5, Model construction: Based on the composite risk feature vector, construct a prediction model for the incidence risk of geriatric caries;
[0070] S6, Cross-period verification and risk prediction: Use the cross-period verification strategy to verify the prediction model for the incidence risk of geriatric caries, and output a risk level map with time dimension prediction ability;
[0071] Through the above content, the comprehensive information of the occlusal surface, microbial metabolism and restorative materials is effectively integrated, and a prediction model for the incidence risk of geriatric caries can be accurately constructed. By adopting the sliding time window mechanism and time-varying feature fusion technology, combined with the cross-period verification strategy, not only the accuracy and stability of the model are improved, but also its adaptability to time changes is enhanced, and the accurate prediction of the incidence risk of geriatric caries can be achieved.
[0072] The oral data collection in S1 includes:
[0073] S11, Collection of three-dimensional topographic features of the occlusal surface: Use a three-dimensional scanner to perform three-dimensional scanning on the occlusal surface of the target population, obtain the three-dimensional morphological data of the oral occlusal surface, and generate the three-dimensional topographic features of the occlusal surface;
[0074] S12, Collection of time-series data of the salivary microbial metabolism map: Use a saliva sample collector to collect saliva samples from the target population, use mass spectrometry analysis technology to analyze the microbial community and its metabolites in saliva, generate a salivary microbial metabolism map, and record its time-series changes;
[0075] S13, Collection of bioelectrochemical data of dental restorative materials: Perform electrochemical tests on the surface of oral restorative materials of the target population through an electrochemical workstation, measure the electrochemical characteristics of the materials, including conductivity, current response, ion precipitation rate, and collect the bioelectrochemical reaction data during the use of the restorative materials;
[0076] Through the above content, multi-dimensional data related to oral health are comprehensively collected, including occlusal surface morphology, microbial metabolism and electro-chemical reactions of restorative materials. This diversified data collection method can not only accurately depict the individual's oral health status, but also reveal the biological processes at the microscopic level, provide a more comprehensive basis for risk assessment, contribute to more accurate disease prediction, improve the level of personalized treatment and health management, and thus play a key role in the prevention and treatment of oral diseases such as geriatric caries.
[0077] The data preprocessing in S2 includes:
[0078] S21, Occlusal stress distribution reconstruction: Conduct a mechanical analysis on the three-dimensional topographical features of the occlusal surface of the target population, reconstruct the occlusal stress distribution map, calculate the stress distribution data under different occlusal states, and smooth the stress distribution data to eliminate noise and errors, ensuring the accuracy and repeatability of the stress distribution;
[0079] S22, Microbial metabolic pathway feature alignment: Conduct multi-level alignment on the time-series data of the saliva microbial metabolic map, and eliminate time delays and metabolic cycle differences, ensuring that the metabolic pathway data between different samples can be compared on a unified time scale, generating comparable metabolic pathway features;
[0080] S23, Standardization of the ion release rate of restorative materials: Standardize the bioelectrochemical data of dental restorative materials. Use the ion release rate data collected by an electrochemical workstation and remove the differences between materials through a normalization algorithm, ensuring that the electrochemical behaviors of all restorative materials are analyzed under the same standard, thereby eliminating the influence of different materials;
[0081] S24, Generate a preprocessed dataset with timestamp markings: Organize all preprocessed multi-dimensional oral data according to the time series, and add timestamp markings to the data points of each multi-dimensional oral data, generating a preprocessed dataset with timestamp markings;
[0082] Through the above content, it is possible to effectively integrate multi-dimensional oral data from different sources and different dimensions. It not only removes the noise and errors in the data, ensuring the accuracy and repeatability of the data, but also eliminates the influence of material and time differences through a unified time scale and standardization process. This comprehensive and refined preprocessing process improves the comparability and stability of the data, thereby enhancing the prediction accuracy and reliability of the overall model.
[0083] The occlusal stress distribution reconstruction in S21 includes:
[0084] S211, Three-dimensional topographical data modeling: Use a three-dimensional scanning device to scan the occlusal surface of the target population to obtain three-dimensional point cloud data, and convert the three-dimensional point cloud data into a three-dimensional geometric model through data processing software;
[0085] S212, Mechanical analysis and stress distribution calculation: According to the three-dimensional geometric model of the target population, apply the finite element analysis (FEA) method to conduct a mechanical simulation on the occlusal surface, and calculate the stress distribution under different occlusal states, expressed as:
[0086]
[0087] where σ is the stress, F is the acting force, and A is the area of the force-bearing region;
[0088] S213, Occlusal stress distribution reconstruction: Combining the mechanical analysis results under different occlusal states, the stress distribution is reconstructed using numerical integration methods. Under different occlusal states, the local stress distribution data of different parts are calculated and synthesized into an occlusal stress distribution map to reflect the stress transfer and concentration during the overall occlusal process, expressed as:
[0089]
[0090] Among them, σ total is the reconstructed total stress distribution, F i is the acting force of each region i, A i is the stress-bearing area of each region i, and n is the number of local regions;
[0091] S214, Smoothing processing: Due to measurement errors or scanning noise, the stress distribution map data reconstructed by the moving average method are smoothed, expressed as:
[0092]
[0093] Among them, σ smooth (i) is the smoothed stress value, σ total (j) is the reconstructed total stress distribution, and N is the size of the sliding window;
[0094] Through the above content, local fluctuations caused by noise or measurement errors can be effectively eliminated, ensuring the accuracy and repeatability of the occlusal stress distribution. It can not only accurately reconstruct the stress transfer and concentration during the occlusal process, but also remove high-frequency noise through smoothing processing, making the stress distribution data more stable and reliable. It can help better understand and optimize the stress changes during the occlusal process and improve the accuracy of diagnosis and treatment.
[0095] The microbial metabolic pathway feature alignment in S22 includes:
[0096] S221, Preliminary time alignment: For the time-series data of salivary microbial metabolic profiles collected at different time points, preliminary time alignment is performed through linear interpolation, expressed as:
[0097]
[0098] Among them, y(t) is the concentration of the metabolite after interpolation, t0 and t1 are known time points, t is the time point to be interpolated, and y(t1), y(t1) are the metabolite concentrations at t0 and t1 respectively;
[0099] S222, Time Delay Elimination: To eliminate the differences caused by time delay, the Dynamic Time Warping (DTW) algorithm is used to align the time series data of the salivary microbial metabolic profiles of different samples. The DTW algorithm finds the optimal alignment of two time series by calculating the shortest path between them, expressed as:
[0100] D(i,j) = ‖x i -y j ‖ 2 +min{D(i - 1,j), D(i,j - 1), D(i - 1,j - 1)};
[0101] where D(i,j) is the metabolic feature difference between samples x i and y j , ‖x i -y j ‖ 2 represents the Euclidean distance between metabolic features, d(i - 1,j) is the cost (distance) between the previous time point i - 1 and the current time point j, D(i,j - 1) is the cost (distance) between the current time point i and the previous time point j - 1, and D(i - 1,j - 1) is the cost (distance) between the previous time point i - 1 and the previous time point j - 1;
[0102] S223, Metabolic Cycle Difference Elimination: There are differences in the metabolic cycles of different individuals. To eliminate this cycle difference, the Fast Fourier Transform method is used to transform the time domain data into frequency domain data and remove the differences in periodic changes of different samples, expressed as:
[0103]
[0104] where X(f) is the frequency domain representation of the metabolic signal, x(n) is the time domain data, N is the total number of data points, and f is the frequency;
[0105] Through the above content, the time delay and metabolic cycle differences are effectively eliminated, ensuring that the metabolic profiles of different samples can be compared on the same time scale. By combining linear interpolation and the Dynamic Time Warping (DTW) algorithm, the metabolic data at different time points can be accurately aligned, significantly improving the consistency and comparability of the data in time series. In addition, frequency domain analysis further eliminates the fluctuations caused by differences in individual metabolic cycles, ensuring the unity of metabolic pathway characteristics.
[0106] The extraction of the risk feature matrix in S3 includes:
[0107] S31, Define the sliding time window: Define the size W and the sliding step S of the sliding time window. Let the preprocessed data set include N time points;
[0108] S32, Extraction of mechanical wear cumulative amount: Within each time window, calculate the mechanical wear cumulative amount P(t) by multiplying the stress and the contact area within each time window, expressed as:
[0109]
[0110] where σ i is the biting stress at the i-th time point, A i is the contact area at the i-th time point, and W is the window size (time period);
[0111] S33, Extraction of microbial community metabolic entropy change rate: Within each sliding window, calculate the microbial community metabolic entropy change rate ΔH(t), expressed as:
[0112]
[0113] where H(t) is the entropy value at time window t, p j (t) is the proportion of the metabolite concentration of metabolic pathway j at time point t, and m is the number of metabolic pathways;
[0114]
[0115] where ΔH(t) is the entropy change rate, Δt is the time step, and H(t + 1) is the entropy value at time window t + 1;
[0116] S34, Extraction of the evolution degree of the repair interface micro-leakage: Within each sliding window, calculate the evolution degree of the repair interface micro-leakage L(t) by calculating the ion precipitation rate and the electrochemical characteristics of the material, expressed as:
[0117]
[0118] where L(t) is the evolution degree of the micro-leakage of the repair interface, and R(t) is the ion precipitation rate;
[0119] S35, Construction of the risk characteristic matrix: Within each time window, construct the risk characteristic matrix F(t) by extracting the mechanical wear cumulative amount, the microbial community metabolic entropy change rate, and the evolution degree of the repair interface micro-leakage, expressed as:
[0120] F(t) = [P(t), ΔH(t), L(t)];
[0121] Through the above, the dynamic changes of oral health status can be comprehensively reflected. It not only converts time-series data into quantifiable risk characteristics, but also provides a more detailed and comprehensive basis for risk assessment by integrating data from different sources. It can capture the long-term change trends of oral health more accurately, improve the accuracy and reliability of the prediction model, and thus promote more scientific and effective oral health management.
[0122] The feature fusion in S4 includes:
[0123] S41, calculating the time-varying correlation degree: By calculating the time-varying correlation degree between different risk feature matrices, the importance of each feature at different time points is measured. The time-varying correlation degree is calculated using the Pearson correlation coefficient and is expressed as:
[0124]
[0125] where, F i (t) and G i (t) are the values of two features at time point t, and are the means of these two features within the time window, ρ(t) is the correlation coefficient between the features, reflecting their correlation degree at time t, and p is the total number of features;
[0126] S42, generating the feature fusion weight coefficient: Generate the weight coefficient w i (t) of feature fusion according to the time-varying correlation degree, and use the time-varying correlation degree as the weight of feature fusion, which is expressed as:
[0127]
[0128] where, w i (t) is the weight coefficient of the i-th feature, ρ i (t) is the time-varying correlation degree of feature i at time point t, and q is the total number of risk feature matrices;
[0129] S43, fusing the risk feature vectors: Weightedly fuse all risk feature matrices through the corresponding weight coefficients to obtain the composite risk feature vector F composite (t), which is expressed as:
[0130]
[0131] where, F composite (t) is the fused composite risk feature vector, w i (t) is the weight coefficient of the i-th feature, and F i (t) is the value of the i-th risk feature matrix at time point t;
[0132] Through the above, multiple risk features can be effectively integrated to generate a composite risk feature vector. Through this weighted integration, the features that are most important for risk assessment within a specific time period can be highlighted, thereby improving the prediction accuracy and stability of the model. Compared with the analysis method of single features, the feature fusion method can comprehensively consider the time-series changes of multi-dimensional data, avoiding the limitation that single features are insufficient to comprehensively reflect risks, and thus providing strong data support for more accurate risk assessment and personalized health management.
[0133] The prediction model for the incidence risk of geriatric caries in S5 adopts the Gaussian process regression (GPR) model. The Gaussian process regression (GPR) model includes:
[0134] S51, defining the multi-dimensional kernel function of the composite risk feature vector: For the composite risk feature vector, a weighted radial basis function (RBF) kernel is introduced to handle the non-linear relationship between features, expressed as:
[0135]
[0136] where, x i and x i ' are the values of the i-th feature in the composite risk feature vectors x and x', is the length scale of the i-th feature, controlling the similarity between features, σ 2 is the global scale parameter of the kernel function, representing the variance of the data, and k is the number of features;
[0137] S52, optimizing the kernel function parameters: Optimize the parameters of the multi-dimensional kernel function (such as σ 2 and ) by maximizing the log marginal likelihood function, expressed as:
[0138]
[0139] where, y is the vector of the target variable (incidence risk of geriatric caries), K is the covariance matrix calculated through the kernel function, and K ij = k(x i , x j ) is the similarity between the i-th sample and the j-th sample, log det K is the determinant of the covariance matrix, and l is the number of samples;
[0140] S53, introducing a noise term: To handle the noise in the actual data, a noise term is added to the covariance matrix to make the model more robust, expressed as:
[0141]
[0142] where, is the variance of the noise term, representing the inevitable error in the data, δ xx' is the Kronecker delta function, which is 1 when x = x', representing self-connection;
[0143] S54, Risk Prediction: Given the training dataset X, the corresponding target values y, and a new test point X * perform prediction to obtain the predicted mean μ * and the predicted covariance ∑ * , expressed as:
[0144]
[0145] where μ * is the predicted mean (i.e., the predicted value of the incidence risk of senile caries), K(X * , X) is the covariance matrix between the test point and the training set, K(X, X) is the covariance matrix within the training set, is the noise term diagonal matrix;
[0146]
[0147] where ∑ * is the predicted covariance matrix, i.e., the uncertainty of the prediction result, K(X * , X * ) is the covariance matrix between the test point X * and itself, K(X * , X) is the covariance matrix between the test point X * and the training data point X, representing the similarity between the test point and the training data, K(X, X) is the covariance matrix between the training data points X, representing the similarity between the training data points, is the noise variance, representing the noise level in the data, I is the identity matrix, is the inverse matrix after adding the covariance matrix and the noise term, representing the inverse correlation between the training data points considering the noise;
[0148] S55, Result Output: Output the predicted value of the incidence risk of senile caries (the predicted mean μ * ) and the confidence level (the predicted covariance ∑ * ), and calculate the corresponding confidence interval
[0149] Through the above, it is possible to effectively process the composite risk feature vectors in the prediction of the incidence risk of geriatric caries, capture the complex non-linear relationships between features, and provide accurate prediction results in the high-dimensional feature space. By maximizing the log marginal likelihood function, the model automatically adjusts the parameters and learns in the way that best suits the data. At the same time, the added noise term improves the robustness of the model to the uncertainties in the data, enabling it to better cope with measurement errors and data noise in practical applications. In addition, by providing the prediction mean and confidence interval, the GPR model can not only output accurate prediction values for the incidence risk of geriatric caries, but also quantify the uncertainty of the prediction results, providing a solid basis for personalized health management and the formulation of treatment plans, and being able to provide more reliable predictions and uncertainty assessments, providing scientific support for medical decision-making, thereby improving the accuracy and interpretability of the prediction of the incidence risk of geriatric caries.
[0150] The cross-period validation and risk prediction in S6 include:
[0151] S61, defining the time period and data splitting: Define the time period (half a year or one year), and split the preprocessed data set into multiple time segments to verify the prediction ability of the geriatric caries incidence risk prediction model in future time segments;
[0152] S62, evaluating the cross-period prediction ability of the model: By evaluating the prediction ability of the geriatric caries incidence risk prediction model on the preprocessed data set in each time period, calculate its evaluation metrics in different periods, including the mean squared error (MSE) and the root mean squared error (RMSE), expressed as:
[0153]
[0154]
[0155] where, y i is the true incidence risk, μ * is the incidence risk predicted by the model, and E is the number of samples in the test data;
[0156] S63, generating a risk level map: According to the results of the cross-period validation, generate a risk level map with a time dimension to reflect the prediction results in different time periods;
[0157] Through the above, it is possible to effectively evaluate the performance of the geriatric caries incidence risk prediction model in different time periods and ensure that the model has strong prediction ability in the time dimension. By splitting the data into multiple periods, the cross-period validation not only examines the long-term stability and generalization ability of the model, but also ensures the prediction ability of the model in future periods, thereby improving the reliability of the prediction, being able to accurately capture the risk change trends at different time points, and by generating a risk level map, helping health managers to adjust the intervention strategies in real time.
[0158] The generation of the risk level map in S63 includes:
[0159] S631, risk level division: Based on the predicted value of the incidence risk of geriatric caries, i.e., the predicted mean μ * , the risk level is divided. When μ * ≤0.25, the risk level is low risk. When 0.25 < μ * ≤0.5, the risk level is medium risk. When 0.5 < μ * ≤0.75, the risk level is high risk. When μ * >0.75, the risk level is extremely high risk;
[0160] S632, generating the risk level map: Based on the divided risk levels, draw the risk level map, and visualize different risk levels by using color coding;
[0161] Through the above content, the complex prediction results can be presented in an intuitive way, helping doctors and health managers better understand the changing trend of the incidence risk of geriatric caries in patients over time. By distinguishing different risk levels with colors, the map makes the high-risk and low-risk areas clear at a glance, thus effectively supporting the decision-making process. The risk level map can not only provide real-time risk assessment, but also reflect the risk fluctuations in the time dimension, helping to predict possible future risk changes, and further providing a scientific basis for the formulation of personalized health intervention and treatment plans. Through this visualization method, medical staff can implement early intervention more precisely, optimize health management, and improve the treatment effect and quality of life of patients.
[0162] This invention covers any substitutions, modifications, equivalent methods, and solutions made within the essence and scope of this invention. To enable the public to have a thorough understanding of this invention, specific details are described in detail in the following preferred embodiments of this invention. However, those skilled in the art can fully understand this invention without the description of these details. Additionally, to avoid unnecessary confusion to the essence of this invention, well-known methods, processes, procedures, components, and circuits are not described in detail.
[0163] The above description is only a preferred embodiment of this invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of this invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this invention.
Claims
1. A method for constructing a prediction model for the risk of geriatric caries, characterized in that, It includes the following steps: S1, Oral data collection: Collect multi-dimensional oral data of the target population, including three-dimensional topographic features of the occlusal surface, time-series data of the salivary microbial metabolic profile, and bioelectrochemical data of dental restorative materials; S2, Data preprocessing: Dynamically preprocess the collected multi-dimensional oral data, including reconstructing the occlusal stress distribution, aligning the microbial metabolic pathway features, standardizing the ion release rate of the restorative material, and generating a preprocessed dataset with timestamp markings; S3, Risk feature matrix extraction: Extract the risk feature matrix from the preprocessed dataset based on the sliding time window mechanism, including the cumulative amount of mechanical wear, the entropy change rate of the microbial community metabolism, and the evolution degree of the microleakage at the restoration interface; S4, Feature fusion: Generate the feature fusion weight coefficient according to the time-varying correlation degree of the risk feature matrix, and fuse the risk feature matrix into a composite risk feature vector; S5, Model construction: Based on the composite risk feature vector, construct a prediction model for the incidence risk of geriatric caries; S6, Cross-cycle verification and risk prediction: Use the cross-cycle verification strategy to verify the prediction model for the incidence risk of geriatric caries, and output a risk level map with time-dimensional prediction ability.
2. The method for constructing a prediction model for the onset risk of geriatric caries according to claim 1, wherein The oral data collection in S1 includes: S11, Collection of three-dimensional topographic features of the occlusal surface: Use a three-dimensional scanner to perform three-dimensional scanning on the occlusal surface of the target population, obtain the oral occlusal surface morphology data, and generate the three-dimensional topographic features of the occlusal surface; S12, Collection of time-series data of the salivary microbial metabolic profile: Use a saliva sample collector to collect saliva samples of the target population, use mass spectrometry analysis technology to analyze the microbial community and its metabolites in the saliva, generate the salivary microbial metabolic profile, and record its time-series changes; S13, Collection of bioelectrochemical data of dental restorative materials: Perform electrochemical tests on the surface of the oral restorative materials of the target population through an electrochemical workstation, measure the electrochemical properties of the materials, including conductivity, current response, and ion release rate, and collect the bioelectrochemical reaction data during the use of the restorative materials.
3. A method for constructing a prediction model for the risk of geriatric caries according to claim 1, characterized in that, The data preprocessing in S2 includes: S21, Reconstruction of the occlusal stress distribution: Perform mechanical analysis on the three-dimensional topographic features of the occlusal surface of the target population, reconstruct the occlusal stress distribution map, calculate the stress distribution data under different occlusal states, and smooth the stress distribution data; S22, Alignment of microbial metabolic pathway features: Perform multi-level alignment on the time-series data of the salivary microbial metabolic profile, and eliminate time delay and metabolic cycle differences; S23, Standardization of the ion release rate of the restorative material: Perform standardization processing on the bioelectrochemical data of the dental restorative materials, and use the ion release rate data collected by the electrochemical workstation to remove the differences between materials through a normalization algorithm; S24, Generation of a preprocessed dataset with timestamp markings: Organize all the preprocessed multi-dimensional oral data according to the time series, and add timestamp markings to each data point of the multi-dimensional oral data to generate a preprocessed dataset with timestamp markings.
4. The method for constructing a prediction model for the risk of geriatric caries according to claim 3, characterized in that, The reconstruction of the occlusal stress distribution in S21 includes: S211, 3D Topography Data Modeling: Use a 3D scanning device to scan the occlusal surface of the target population to obtain 3D point cloud data, and convert the 3D point cloud data into a 3D geometric model through data processing software; S212, Mechanical Analysis and Stress Distribution Calculation: According to the 3D geometric model of the target population, apply the finite element analysis method to perform mechanical simulation on the occlusal surface and calculate the stress distribution under different occlusion states; S213, Reconstruction of Occlusal Stress Distribution: Combine the mechanical analysis results under different occlusion states, and use the numerical integration method to reconstruct the stress distribution. Under different occlusion states, calculate the local stress distribution data of different parts and synthesize them into an occlusal stress distribution map; S214, Smoothing Processing: Use the moving average method to smooth the data of the reconstructed stress distribution map.
5. The method for constructing a prediction model for the onset risk of geriatric caries according to claim 4, wherein, The microbial metabolic pathway feature alignment in S22 includes: S221, Preliminary Time Alignment: For the time-series data of saliva microbial metabolic maps collected at different time points, perform preliminary time alignment through linear interpolation; S222, Time Delay Elimination: Use the dynamic time warping algorithm to align the time-series data of saliva microbial metabolic maps of different samples; S223, Elimination of Metabolic Cycle Differences: Use the fast Fourier transform method to convert the time-domain data into frequency-domain data and remove the differences in periodic changes among different samples.
6. The method for constructing a prediction model for the risk of geriatric caries according to claim 1, wherein The risk feature matrix extraction in S3 includes: S31, Define a Sliding Time Window: Define the size W and sliding step S of the sliding time window, and assume that the preprocessed data set includes N time points; S32, Extraction of Mechanical Wear Accumulation: Within each time window, calculate the mechanical wear accumulation P(t) by multiplying the stress and contact area within each time window; S33, Extraction of Microbial Community Metabolic Entropy Change Rate: Within each sliding window, calculate the microbial community metabolic entropy change rate ΔH(t); S34, Extraction of the Evolution Degree of Restoration Interface Microleakage: Within each sliding window, calculate the evolution degree L(t) of the restoration interface microleakage by calculating the ion precipitation rate and the electrochemical characteristics of the material; S35, Construct a Risk Feature Matrix: Within each time window, construct a risk feature matrix F(t) by extracting the mechanical wear accumulation, microbial community metabolic entropy change rate, and the evolution degree of restoration interface microleakage.
7. A method for constructing a prediction model for the risk of geriatric caries according to claim 6, characterized in that, The feature fusion in S4 includes: S41, Calculate the Time-Varying Correlation Degree: By calculating the time-varying correlation degree between different risk feature matrices, measure the importance of each feature at different time points. The time-varying correlation degree is calculated using the Pearson correlation coefficient; S42, Generate the feature fusion weight coefficient: Generate the weight coefficient w i (t) of feature fusion according to the time-varying correlation degree, and use the time-varying correlation degree as the weight of feature fusion; S43, Fusion of risk feature vectors: All risk feature matrices are weighted and fused with corresponding weight coefficients to obtain a composite risk feature vector F composite (t).
8. A method for constructing a prediction model for the risk of geriatric caries according to claim 7, characterized in that, The elderly caries incidence risk prediction model in S5 uses a Gaussian process regression model, and the Gaussian process regression model includes: S51, Define the Multidimensional Kernel Function of the Composite Risk Feature Vector: For the composite risk feature vector, introduce a weighted radial basis function kernel to handle the nonlinear relationship between features; S52, Optimize the Kernel Function Parameters: Optimize the parameters of the multidimensional kernel function by maximizing the log marginal likelihood function; S53, Introduce a noise term: Add a noise term to the covariance matrix S54, Risk Prediction: Given a training dataset X, corresponding target values y, and a new test point X * make a prediction to obtain the predicted mean μ * and the predicted covariance ∑ * ; S55, Result Output: Output the elderly caries incidence risk prediction value and confidence level, and calculate the corresponding confidence interval.
9. A method for constructing a prediction model for the risk of geriatric caries according to claim 8, characterized in that, The cross-cycle verification and risk prediction in S6 include: S61, Define time period and data segmentation: Define a time period and segment the preprocessed dataset into multiple time segments for verifying the prediction ability of the geriatric caries incidence risk prediction model in future time segments; S62, Evaluate the cross-cycle prediction ability of the model: Evaluate the prediction ability of the geriatric caries incidence risk prediction model on the preprocessed dataset of each time period, and calculate its evaluation metrics in different periods, including mean squared error and root mean squared error; S63, Generate a risk level map: According to the results of cross-cycle verification, generate a risk level map with a time dimension to reflect the prediction results in different time periods.
10. A method for constructing a prediction model for the risk of geriatric caries according to claim 9, characterized in that, The generation of the risk level map in S63 includes: S631, Risk level classification: Based on the predicted value of the incidence risk of geriatric caries, i.e., the predicted mean μ * , classify the risk level. When μ * ≤0.25, the risk level is low risk. When 0.25 < μ * ≤0.5, the risk level is medium risk. When 0.5 < μ * ≤0.75, the risk level is high risk. When μ * >0.75, the risk level is extremely high risk; S632, Generate a risk level map: Based on the divided risk levels, draw a risk level map and visualize different risk levels by using color coding.