Method, device and storage medium for caries risk assessment
By combining the features of genetic data and oral imaging data, the problem of inaccurate caries assessment caused by insufficient resolution of oral imaging data is solved, and a more accurate and comprehensive caries risk assessment is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PEKING UNIV SCHOOL OF STOMATOLOGY
- Filing Date
- 2026-04-27
- Publication Date
- 2026-07-24
AI Technical Summary
In existing technologies, when caries risk assessment is performed solely based on oral imaging data, the image resolution is insufficient, resulting in unclear early pathological features of caries and inaccurate assessment results.
By combining genetic data, risk alleles at multiple gene loci associated with susceptibility to dental caries are numerically encoded to obtain gene features, which are then fused with oral imaging features to form a fused feature for risk assessment.
It improves the accuracy and comprehensiveness of caries risk assessment, and can reflect the patient's caries risk from both congenital susceptibility and acquired lesions.
Smart Images

Figure CN122455337A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of oral image processing technology, and in particular relates to a method, device and storage medium for caries risk assessment. Background Technology
[0002] Dental caries is a common chronic disease in oral clinics, and caries risk assessment is an important technical means for the prevention and intervention of oral diseases. Current technologies primarily rely on oral imaging data to detect the presence of caries. For example, image analysis techniques are used to identify caries lesions in the oral cavity, determine the severity of caries, and use this as the result of caries risk assessment.
[0003] However, relying solely on oral imaging data for caries assessment results in inaccurate results because the image resolution is insufficient, making it difficult to clearly display the pathological features of early caries in the images. Summary of the Invention
[0004] This application provides a method, device, and storage medium for caries risk assessment, which can improve the accuracy of caries risk assessment.
[0005] In a first aspect, embodiments of this application provide a method for assessing the risk of dental caries, the method comprising: Obtain multi-source data corresponding to the caries sample to be evaluated. The multi-source data includes at least: oral imaging data and genetic data. Based on the risk alleles of multiple gene loci associated with susceptibility to dental caries, the gene data are numerically encoded to obtain gene characteristics; Feature extraction is performed on oral imaging data to obtain oral imaging features; Feature fusion is performed on genetic features and oral imaging features to obtain fused features; The caries risk was assessed based on the fusion characteristics to obtain the caries risk assessment results for the caries samples to be evaluated.
[0006] In some possible implementations, gene data is numerically encoded based on risk alleles at multiple gene loci associated with susceptibility to dental caries, resulting in gene characteristics, including: Obtain genotypic information of multiple gene loci associated with susceptibility to dental caries from genetic data; The genotype information of each gene locus is compared with the risk alleles corresponding to that gene locus, and the number of risk alleles carried in the genotype information is counted. The statistically obtained number is used as the coding value of the gene locus; The coding values of each gene locus are combined to form gene features.
[0007] In some possible implementations, before combining the coding values of each gene locus into a gene feature, the following steps are also included: Calculate the risk weight corresponding to the risk allele at each gene locus; whereby the risk weight is used to characterize the strength of the association between a single risk allele and the risk of dental caries. The coding values and risk weights of each gene locus are combined to form gene features.
[0008] In some possible implementations, the risk weights corresponding to risk alleles at each gene locus are calculated, including: We obtained sample data from the exposed group and the control group corresponding to the risk allele. The exposed group sample data included dental caries prevalence information of the first population carrying the risk allele, and the control group sample data included dental caries prevalence information of the second population not carrying the risk allele. The disease advantage in the exposed group was calculated based on sample data from the exposed group, and the disease advantage in the control group was calculated based on sample data from the control group. Divide the disease prevalence of the exposed group by the disease prevalence of the control group to obtain the risk weights corresponding to the risk alleles.
[0009] In some possible implementations, before fusing genetic features and oral imaging features to obtain the fused features, the following steps are also included: Acquire oral health record data and lifestyle habit data, extract features from the oral health record data and lifestyle habit data to obtain oral health record features and lifestyle habit data features; Non-image fusion features are obtained by splicing features from oral health record features, lifestyle data features, and genetic features. The non-image fusion features and oral image features are fused to obtain the fused features.
[0010] In some possible implementations, before performing feature fusion on the multimodal features to obtain the fused features, the following steps are also included: Image enhancement processing is performed on oral imaging data to obtain enhanced oral imaging data, and oral imaging features are extracted from the enhanced oral imaging data; the image enhancement processing includes at least one of geometric transformation and pixel-level transformation.
[0011] Non-image feature enhancement processing is performed on the features of oral health records, lifestyle data features, and genetic features to obtain enhanced non-image features. The non-image feature enhancement processing includes at least one of the following: standardization processing, normalization processing, nonlinear transformation, and categorical feature enhancement.
[0012] In some possible implementations, after assessing the risk of fusion features to obtain a caries risk score for the caries sample to be evaluated, the following is also included: Obtain follow-up data for the caries sample to be evaluated. The follow-up data includes at least one of the following: updated oral imaging data, updated oral health record data, and updated lifestyle data. Feature extraction was performed on gene data and follow-up data to obtain updated multimodal features; The updated multimodal features are fused to obtain the updated fused features; The caries risk was assessed based on the fusion characteristics after the update, and the caries risk assessment results of the caries sample to be evaluated after the update were obtained.
[0013] In some possible implementations, multiple gene loci associated with susceptibility to dental caries include at least one of the following: gene loci encoding enamel matrix proteins, gene loci encoding salivary proteins, and gene loci encoding immunomodulatory factors and metabolic receptors.
[0014] Secondly, embodiments of this application provide a dental caries risk assessment device, the device comprising: The acquisition module is used to acquire multi-source data corresponding to the caries sample to be evaluated. The multi-source data includes at least: oral imaging data and genetic data. The coding module is used to numerically encode gene data based on the risk alleles of multiple gene loci associated with susceptibility to dental caries, thereby obtaining gene characteristics; The feature extraction module is used to extract features from oral imaging data to obtain oral imaging features; The feature fusion module is used to fuse gene features and oral imaging features to obtain fused features; The risk assessment module is used to assess the caries risk of fusion features and obtain the caries risk assessment results of the caries sample to be assessed.
[0015] Thirdly, embodiments of this application provide an electronic device, the device comprising: A processor and a memory storing computer program instructions; a method for assessing the risk of dental caries that implements any of the above when the processor executes the computer program instructions.
[0016] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement any of the above-mentioned methods for assessing the risk of dental caries.
[0017] Fifthly, embodiments of this application provide a computer program product in which the instructions are executed by the processor of an electronic device, enabling the electronic device to perform any of the above-mentioned methods for assessing dental caries risk.
[0018] The dental caries risk assessment method, device, and storage medium provided in this application embodiment numerically encode gene data based on risk alleles of multiple gene loci related to dental caries susceptibility to obtain gene features. Then, the gene features and oral imaging features are fused to obtain fused features, which can comprehensively reflect the patient's dental caries risk from the dimensions of congenital susceptibility and acquired lesions, thereby effectively improving the accuracy and comprehensiveness of dental caries risk assessment. Attached Figure Description
[0019] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 A flowchart illustrating the caries risk assessment method provided in the first embodiment of this application is shown; Figure 2 A flowchart illustrating the caries risk assessment method provided in the second embodiment of this application is shown; Figure 3 A flowchart illustrating the caries risk assessment method provided in the third embodiment of this application is shown; Figure 4 A flowchart illustrating the caries risk assessment method provided in the third embodiment of this application is shown; Figure 5 This illustration shows a structural schematic diagram of a caries risk assessment device provided in an embodiment of this application; Figure 6 A schematic diagram of the hardware structure of the electronic device provided in this application is shown. Detailed Implementation
[0021] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.
[0022] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0023] It should be noted that the acquisition, storage, use, and processing of data in this application embodiment all comply with the relevant provisions of national laws and regulations.
[0024] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.
[0025] In existing technologies, dental caries risk assessment is primarily conducted using oral imaging data. This involves manually identifying caries-affected areas in the images or using image analysis techniques to identify pathological features and determine the severity of caries. However, due to insufficient resolution of oral imaging data, early-stage caries pathological features are not clearly visible on the images. Therefore, risk assessment based solely on oral imaging data is inaccurate.
[0026] To address the problems in the prior art, this application provides a method, device, and storage medium for caries risk assessment. When assessing caries based on oral imaging data, combining genetic data reduces the limitations of single imaging data in representing caries risk and compensates for the shortcomings of imaging data in capturing early caries features. This allows fused features to comprehensively reflect the patient's caries risk from both congenital susceptibility and acquired lesions, thereby assessing the caries risk of the caries sample to be evaluated. This effectively improves the accuracy and comprehensiveness of caries risk assessment.
[0027] The dental caries risk assessment method provided in the embodiments of this application will be introduced first below.
[0028] Figure 1 A flowchart illustrating the caries risk assessment method provided in the first embodiment of this application is shown. Figure 1As shown, the method may include the following steps: S101: Obtain multi-source data corresponding to the caries sample to be evaluated. The multi-source data shall include at least: oral imaging data and genetic data.
[0029] In this embodiment, the multi-source data of the caries sample to be evaluated refers to data related to the occurrence of dental caries, including: oral imaging data and genetic data. Oral imaging data are imaging data reflecting the current state of dental caries in the patient's oral cavity, such as: panoramic radiographs, periapical radiographs, or cone beam computer-to-tomography (CBCT) data. Genetic data are gene locus detection data reflecting the patient's congenital susceptibility to dental caries.
[0030] It should be noted that since the occurrence of dental caries is the result of the combined effects of congenital genetic susceptibility and acquired dental lesions, single imaging data can only reflect the current state of dental caries and cannot predict an individual's risk of congenital dental caries. Therefore, this application introduces genetic data, which can improve the accuracy of dental caries risk assessment from a biological perspective.
[0031] S102: Based on the risk alleles of multiple gene loci associated with susceptibility to dental caries, the gene data are numerically encoded to obtain gene characteristics.
[0032] In the embodiments of this application, dental caries susceptibility refers to an individual's innate tendency to develop dental caries in the same caries-prone environment, which is determined by genetic factors; risk alleles refer to allele forms at specific gene loci that are associated with an increased risk of dental caries.
[0033] In one example, the multiple gene loci associated with susceptibility to dental caries include at least one of the following: a gene locus encoding enamel matrix proteins, a gene locus encoding salivary proteins, and a gene locus encoding immunomodulatory factors and metabolic receptors.
[0034] In this embodiment, the genetic susceptibility to dental caries involves multiple biological processes, including enamel development, salivary secretion and antibacterial activity, and immunity and metabolism. Enamel development determines the tooth structure's resistance to caries, salivary secretion and antibacterial activity determine the cleanliness and buffering capacity of the oral microenvironment, and immunity and metabolism determine the host's response to cariogenic bacteria and its ability to regulate glucose metabolism. Therefore, it is necessary to comprehensively select gene loci encoding enamel matrix proteins, gene loci encoding salivary proteins, and gene loci encoding immune regulatory factors and metabolic receptors to fully obtain the innate genetic susceptibility characteristics for dental caries, thereby mining the patient's innate predisposition to dental caries at the genetic level and providing a genetic basis for dental caries risk assessment.
[0035] Among them, gene loci encoding enamel matrix proteins are involved in the formation and mineralization of tooth enamel. Mutations in these loci can lead to enamel structural defects and increase the risk of dental caries. Specific gene loci may include: AMELX (SNP number rs17878486), which encodes amelogenin and participates in enamel development. Its risk allele G causes abnormal amelogenin function, reducing enamel hardness; ENAM (SNP number rs12640848), which encodes ameloprotein and participates in enamel mineralization. Its risk allele A causes abnormal ameloprotein function, resulting in insufficient enamel mineralization, reduced enamel hardness, and increased susceptibility to corrosion by acidic substances produced by cariogenic bacteria, thus increasing the probability of dental caries; and MMP20 (SNP number rs1784418), which encodes matrix metalloproteinase and participates in enamel maturation. Its risk allele T causes abnormal enzyme activity, hindering the enamel maturation process and reducing enamel hardness. Disordered crystal arrangement, loose enamel structure, and decreased resistance to caries; KLK4 (SNP number rs198968): encodes kallikrein, which participates in enamel maturation. Its risk allele A leads to impaired hydrolytic function of kallikrein, incomplete removal of enamel matrix, impaired enamel maturation, insufficient enamel density, and easy demineralization of tooth structure, which in turn leads to caries; DSPP (SNP number rs36094464): encodes dentin sialophosphoprotein, which affects dentin formation. Its risk allele T leads to loss of dentin sialophosphoprotein function, reducing the integrity of dentin and indirectly accelerating the development of caries.
[0036] Gene loci encoding salivary proteins are responsible for encoding proteins in saliva, affecting saliva flow rate, antibacterial activity, and buffering capacity. Mutations in these loci can impair the physiological functions of saliva, disrupt the homeostasis of the oral microenvironment, and prevent saliva from effectively performing its cleaning, antibacterial, and acid-neutralizing functions, thereby increasing the risk of dental caries. Specific gene loci may include: AQP5 (SNP number rs296763): encodes aquaporin 5, which regulates saliva secretion. Its risk allele C leads to abnormal water transport function of aquaporin 5, reduced saliva secretion, and weakened oral self-cleaning ability, thus increasing the risk of dental caries; MUC7 (SNP number rs1104022): encodes mucin 7, which participates in saliva antibacterial activity. Its risk allele A leads to decreased mucin activity, weakening the adhesion, inhibition, and killing effects of saliva on cariogenic bacteria, thus increasing the risk of dental caries; CA6 (S... NP number rs2274327: Encodes carbonic anhydrase 6, which participates in salivary buffering. Its risk allele T reduces the ability of saliva to produce bicarbonate, resulting in a lower pH value in the oral cavity. The hard tissues of the teeth are in an acidic environment for a long time, which easily leads to demineralization and thus caries. LYZ (SNP number rs1800014): Encodes lysozyme, which participates in antibacterial defense. Its risk allele A leads to reduced lysozyme activity, which cannot effectively clear cariogenic bacteria in the oral cavity. Cariogenic bacteria multiply and form dental plaque, accelerating the demineralization process of the teeth and thus increasing the risk of caries.
[0037] Gene loci encoding immunomodulatory factors and metabolic receptors are involved in immune responses, inflammation regulation, and glucose metabolism, affecting the host's defense against cariogenic bacteria and dietary habits. Mutations in these loci can lead to immune response dysregulation or metabolic abnormalities, increasing the risk of dental caries. Specific gene loci may include: DEFB1 (SNP number rs11362): encoding β-defensin 1, an innate immune antimicrobial peptide. Its risk allele G leads to decreased expression of the antimicrobial peptide, weakening the local defense of the oral mucosa and thus causing dental caries; TAS1R2 (SNP number rs35874116): encoding a sweet taste receptor, affecting the perception and preference for sweetness. Its risk allele T enhances sweet taste preference, prompting individuals to consume more sugary foods, indirectly increasing the risk of dental caries; SLC2A2 (SNP number rs35874116): encoding a sweet taste receptor, affecting the perception and preference for sweetness. Its risk allele T enhances sweet taste preference, prompting individuals to consume more sugary foods, indirectly increasing the risk of dental caries; SLC2A2 (SNP number rs35874116): encoding a sweet taste receptor, affecting the perception and preference for sweetness. Its risk allele T enhances sweet taste preference, prompting individuals to consume more sugary foods, indirectly increasing the risk of dental caries; IL1B (SNP number rs5400): Encodes a glucose transporter involved in glucose metabolism. Its risk allele A leads to abnormal glucose transport function, decreased efficiency in the body's glucose metabolism, and high glucose concentrations in the blood and oral cavity. The high-sugar environment in the oral cavity promotes the proliferation of cariogenic bacteria and the secretion of acidic substances, thereby increasing the risk of dental caries. IL1B (SNP number rs1143634): Encodes interleukin-1β involved in inflammatory responses. Its risk allele T leads to excessive inflammatory responses, accelerating the destruction of tooth tissue and the progression of caries.
[0038] In one example, to achieve rapid encoding of gene data and improve the efficiency of encoding and subsequent risk assessment, adapting to scenarios such as rapid initial screening in primary healthcare institutions and large-scale population screening for dental caries susceptibility, the method for numerically encoding gene data to obtain gene features can be as follows: First, determine the risk alleles corresponding to each caries susceptibility-related gene locus. Then, determine the genotype of each gene locus one by one. If the genotype of a certain locus in the gene data carries at least one risk allele, then the locus is encoded as 1; if it does not carry any risk alleles, it is encoded as 0. Finally, the encoded values of each locus are concatenated sequentially according to the preset gene locus order to form a one-dimensional gene feature vector. The concatenation of the encoded values can be done in ascending order of chromosome position or in alphabetical order of gene name.
[0039] In another example, to more accurately capture the impact of genotype information at each gene locus on caries risk assessment—for instance, given the non-linear effect of genotype information on caries risk—the three possible genotypes at each gene locus can be encoded into an independent binary dimension, forming a unique binary encoding vector for that locus. For example, for the locus AMELX, with the risk allele being G, the genotype "A / A" is encoded as [1,0,0], the genotype "A / G" as [0,1,0], and the genotype "G / G" as [0,0,1]. After completing the binary encoding of a single locus, the binary encoding vectors of each locus are concatenated sequentially according to a preset gene locus order to form a high-dimensional gene feature vector. This gene feature vector can be directly used as input to a multimodal fusion model to participate in subsequent feature fusion and caries risk assessment calculations.
[0040] S103: Extract features from oral imaging data to obtain oral imaging features.
[0041] In this embodiment of the application, the original oral imaging data is pixel-level image information, which contains a large amount of invalid information such as soft tissue and background. Directly using oral imaging data for feature fusion will lead to a decrease in the accuracy and efficiency of fusion calculation. Therefore, by feature extraction, the core imaging indicators related to caries are accurately screened out, and the image information is transformed into quantitative features to improve the accuracy and efficiency of subsequent fusion calculation.
[0042] In one example, feature extraction from oral imaging data can be performed using a pre-trained U-Net segmentation model to segment the tooth region in a panoramic image, obtaining an independent mask for each tooth. Then, for each tooth's mask region, a classification network is used to calculate the probability and type of caries, and a threshold is set to determine whether it is a removable caries tooth (a caries probability greater than or equal to the set threshold is considered a removable caries tooth). After obtaining the caries probability for each tooth, the following image features can be automatically extracted based on the statistical results of the entire mouth: the dental caries index (DMF index), the number of removable caries, the number of deep caries, and the number of proximal caries. Then, the obtained image features are concatenated into a one-dimensional vector in a fixed order; this vector is the oral imaging feature. It should be noted that, to eliminate the dimensional differences between different indicators, the obtained image features can be normalized to ensure that different features are on the same scale, and finally, the normalized image features are concatenated into a vector in a fixed order.
[0043] The DMF index measures the number of decayed, lost, and filled teeth. "Caries" refers to untreated, active caries; "lost" refers to teeth extracted due to caries; and "filled" refers to teeth that have been filled. To more precisely reflect the severity of caries, the DMF index can be expanded to DMFT (Dental Mean Forecast) and DMFS (Dental Surface Mean Forecast). DMFT counts the total number of decayed, lost, and filled teeth in the entire mouth, reflecting the extent of caries. DMFS counts the total number of decayed, lost, and filled tooth surfaces in the entire mouth, using the five measurable surfaces of the teeth (occlusal, buccal, lingual, mesial, and distal), providing a more sensitive reflection of the extent and progression of caries.
[0044] In another example, image features, in addition to the aforementioned statistical metrics, can also include high-dimensional abstract features directly extracted using a deep learning model. For instance, a pre-processed panoramic image can be input into a pre-trained ResNet-50 network, and the feature map before the fully connected layers (e.g., a 2048-dimensional vector) can be used as the image features. This high-dimensional feature can capture the fine structural and textural information of early caries lesions that are difficult for the human eye to recognize, complementing the aforementioned statistical metrics. In practical applications, statistical metrics and high-dimensional abstract features can be concatenated or weighted and fused to balance interpretability and representational ability.
[0045] S104: Feature fusion is performed on gene features and oral imaging features to obtain fused features.
[0046] In this embodiment, since genetic features and oral imaging features originate from data of different modalities and each reflects different dimensions of dental caries risk, information from a single modality is insufficient to comprehensively characterize an individual's overall risk. Therefore, this step integrates the two types of features through feature fusion to form a more comprehensive fused feature.
[0047] Specifically, rule-based fusion or machine learning fusion (multilayer perceptron fusion) can be used to fuse genetic features and oral imaging features. Rule-based fusion typically uses a weighted summation method, which is transparent and highly interpretable, making it suitable for clinical scenarios where interpretability is a high priority. MLP fusion, on the other hand, learns the nonlinear interaction relationships between features based on neural networks and other models, and can achieve higher prediction accuracy when data is abundant.
[0048] In one example, rule fusion employs a weighted summation approach, assigning different weights to genetic features and oral imaging features. Clinicians can customize these weights based on their clinical experience and the specific needs of the assessment scenario. Alternatively, weights can be learned autonomously from the training data using optimization algorithms, such as grid search or Bayesian optimization, targeting the evaluation metrics on the validation set to find the optimal weight combination.
[0049] In another example, the structure of a multilayer perceptron (MLP) is as follows: the input layer receives the spliced gene features and image features, which pass through the first hidden layer (32 neurons), the second hidden layer (16 neurons), and finally through the output layer (1 neuron) to output the fused features.
[0050] In another example, to accurately extract key information from genetic features and oral imaging features that is more relevant to caries risk assessment, strengthen the weight of high-contribution features, weaken the interference of low-relevance features, and make the feature fusion result more closely reflect the patient's actual caries risk status, an attention mechanism can be used to adaptively weight and fuse genetic features and oral imaging features. Specifically: First, the genetic feature vector and the oral imaging feature vector are aligned and standardized to be on the same computational scale. Then, the two types of features are input into an attention weight allocation network. Based on a caries clinical annotation dataset, this network autonomously learns and quantifies the correlation strength between each feature dimension and caries risk, assigning differentiated attention weights to different feature dimensions, giving higher weights to feature dimensions that are highly correlated with caries risk and lower weights to feature dimensions that are not highly correlated. Subsequently, the attention weights are used to weight and enhance the genetic features and oral imaging features respectively. Then, the weighted two types of features are concatenated and fused to obtain the final fused feature.
[0051] In another example, since image features typically have high dimensionality (e.g., image features extracted through deep learning can reach 2048 dimensions), while gene features have low dimensionality (e.g., 13 dimensions, i.e., the 13 gene loci related to caries susceptibility provided in this application), direct splicing will lead to dimensionality mismatch, with high-dimensional image features dominating the calculation results and masking the information of gene features. Therefore, principal component analysis (PCA) can be performed on the image features for dimensionality reduction. By extracting the principal components of the image features, the core information related to caries is preserved while compressing the feature dimensions, making the dimensionality of the dimensionality-reduced image features the same as that of the gene features. Then, rule-based fusion or MLP fusion can be performed between the image features and the gene features.
[0052] S105: Perform caries risk assessment on fusion characteristics to obtain the caries risk assessment results of the caries sample to be assessed.
[0053] In this embodiment, although the fusion feature integrates multidimensional information from genetic and oral imaging features, it remains an abstract feature vector and is difficult to directly apply to clinical decision-making. Therefore, this step maps the fusion feature to a trained classifier or regressor, resulting in a caries risk assessment with clear clinical significance. Since the fusion feature simultaneously contains imaging information reflecting current lesions and genetic information predicting future risks, this application can support the construction of both current caries risk assessment models and future caries risk prediction models.
[0054] In one example, the assessment result can be the current risk of caries, reflecting the severity of caries lesions at the current point in time; or it can be the future risk of caries, predicting the probability or number of new caries developments in an individual within a future period (e.g., 6 months, 1 year, or 2 years). For different assessment objectives, the fusion model can be adapted by setting different modality weights or using different classifiers.
[0055] Specifically, when the assessment result indicates a future risk of developing dental caries, the weight of genetic features is higher than that of oral imaging features during the feature fusion stage because imaging features primarily reflect existing caries lesions, while genetic features reveal innate susceptibility and have a stronger predictive ability for future disease development. Conversely, when the assessment result indicates a current risk of dental caries, the weight of oral imaging features is higher than that of genetic features because imaging features primarily reflect existing caries lesions.
[0056] When using MLP fusion or attention-based fusion mechanisms, parameter optimization can be performed through label-supervised learning to ensure that the output fused features are effective for subsequent risk assessment tasks. Specifically, the fusion network and subsequent classifiers or regressors can be optimized through joint training or staged training.
[0057] In one example, an end-to-end joint training approach is used: the fusion network and the classifier (or regressor) are concatenated to form a complete network, and the labels of the labeled samples are used as supervision signals. The parameters of the fusion network and the classifier are optimized simultaneously through the backpropagation algorithm. If it is a current caries risk model, the training labels can be: whether there are active caries, i.e., binary labels, where 0 indicates no active caries and 1 indicates at least one active caries tooth; or caries severity classification, i.e., multi-class labels, such as 0: no caries, 1: mild caries (1-2 teeth), 2: moderate caries (3-4 teeth), 3: severe caries (≥5 teeth). If it is a future caries risk model, the training labels can be: whether new caries will appear in the future, i.e., binary labels, such as whether at least one new caries tooth will appear in the next 12 months; or the number of new caries teeth in the future, i.e., regression labels, such as the specific number of new caries teeth in the next 12 months.
[0058] In one example, the appropriate classifier can be selected based on the output format of the evaluation results: If a quantified risk probability is required, a logistic regression classifier can be used. After inputting the fused features into the logistic regression classifier, it outputs a continuous probability value between 0 and 1, which directly reflects the patient's risk level of caries. If a multi-level risk probability distribution is needed directly, a Softmax multi-classifier can be used to directly output the risk category. If a quantitative prediction of caries incidence is required, an XGBoost regression model can be used. After inputting the fused features, the XGBoost regression model can directly output the future disease trend. It should be noted that if MLP fusion is used, the classifier can be Sigmoid.
[0059] In this embodiment, compared to the prior art which relies solely on oral imaging data for caries assessment, the insufficient image resolution leads to unclear early pathological features of caries on the images, resulting in inaccurate caries risk assessment results. This application, however, combines genetic data with caries risk assessment. By numerically encoding the genetic data based on risk alleles at multiple gene loci related to caries susceptibility, gene features are obtained. These gene features and oral imaging features are then fused to obtain fused features. This integrates congenital genetic susceptibility information for caries with current oral lesion information, reducing the limitations of single imaging data in representing caries risk and compensating for the shortcomings of imaging data in capturing early caries features. The fused features comprehensively reflect the patient's caries risk from both congenital susceptibility and acquired lesions. Furthermore, caries risk assessment is performed on the fused features to obtain the caries risk assessment results for the caries sample to be assessed, effectively improving the accuracy and comprehensiveness of caries risk assessment.
[0060] Figure 2 A flowchart illustrating the caries risk assessment method provided in the second embodiment of this application is shown. Figure 2As shown, one specific implementation of step S102 is as follows: S201: Obtain genotypic information of multiple gene loci associated with susceptibility to dental caries from genetic data.
[0061] In this embodiment, genotype information refers to the specific base composition of a pair of alleles at a specific gene locus. This information can be obtained based on a gene testing report.
[0062] In one example, the following is the genotypic information for the caries sample to be evaluated. AMELX (SNP number rs17878486): A / G; ENAM (SNP number rs12640848): G / G; MMP20 (SNP number rs1784418): T / T; KLK4 (SNP number rs198968): G / A; DSPP (SNP number rs36094464): C / C; AQP5 (SNP number rs296763): C / T; MUC7 (SNP number rs198968): G / A; DSPP (SNP number rs36094464): C / C; AQP5 (SNP number rs296763): C / T; MUC7 (SNP number rs198968): C / G; MMP20 (SNP number rs1784418): T / T; MMP20 (SNP number rs1784418): T / T; MMP20 (SNP number rs198968): T / T; MMP20 (SNP number rs198968): T / G ... rs1104022): A / A; CA6 (SNP number rs2274327): C / T; LYZ (SNP number rs1800014): G / G; DEFB1 (SNP number rs11362): G / G; TAS1R2 (SNP number rs35874116): C / T; SLC2A2 (SNP number rs5400): G / G; IL1B (SNP number rs1143634): C / T.
[0063] S202: Compare the genotype information of each gene locus with the risk alleles corresponding to that gene locus, and count the number of risk alleles carried in the genotype information.
[0064] In the embodiments of this application, for each caries susceptibility-related gene locus, its corresponding risk allele is first determined, and then the two alleles in the genotype information of the locus are compared with the risk allele one by one to accurately count the number of risk alleles in the genotype information. The statistical result is only 0, 1 or 2, which correspond to three cases: no carrier, heterozygous carrier, and homozygous carrier of risk alleles, respectively.
[0065] S203: The statistically obtained number is used as the coding value of the gene locus.
[0066] In the embodiments of this application, the number of risk alleles carried directly reflects the innate susceptibility to dental caries at that locus. Homozygous carriers have a higher susceptibility than heterozygous carriers, and those who do not carry the alleles have no innate susceptibility risk at that locus. Therefore, directly using the statistical number as the coding value of the gene locus can preserve the innate susceptibility information of the gene locus, avoid information loss during the coding process, and make the coding value directly correspond to the susceptibility level, thereby improving the intuitiveness of gene coding.
[0067] In one example, based on the genotype information of the caries sample to be evaluated, the encoded values are shown in Table 1: Table 1. Genotypic information and gene locus coding values of carious dental samples to be evaluated.
[0068] S204: Combine the coding values of each gene locus into a gene feature.
[0069] In this embodiment of the application, the coding value of each gene locus is used as the feature value of that locus. The coding values of all loci are arranged in a preset order (e.g., ascending order by chromosome position, alphabetical order by gene name, or order by SNP number) and combined into a one-dimensional vector, which is the gene feature. For example, according to the order in Table 1, the gene feature is: [1,0,2,1,0,1,2,1,2,0,1,0,1].
[0070] In one example, for cases where genotypes at certain loci are missing, such as when genotype information for a particular locus cannot be obtained due to limitations in detection technology or sample quality issues, 0 (i.e., risk-free alleles) or the population average coding value can be used to fill the gaps.
[0071] In this embodiment, by obtaining genotype information of multiple gene loci related to susceptibility to dental caries in the gene data, and then counting the number of risk alleles carried in the genotype information, the counted number is used as the encoding value of the gene locus. The encoding values of each gene locus are then combined to form the gene feature, thereby quantifying the gene features of dental caries-related gene data. Furthermore, since the encoding result is directly based on the number of risk alleles carried, the encoding result corresponds to the degree of innate susceptibility to dental caries, preserving the innate susceptibility information of the gene locus and avoiding information loss during the encoding process.
[0072] In one example, to further improve the quantification accuracy of genetic characteristics on dental caries risk and eliminate the interference of differences in the intensity of effects at different gene loci on risk assessment, the following steps are included before step S204: Calculate the risk weight corresponding to the risk allele at each gene locus; wherein the risk weight is used to characterize the association strength between a single risk allele and the risk of dental caries.
[0073] In this embodiment, since the association strength between risk alleles at different gene loci and dental caries varies, the coding value alone cannot reflect the difference in the association strength with the onset of dental caries. Therefore, in order to further explore the deeper risk information of gene data, make gene features more consistent with the actual dental caries pathogenesis, and improve the accuracy of subsequent fusion and assessment, it is necessary to calculate the risk weight corresponding to the risk allele at each gene locus to quantify the importance of gene data at different loci to dental caries assessment, so that high-association-strength loci can play a more core role in gene features.
[0074] The coding values and risk weights of each gene locus are combined to form the gene feature.
[0075] In this embodiment of the application, the coding value of each gene locus and the risk weight can be multiplied by vector dot product to obtain the weighted coding value of each gene locus. The weighted coding value reflects both the number of risk alleles and the strength of the association between the locus and dental caries, making the feature value of a single locus more biologically meaningful and clinically valuable. Then, all weighted contribution values are combined in a preset order to form the weighted gene feature.
[0076] In another example, in order to preserve the original information of the encoded values and risk weights, and to allow the subsequent fusion model to learn the synergistic relationship between the two autonomously, avoiding information compression that may be caused by dot product operations, the risk weights corresponding to the risk alleles in each gene locus can be arranged in a predetermined gene locus order to form a risk weight vector. Then, the risk weight vector and the encoded value vector are concatenated to form a higher-dimensional gene feature vector. This allows the model to autonomously explore the interaction patterns between the encoded values and risk weights during training, further improving the gene feature expression ability.
[0077] In this embodiment, since the degree of influence of risk alleles at different gene loci on dental caries varies fundamentally, the coding value alone cannot distinguish this difference. In order to make the gene features more accurately and comprehensively reflect the patient's congenital susceptibility to dental caries, the risk weight corresponding to the risk allele at each gene locus is calculated, and the coding value and the risk weight of each gene locus are combined into the gene feature. This allows the gene feature to simultaneously integrate the number of risk alleles carried and the information on the strength of the locus association, which can more realistically reflect the patient's level of congenital dental caries susceptibility, thereby improving the accuracy of dental caries risk assessment.
[0078] In one example, a specific implementation for calculating the risk weight corresponding to the risk allele at each gene locus is as follows: Obtain sample data from the exposed group and the control group corresponding to the risk allele, wherein the exposed group sample data includes dental caries prevalence information of a first population carrying the risk allele, and the control group sample data includes dental caries prevalence information of a second population not carrying the risk allele.
[0079] In this embodiment of the application, the exposed group sample data is a set of dental caries prevalence information of a first population carrying a certain risk allele, including the total number of people in the population, the number of people with dental caries, and the number of people without dental caries; the control group sample data is a set of dental caries prevalence information of a second population not carrying the risk allele, including the total number of people in the population, the number of people with dental caries, and the number of people without dental caries.
[0080] In one example, to eliminate interference from other factors and ensure the validity of the comparative analysis of the two groups of data, the age distribution, gender ratio, oral hygiene habits, geographical environment, and dietary structure of the first and second groups need to be consistent. This is to eliminate the influence of non-genetic factors on the caries morbidity outcome and ensure that the risk weights calculated subsequently are determined only by the pathogenicity differences of the risk alleles themselves, thereby improving the accuracy and reliability of the weight calculation results.
[0081] The disease advantage of the exposed group was calculated based on the sample data of the exposed group, and the disease advantage of the control group was calculated based on the sample data of the control group. In this embodiment, the disease advantage in the exposed group refers to the ratio of the number of caries-affected patients to the number of caries-free patients in the first population; the disease advantage in the control group refers to the ratio of the number of caries-affected patients to the number of caries-free patients in the second population. The prevalence index of the number of caries-affected patients / total population is easily affected by population size, while the disease advantage can more accurately reflect the relative probability of disease in the population and is not affected by population size. By calculating the disease advantage of the two groups, the caries prevalence of the two groups can be transformed into a directly comparable quantitative indicator.
[0082] The risk weight corresponding to the risk allele is obtained by dividing the disease advantage of the exposed group by the disease advantage of the control group.
[0083] In this embodiment, the disease advantage ratio between the exposed group and the control group can directly quantify the influence of risk alleles on the risk of dental caries. The calculated disease advantage of the exposed group is used as the numerator, and the disease advantage of the control group is used as the denominator. The division is performed, and the quotient is the risk weight corresponding to the risk allele. In this application, the risk weight is the OR value (odds ratio). An OR value > 1 indicates that the risk allele is a risk factor for dental caries. The larger the OR value, the stronger the association.
[0084] In one example, the risk weights corresponding to risk alleles at each gene locus can be obtained based on publicly available GWAS research literature. Specifically: AMELX (SNP number rs17878486): OR value 1.8; ENAM (SNP number rs12640848): OR value 1.5; MMP20 (SNP number rs1784418): OR value 1.6; KLK4 (SNP number rs198968): OR value 1.4; DSPP (SNP number rs36094464): OR value 1.4; AQP5 (SNP number rs296763): OR value 1.5; MUC7 (SNP number rs198968 ...UC7 (SNP number rs198968): OR value 1.5; MUC7 (SNP number rs198968): OR value 1.5; MUC7 (SNP number rs198968): OR value 1.5; MUC7 (SNP number rs198968): OR value 1.5; MUC7 (SNP number rs198968): OR value 1.5; MUC7 (SNP number rs198968): OR value 1.5; The OR value for SNP 1104022 was 1.6; for CA6 (SNP 2274327): 1.4; for LYZ (SNP 1800014): 1.3; for DEFB1 (SNP 11362): 1.7; for TAS1R2 (SNP 35874116): 1.5; for SLC2A2 (SNP 5400): 1.3; and for IL1B (SNP 1143634): 1.6.
[0085] In one example, the OR value may differ among different regions and ethnic groups. Therefore, the OR value can be calculated separately for different regions and ethnic groups to make the genetic characteristics more closely match the genetic characteristics of the target population, thereby improving the application effect of the risk assessment model in specific populations.
[0086] In this embodiment, the disease prevalence of the exposed group is calculated based on the sample data of the exposed group, and the disease prevalence of the control group is calculated based on the sample data of the control group. The disease prevalence of the exposed group is divided by the disease prevalence of the control group to obtain the risk weight corresponding to the risk allele. This can intuitively reflect how many times the individual carrying a certain risk allele has the advantage of developing dental caries compared to non-carriers. In turn, it can intuitively reflect the association strength between a single risk allele and the risk of dental caries, providing more realistic gene feature input for subsequent multimodal fusion, thereby improving the accuracy of dental caries risk assessment.
[0087] In another example, the genetic susceptibility to dental caries is not determined independently by a single gene locus, but rather by the combined effects and interactions of multiple gene loci. For instance, enamel development-related genes (such as AMELX) and matrix metalloproteinase genes (such as MMP20) both participate in the formation of tooth enamel. If both loci simultaneously carry risk alleles, it may lead to the superposition of enamel structural defects. Therefore, to further explore the associations and interactions between different gene loci and improve the comprehensiveness of gene characteristics, this application also includes the following embodiments: After obtaining the coding values (0, 1, 2) for each gene locus, the interaction terms between any two gene loci are further calculated. The interaction term can be calculated as follows: if both loci carry at least one risk allele, the interaction term is set to 1; otherwise, it is 0. All calculated interaction terms are concatenated or combined with the original coding values to form an expanded gene feature vector. This expanded gene feature vector contains both the independent effects of each gene locus and the joint effects between loci, providing richer genetic information for subsequent fusion models.
[0088] This allows obtaining sample data from the exposed group and control group corresponding to any two gene loci. The exposed group sample data includes dental caries prevalence information for a third population where at least one of the two gene loci carries a risk allele; specifically, this may include the total number of people in this population, the number of people with dental caries, and the number of people without dental caries. The control group sample data includes dental caries prevalence information for a fourth population where at least one of the two gene loci does not carry a risk allele; specifically, this may include the total number of people in this population, the number of people with dental caries, and the number of people without dental caries. Then, the disease prevalence advantage corresponding to the exposed group sample data and the disease prevalence advantage corresponding to the control group sample data are calculated. Dividing the exposed group disease prevalence advantage by the control group disease prevalence advantage yields the risk weight corresponding to the interaction term. Therefore, the risk weight and the encoded value of the interaction term can be combined into a gene feature and concatenated with the gene feature corresponding to the original single gene locus.
[0089] Figure 3 A flowchart illustrating the caries risk assessment method provided in the third embodiment of this application is shown. Figure 3 As shown, prior to step S104, the method may further include the following steps: S301: Obtain oral health record data and lifestyle habit data, extract features from the oral health record data and lifestyle habit data, and obtain oral health record features and lifestyle habit data features.
[0090] In this embodiment, to more accurately assess the risk of dental caries, oral health record data and lifestyle data are further introduced in addition to genetic data and oral imaging data to more comprehensively reflect the multidimensional factors influencing the occurrence of dental caries. Dental caries is a multifactorial disease; besides congenital genetic susceptibility and existing structural lesions, a patient's past oral history, oral microecological status, and daily habits all have a significant impact on the occurrence and development of dental caries. Therefore, this application, by acquiring and extracting features from these two types of data, enables the risk assessment to comprehensively consider four dimensions: "congenital genetics," "acquired lesions," "clinical status," and "behavioral factors," thereby improving the comprehensiveness and accuracy of the assessment.
[0091] It should be noted that oral health record data is a collection of clinical data obtained based on the patient's oral clinical examination results. This data may include: dental caries history, saliva parameters, plaque index, and cariogenic bacteria detection results. The oral health record data can be past and present time-series data. By extracting trend features from the series, it can more accurately reflect the dynamic patterns of dental caries progression in patients, providing richer information for future risk prediction. Lifestyle data is obtained based on standardized questionnaires and may include: frequency of sugary food intake, brushing habits, systemic diseases, etc., reflecting the acquired behavioral characteristics of patients that influence the occurrence of dental caries.
[0092] S302: Non-image fusion features are obtained by splicing features from oral health record features, lifestyle data features, and genetic features.
[0093] In this embodiment, since the oral health record features, lifestyle data features, and genetic features are all non-image features, in order to improve the efficiency of subsequent multimodal fusion, simplify the model structure, and avoid increasing computational complexity due to the introduction of too many modal branches too early, this embodiment adopts an early fusion strategy, which splices the oral health record features, lifestyle data features, and genetic features to obtain non-image fusion features.
[0094] For example, after extracting and encoding the features of oral health records, the following feature vectors are obtained: caries history feature values, saliva index feature values, plaque index feature values, and cariogenic bacteria detection result feature values, totaling 5 dimensions. After extracting and encoding the lifestyle data features, the following features are obtained: sugary food frequency feature values, brushing habit feature values, and systemic disease feature values, totaling 3 dimensions. Genetic features are obtained using additive encoding of the aforementioned 13 caries-related gene loci, with a dimension of 13. Therefore, after concatenating these three types of features, a non-image fusion feature vector of 5 + 3 + 13 = 21 dimensions can be obtained.
[0095] S303: Perform feature fusion between non-image fusion features and oral image features to obtain fused features.
[0096] In this embodiment, late-stage fusion involves performing cross-modal deep fusion of structured non-image fusion features and oral imaging features to obtain the final fused feature. Late-stage fusion can employ weighted fusion, MLP fusion, or attention fusion strategies.
[0097] In this embodiment, to improve the accuracy of caries risk assessment, oral health record data and lifestyle data are acquired, and features are extracted from these data to obtain oral health record features and lifestyle data features. This increases the feature dimensions of caries risk assessment, supplementing core information reflecting the patient's oral clinical health status, past medical history, and acquired behavioral risk factors, thus achieving comprehensive coverage of caries causative factors. Furthermore, to achieve the fusion of different types of features and avoid fusion chaos caused by feature modal heterogeneity and complex dimensions, and to improve the efficiency of subsequent multimodal fusion, this application adopts a hierarchical fusion strategy combining early and late fusion. First, the oral health record features, lifestyle data features, and genetic features are spliced together to obtain non-image fusion features. Then, the non-image fusion features and the oral imaging features are fused to obtain the fused features.
[0098] In one example, prior to step S503, the following is also included: The oral cavity image data is subjected to image enhancement processing to obtain enhanced oral cavity image data, and the oral cavity image features are extracted from the enhanced oral cavity image data; the image enhancement processing includes at least one of geometric transformation and pixel-level transformation.
[0099] In this embodiment of the application, the original oral imaging data is easily affected by factors such as device perspective, patient shooting posture, soft tissue obstruction in the oral cavity, and device noise during the acquisition process, resulting in low image feature recognition and large model recognition error. Therefore, in order to improve the quality of oral imaging data, enhance the recognition and integrity of lesion features in the images, reduce the influence of various interference factors during the acquisition process, and provide high-quality input data for subsequent image feature extraction, it is necessary to perform image enhancement processing on the oral imaging data to obtain enhanced oral imaging data.
[0100] Geometric transformation refers to adjusting the spatial structure of an image, changing the position, orientation, size, or shape of objects within the image without altering its content information. Geometric transformations can simulate various spatial variations that may occur during clinical imaging, enhancing the model's adaptability to differences in patient head position, shooting angle, and tooth alignment. Specifically, it can include at least one of the following: horizontal flipping, vertical flipping, random 90-degree rotation, affine transformation, elastic transformation, and mesh distortion.
[0101] Pixel-level transformation refers to adjusting the pixel values of an image to change its color, brightness, texture, or add noise without altering its spatial structure. Pixel-level transformation can simulate factors that may occur during clinical imaging, such as exposure differences, X-ray dose variations, and sensor noise, enhancing the model's adaptability to fluctuations in image quality. Specifically, it can include at least one of Gaussian blur, Gaussian noise, and coarse-grained dropout.
[0102] During the training phase, one or more combinations of the aforementioned geometric and pixel-level transformations can be randomly applied to the oral cavity image data. For example, for each original oral cavity image data, 3-5 transformations can be randomly selected and combined with a certain probability during each training iteration, such as "horizontal flip + Gaussian blur + elastic transformation + coarse-grained discarding," to generate enhanced image data for model training. This online enhancement strategy allows the model to learn more diverse image variants during training, effectively expanding the diversity of training samples and preventing overfitting.
[0103] The oral health record features, lifestyle data features, and genetic features are subjected to non-image feature enhancement processing to obtain enhanced non-image features; the non-image feature enhancement processing includes at least one of the following: standardization processing, normalization processing, nonlinear transformation, and categorical feature enhancement.
[0104] In this embodiment, the oral health record features, lifestyle data features, and genetic features are derived from clinical examinations, questionnaires, and gene testing, respectively. These data vary in form, scale, and distribution characteristics. If these raw features are directly input into the fusion model, the differences in feature scale will lead to slow gradient descent convergence, skewed distribution will affect the model's fitting performance, and categorical variables cannot directly participate in numerical calculations. Therefore, to eliminate dimensional differences, address skewed distribution, convert categorical variables into numerical forms, and improve the expressive power of non-image features and the stability of model training, it is necessary to enhance the non-image features to obtain enhanced non-image features.
[0105] Standardization refers to converting features into a distribution with a mean of 0 and a standard deviation of 1 to eliminate the influence of different feature dimensions. For example, after standardization, the saliva secretion volume in oral health records and the frequency of sweets intake in lifestyle data can be unified in dimension, which facilitates subsequent feature splicing and fusion calculation.
[0106] Normalization refers to scaling features to a fixed range to reduce the interference of extreme outliers on feature representation and improve the convergence speed of model training.
[0107] Nonlinear transformations adjust the distribution of features through mathematical transformations, strengthening the nonlinear association between features and dental caries risk. These transformations can include at least one of the following: logarithmic transformation, square root transformation, and binning. Logarithmic transformation takes the logarithm of non-image feature values, compressing the numerical range of features, mitigating the impact of extreme values, and enhancing the discriminative power of low-value features. For example, cariogenic bacteria counts typically follow a log-normal distribution; after logarithmic transformation, this distribution becomes more symmetrical, allowing the model to better capture its association with dental caries. Square root transformation takes the square root of non-image feature values, moderately smoothing feature value fluctuations, improving the stability of feature data, and preserving the core distribution patterns of features. For example, for the number of times a sugary food is consumed daily, square root transformation can reduce the impact of extreme values while preserving discriminative power. Binning divides continuous non-image features into multiple discrete intervals, converting continuous values into discrete categories, reducing the impact of outliers and strengthening the interval correlation of features.
[0108] Category feature enhancement involves encoding and transforming discrete category features in non-image features, converting textual or categorical features into computable numerical features. This can include at least one of one-hot encoding, ordinal encoding, and embedding. One-hot encoding avoids interference from the numerical magnitude of category features, ordinal encoding preserves the hierarchical relationship between categories, and embedding maps category features to low-dimensional dense vectors, improving the expressive power of category features and adapting to the computational needs of subsequent models.
[0109] In this embodiment, enhanced oral imaging data is obtained by performing image enhancement processing on oral imaging data. Oral imaging features are extracted from the enhanced oral imaging data. Non-image feature enhancement processing is performed on the oral health record features, lifestyle data features, and genetic features to obtain enhanced non-image features. This can improve the quality and expressive power of various features, eliminate the influence of various interference factors during the collection and generation process, provide high-quality feature input for subsequent multimodal feature fusion and caries risk assessment, and thus improve the accuracy of caries risk assessment.
[0110] Figure 4 A flowchart illustrating the caries risk assessment method provided in the third embodiment of this application is shown. Figure 3 As shown, prior to step S104, the method may further include the following steps: S401: Obtain follow-up data for the caries sample to be evaluated. The follow-up data includes at least one of the following: updated oral imaging data, updated oral health record data, and updated lifestyle data.
[0111] In this embodiment, follow-up data refers to multi-source data reflecting changes in the patient's condition collected again after obtaining the caries risk assessment results for the caries sample to be evaluated and after providing early intervention recommendations to the patient. By introducing follow-up data, this application can achieve dynamic monitoring of caries risk and iterative updates of assessment results, enabling the risk assessment model to adjust its output in real time according to changes in the patient's condition, thereby providing continuous and accurate support for clinical decision-making. The follow-up data includes at least one of updated oral imaging data, updated oral health record data, and updated lifestyle habit data.
[0112] In one example, early intervention recommendations include at least one of the following: oral hygiene instructions, dietary adjustment plans, and reminders for regular check-ups.
[0113] S402: Extract features from gene data and follow-up data to obtain updated multimodal features.
[0114] In this embodiment, since genetic data is determined by an individual's innate genetic information and typically does not change throughout their lifespan, the genetic characteristics from the initial assessment can be directly reused when updating follow-up data, without the need for re-detection or extraction. However, updated oral imaging data, updated oral health record data, and updated lifestyle data in the follow-up data need to be processed using the same feature extraction method as the initial risk assessment to ensure consistency and comparability of the features.
[0115] S403: Perform feature fusion on the updated multimodal features to obtain the updated fused features.
[0116] In this embodiment, to ensure that the updated risk assessment results are consistent with the initial assessment results in terms of scoring scale and meaning, facilitating before-and-after comparisons and effectiveness evaluations, the same fusion strategy and model parameters as the initial assessment are used when fusing the updated multimodal features. That is, if the initial assessment used rule-based fusion, the same weight configuration is used in the follow-up update; if the initial assessment used MLP fusion or attention fusion, the same network structure and trained model parameters are used in the follow-up update. By maintaining consistency in the fusion strategy, the updated fused features and the initial fused features reside in the same feature space, making the risk scores output in the two assessments directly comparable and able to intuitively reflect the changing trend of the patient's risk status and the effectiveness of intervention measures.
[0117] S404: Perform caries risk assessment on the updated fusion features to obtain the updated caries risk assessment results for the caries sample to be evaluated.
[0118] In this embodiment, the updated fusion features are used to assess the risk of dental caries, employing the same classifier or regressor as the initial assessment to output the updated risk assessment results. These updated results objectively reflect the effectiveness of the intervention recommendations: if the updated risk score is lower than the initial assessment, it indicates that the intervention is effective and the patient's risk is controlled; if the risk score shows little change or even increases, it suggests that the current intervention plan may be inapplicable, requiring adjustment of the intervention strategy or enhanced follow-up.
[0119] In one example, a follow-up sequence can be obtained based on the follow-up period. That is, multiple follow-ups are conducted on the same patient to form a time-series follow-up data set. Based on the follow-up sequence, the changing trend of the patient's caries risk can be dynamically tracked, the impact of intervention measures on the patient's caries risk can be accurately captured, and the future caries incidence trend of the patient can be predicted in advance, so as to realize dynamic monitoring and prospective intervention of caries risk.
[0120] For example, the follow-up sequence containing the time dimension can be input into a time-series feature extraction model such as a recurrent neural network, a long short-term memory network, or a gated recurrent unit. By leveraging the model's ability to remember and associate time-series information, the evolution of the patient's risk characteristics at different follow-up time points can be learned, thereby more accurately predicting future caries risk assessment results.
[0121] In another example, after obtaining new follow-up data and updated risk assessment results, the follow-up data, updated multimodal features, and corresponding updated risk assessment results can be used as new training samples to supplement the original training set, retrain the caries risk assessment model, thereby optimizing the model parameters, improving the model's adaptability to dynamic risk changes in patients and the accuracy of assessment, allowing the model to better adapt to the post-intervention risk change patterns of different patients, and further improving the accuracy of subsequent risk assessments.
[0122] In this embodiment of the application, after assessing the risk of the fusion features to obtain the caries risk score of the caries sample to be assessed, the follow-up data of the caries sample to be assessed is obtained, the gene data and the follow-up data are used to extract features to obtain updated multimodal features, and the updated multimodal features are fused to obtain updated fusion features. Thus, caries risk assessment is performed based on the updated fusion features to obtain the updated caries risk assessment result of the caries sample to be assessed, which can realize dynamic monitoring and continuous assessment of caries risk.
[0123] Based on the caries risk assessment method provided in the above embodiments, this application also provides specific implementation methods for the caries risk assessment device. Please refer to the following embodiments.
[0124] Figure 5 This application provides a schematic diagram of the structure of a caries risk assessment device according to an embodiment of the present application. Figure 5 As shown, the caries risk assessment device 500 includes: The acquisition module 501 is used to acquire multi-source data corresponding to the caries sample to be evaluated. The multi-source data includes at least: oral imaging data and genetic data. The encoding module 502 is used to numerically encode gene data based on the risk alleles of multiple gene loci associated with susceptibility to dental caries, thereby obtaining gene characteristics. The feature extraction module 503 is used to extract features from oral imaging data to obtain oral imaging features; The feature fusion module 504 is used to fuse gene features and oral imaging features to obtain fused features; Risk assessment module 505 is used to assess the caries risk of fusion features and obtain the caries risk assessment results of the caries sample to be assessed.
[0125] In one example, encoding module 502 includes: The first acquisition submodule is used to acquire genotype information of multiple gene loci related to susceptibility to dental caries in the gene data; The statistics submodule is used to compare the genotype information of each gene locus with the risk alleles corresponding to that gene locus, and to count the number of risk alleles carried in the genotype information. The first processing submodule is used to take the statistically obtained quantity as the encoding value of the gene locus; The second processing submodule is used to combine the coding values of each gene locus into gene features.
[0126] In one example, encoding module 502 includes: The first calculation submodule is used to calculate the risk weight corresponding to the risk allele at each gene locus; wherein, the risk weight is used to characterize the association strength between a single risk allele and the risk of dental caries. The third processing submodule is used to combine the coding values and risk weights of each gene locus into gene features.
[0127] In one example, encoding module 502 includes: The second acquisition submodule is used to acquire the exposed group sample data and the control group sample data corresponding to the risk allele. The exposed group sample data includes the caries prevalence information of the first population carrying the risk allele, and the control group sample data includes the caries prevalence information of the second population not carrying the risk allele. The second calculation submodule is used to calculate the disease advantage of the exposed group based on the sample data of the exposed group, and to calculate the disease advantage of the control group based on the sample data of the control group. The third calculation submodule is used to divide the disease advantage of the exposed group by the disease advantage of the control group to obtain the risk weight corresponding to the risk allele.
[0128] In one example, the caries risk assessment device 500 also includes: The processing module is used to acquire oral health record data and lifestyle data, extract features from the oral health record data and lifestyle data, and obtain oral health record features and lifestyle data features. The feature stitching module is used to stitch together features from oral health records, lifestyle data, and genetic features to obtain non-image fusion features. The feature fusion module is also used to fuse non-image fusion features and oral image features to obtain fused features.
[0129] In one example, the caries risk assessment device 500 also includes: The image enhancement module is used to perform image enhancement processing on oral image data to obtain enhanced oral image data, and to extract oral image features from the enhanced oral image data; the image enhancement processing includes at least one of geometric transformation and pixel-level transformation.
[0130] The non-image feature enhancement module is used to perform non-image feature enhancement processing on oral health record features, lifestyle data features, and genetic features respectively to obtain enhanced non-image features. The non-image feature enhancement processing includes at least one of the following: standardization processing, normalization processing, nonlinear transformation, and categorical feature enhancement.
[0131] In one example, the caries risk assessment device 500 also includes: The acquisition module is also used to acquire follow-up data of the caries sample to be evaluated. The follow-up data includes at least one of the following: updated oral imaging data, updated oral health record data, and updated lifestyle data. The feature extraction module is also used to extract features from gene data and follow-up data to obtain updated multimodal features; The feature fusion module is also used to fuse the updated multimodal features to obtain the updated fused features; The risk assessment module is also used to assess the caries risk of the updated fusion features, and obtain the caries risk assessment results of the caries sample to be assessed after the update.
[0132] Figure 6 A schematic diagram of the hardware structure of the electronic device provided in this application is shown. The electronic device 600 includes a processor 601 and a memory 602 storing computer program instructions.
[0133] Specifically, the processor 601 may include a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0134] Memory 602 may include mass storage for data or instructions. For example, and not limitingly, memory 602 may include a hard disk drive (HDD), a floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. In one instance, memory 602 may include removable or non-removable (or fixed) media, or memory 602 may be a non-volatile solid-state memory.
[0135] In one instance, memory 602 may be read-only memory (ROM). In one instance, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or flash memory, or a combination of two or more of these.
[0136] Memory 602 may include read-only memory (ROM), random access memory (RAM), disk storage media device, optical storage media device, flash memory device, electrical, optical, or other physical / tangible memory storage device. Therefore, generally, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to one aspect of this application.
[0137] The processor 601 implements a caries risk assessment method in the above-described embodiment by reading and executing computer program instructions stored in the memory 602.
[0138] In one example, the electronic device may also include a communication interface 603 and a bus 604. Wherein, as... Figure 6 As shown, the processor 601, memory 602, and communication interface 603 are connected through bus 604 and complete communication with each other.
[0139] The communication interface 603 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0140] Bus 604 includes hardware, software, or both, that couples components of an online data traffic metering device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 604 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, this application contemplates any suitable bus or interconnect.
[0141] Furthermore, in conjunction with the caries risk assessment method in the above embodiments, this application embodiment can provide a computer-readable storage medium for implementation. The computer-readable storage medium stores computer program instructions; when executed by a processor, these computer program instructions implement any of the caries risk assessment methods in the above embodiments.
[0142] This application also provides a computer program product, including a computer program, which, when executed, implements any of the caries risk assessment methods described in the above embodiments.
[0143] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0144] The functional blocks shown in the above block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, read-only memory (ROM), flash memory, erasable read-only memory (EROM), floppy disks, compact disc read-only memory (CD-ROM), optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0145] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0146] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.
[0147] The above are merely specific embodiments of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and servers described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.
Claims
1. A method for assessing the risk of dental caries, characterized in that, include: Obtain multi-source data corresponding to the caries sample to be evaluated, wherein the multi-source data includes at least: oral imaging data and genetic data; Based on the risk alleles of multiple gene loci associated with susceptibility to dental caries, the gene data is numerically encoded to obtain gene characteristics; Feature extraction is performed on the oral imaging data to obtain oral imaging features; The genetic features and the oral imaging features are fused to obtain fused features; The caries risk is assessed based on the fusion features to obtain the caries risk assessment results of the caries sample to be assessed.
2. The method according to claim 1, characterized in that, The genetic characteristics are obtained by numerically encoding the gene data based on risk alleles at multiple gene loci associated with susceptibility to dental caries, including: Obtain genotype information for multiple gene loci associated with susceptibility to dental caries from the gene data; The genotype information of each gene locus is compared with the risk allele corresponding to that gene locus, and the number of risk alleles carried in the genotype information is counted. The statistically obtained number is used as the coding value of the gene locus; The coding values of each gene locus are combined to form the gene feature.
3. The method according to claim 2, characterized in that, Before combining the coding values of each gene locus into the gene feature, the method further includes: Calculate the risk weight corresponding to the risk allele at each gene locus; wherein, the risk weight is used to characterize the association strength between a single risk allele and the risk of dental caries. The coding values and risk weights of each gene locus are combined to form the gene feature.
4. The method according to claim 3, characterized in that, The calculation of the risk weight corresponding to the risk allele at each gene locus includes: Obtain sample data from the exposed group and the control group corresponding to the risk allele, wherein the exposed group sample data includes dental caries prevalence information of a first population carrying the risk allele, and the control group sample data includes dental caries prevalence information of a second population not carrying the risk allele; The disease advantage of the exposed group was calculated based on the sample data of the exposed group, and the disease advantage of the control group was calculated based on the sample data of the control group. The risk weight corresponding to the risk allele is obtained by dividing the disease advantage of the exposed group by the disease advantage of the control group.
5. The method according to claim 1, characterized in that, Before fusing the genetic features and the oral imaging features to obtain the fused features, the method further includes: Acquire oral health record data and lifestyle habit data, and extract features from the oral health record data and lifestyle habit data to obtain oral health record features and lifestyle habit data features; The oral health record features, the lifestyle data features, and the genetic features are spliced together to obtain non-image fusion features; The non-image fusion features and the oral cavity image features are fused to obtain the fused features.
6. The method according to claim 5, characterized in that, Before performing feature fusion on multimodal features to obtain fused features, the following steps are also included: The oral cavity image data is subjected to image enhancement processing to obtain enhanced oral cavity image data, and the oral cavity image features are extracted from the enhanced oral cavity image data; the image enhancement processing includes at least one of geometric transformation and pixel-level transformation; The oral health record features, lifestyle data features, and genetic features are subjected to non-image feature enhancement processing to obtain enhanced non-image features; the non-image feature enhancement processing includes at least one of the following: standardization processing, normalization processing, nonlinear transformation, and categorical feature enhancement.
7. The method according to claim 5, characterized in that, After assessing the risk of the fusion features to obtain the caries risk score of the caries sample to be evaluated, the method further includes: Obtain follow-up data of the caries sample to be evaluated, the follow-up data including at least one of the following: updated oral imaging data, updated oral health record data, and updated lifestyle data; Feature extraction is performed on the gene data and the follow-up data to obtain updated multimodal features; The updated multimodal features are fused to obtain the updated fused features; The caries risk is assessed based on the updated fusion features to obtain the updated caries risk assessment results for the caries sample to be evaluated.
8. The method according to claim 1, characterized in that, The multiple gene loci associated with susceptibility to dental caries include at least one of the following: a gene locus encoding enamel matrix proteins, a gene locus encoding salivary proteins, and a gene locus encoding immunomodulatory factors and metabolic receptors.
9. An electronic device, characterized in that, The device includes: a processor and a memory storing computer program instructions; the processor, when executing the computer program instructions, implements a caries risk assessment method as described in any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that, when executed by a processor, implement a caries risk assessment method as described in any one of claims 1-8.