Old people oral health behavior monitoring method and system based on multi-source data fusion

By using a multi-source data fusion method for monitoring oral health in the elderly, this method identifies the oral health status of the elderly using multimodal data, performs health index compensation and adjustment and spatiotemporal alignment, and establishes an assessment model. This solves the problem of low accuracy in predicting oral disease risk in existing technologies, and enables personalized health management and efficient oral health management.

CN120954730APending Publication Date: 2025-11-14FOSHAN UNIVERSITY

Patent Information

Application Number
CN202511336198.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing oral health monitoring systems for the elderly mainly rely on user-reported data or limited oral examination data, lacking the integration of multi-dimensional and multi-source data, resulting in low accuracy in predicting oral disease risks and a lack of personalized health management recommendations.

Method used

By acquiring multimodal data of the elderly, including lifestyle data, oral image data, and time-series data of physiological indicators, and combining machine learning and deep learning models, multi-source data fusion is performed, including preliminary identification of health indices, compensation adjustment, spatiotemporal alignment, and modal data fusion processing. An assessment model is then established to predict oral disease risk and provide personalized behavior adjustment strategies.

Benefits of technology

It improves the accuracy of oral health risk prediction for the elderly, provides personalized health management advice, reduces the probability of oral diseases, and ensures the integrity of data and the accuracy of analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954730A_ABST
    Figure CN120954730A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of oral health, and discloses an old people oral health behavior monitoring method and system based on multi-source data fusion. The method comprises the following steps: acquiring multi-modal data of a plurality of old people in a historical time period, preliminarily identifying oral cavity states of the old people based on oral cavity image data, and obtaining health indexes; adjusting the health index based on different age ranges, and if the adjusted health index is smaller than a first threshold value, performing time-space alignment association on the multi-modal data; judging whether the modal data in the historical time period has missing modals or not, performing fusion processing on the same modal data in the historical time period based on a judgment result to obtain preliminary fusion vectors, and fusing the preliminary fusion vectors of different modals to obtain a comprehensive feature vector; and establishing an evaluation model to analyze the comprehensive feature vector, and outputting the risk probability and the disease type of the oral disease of the elderly user in the future. According to the invention, the risk prediction precision of oral health of old people is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of oral health technology, and in particular to a method and system for monitoring oral health behavior of the elderly based on multi-source data fusion. Background Technology

[0002] The incidence of oral diseases is high among the elderly due to the decline in their physiological functions, reduced oral self-care abilities, and lack of knowledge about oral health.

[0003] Existing patent applications, such as Chinese patent application CN119673426A, disclose a survey model reflecting oral health and a method for constructing the model, including an information collection module, an analysis module, an evaluation module, and a customization module. The information collection module is used to collect users' oral health data and transmit it to the survey model. The analysis module is used to analyze factors affecting oral health based on the information collected by the information collection module. The evaluation module is used to evaluate the user's oral health based on the information collected by the information collection module.

[0004] For example, Chinese patent application CN120220932A discloses a self-assessment method and system for oral function decline in the elderly, specifically involving the field of medical and health technology. It includes a login unit, an assessment questionnaire unit, and an intervention guidance unit, and includes the following steps: B1, user registration and login; B2, dynamic logical interactive assessment; B3, multi-dimensional classification assessment; B4, risk grading and intervention generation; B5, data synchronization and remote connection. Through the above system and method, the elderly can complete self-assessment at home using a mobile app, and can monitor the progression of oral deterioration anytime and anywhere.

[0005] However, the two existing technologies mentioned above mainly rely on user self-reported data or limited oral examination data, lacking the integration of multi-dimensional and multi-source data. Therefore, this application proposes a method and system for monitoring oral health behavior of the elderly based on multi-source data fusion. Summary of the Invention

[0006] To address the aforementioned technical issues, this application provides a method and system for monitoring oral health behaviors in the elderly based on multi-source data fusion, which aims to improve the accuracy of risk prediction for oral health in the elderly.

[0007] Firstly, this application provides a method for monitoring oral health behaviors of the elderly based on multi-source data fusion, the method comprising: Step S1: Obtain multimodal data of multiple elderly people over a historical time period. The multimodal data includes lifestyle behavior data, oral image data, and time series data of multiple physiological indicators. Based on the oral image data, the oral condition of the elderly people is initially identified and a health index is obtained. Step S2: Obtain the age range of the elderly, and adjust the health index based on different age ranges. If the adjusted health index is less than the first threshold, perform spatiotemporal alignment association on the multimodal data. Step S3: Determine whether there are missing modalities in the modal data of multiple sub-time periods within the historical time period. Based on the determination result, perform fusion processing on the same modal data within the historical time period to obtain a preliminary fusion vector. Then, fuse the preliminary fusion vectors of different modalities to obtain a comprehensive feature vector. Step S4: Establish an evaluation model. Input the comprehensive feature vector into the evaluation model. The evaluation model performs comprehensive analysis on multiple modal data, outputs the risk probability and disease type of oral diseases that elderly users will experience in the future, sets behavioral adjustment strategies for each disease type, and guides the elderly to improve their oral health.

[0008] In conjunction with the first aspect, in the first implementation of the first aspect of this application, the initial identification of the oral health status of the elderly and the acquisition of a health index include: A first label is assigned to each oral image data. The first label indicates whether the corresponding elderly user has oral problems and the type of problems. A target processing model is selected based on the first label. The target processing model performs image processing on the oral image data to obtain first image data. A recognition model is created based on a machine learning model. The first label, the first image data, and the oral image data are used as training data for the recognition model. The recognition model identifies the oral condition of the elderly user. The oral condition includes whether there are oral problems and the probability of having oral problems. The probability of having oral problems is converted into a health index and output.

[0009] In conjunction with the first aspect, in the second implementation of the first aspect of this application, selecting the target processing model based on the first label includes: The oral cavity image data is input into multiple different machine learning models. Each machine learning model uses a different image processing method to process the first image data and outputs a corresponding feature vector. The image processing method includes grayscale processing, edge detection processing, and color conversion processing. The matching degree between the feature vector output by each machine learning model and the corresponding first label is calculated. The machine learning model corresponding to the matching degree being greater than a first threshold is defined as the target processing model.

[0010] In conjunction with the first aspect, in the third implementation of the first aspect of this application, the health index is adjusted compensatorily based on different age ranges, including: The elderly users are divided into multiple age ranges. User information and the number of various dental categories are obtained for each age range. Dental categories include abnormal teeth, filled teeth, missing teeth, and healthy teeth. The average number of each dental category in each age range is calculated, and a compensation coefficient is set for each dental category. The health index is adjusted based on the compensation coefficient.

[0011] In conjunction with the first aspect, in the fourth implementation of the first aspect of this application, spatiotemporal alignment and association of the multimodal data includes: The historical time period is divided into multiple sub-time periods. A second label is set for the multimodal data of each sub-time period. The second label is used to mark whether the sub-time period is missing and the missing modality type. Each type of modality data corresponding to each sub-time period is uniformly converted into a triple. The triple includes a timestamp, a feature value, and a feature type. The multimodal data is spatiotemporally aligned based on the triple and related associations are performed.

[0012] In conjunction with the first aspect, in the fifth implementation of the first aspect of this application, the data of the same modality within the historical time period are fused based on the judgment result, including: Based on the second label, it is determined whether there is a missing modality in each sub-time period. If not, an extraction model is set for each modality data to extract the key features of each modality data, output a feature value embedding vector of uniform length, and convert the timestamp into a time embedding vector of fixed dimension, and the feature type into a feature type embedding vector of fixed dimension. The average value of the feature value embedding vector, the time embedding vector and the feature type embedding vector of multiple sub-time periods is defined as the preliminary fusion vector of the same modality data.

[0013] In conjunction with the first aspect, in the sixth implementation of the first aspect of this application, the data of the same modality within the historical time period are fused based on the judgment result, including: If a missing modality exists, the modality data present in each sub-time period is defined as valid modality data. Based on the second label, it is determined whether the missing modality data is oral image data. If so, a virtual oral feature map is generated based on the valid modality data. The virtual oral feature map is used as the modality data of the oral image data. A corresponding time embedding vector, feature value embedding vector, or feature type embedding vector is generated for each modality data. Different modality weights are set for the modality data labeled with the second label and other modality data in each sub-time period. The average value of the embedding vectors corresponding to all sub-time periods in the historical time period under the same modality data is defined as the preliminary fusion vector of the same modality data.

[0014] In conjunction with the first aspect, in the seventh implementation of the first aspect of this application, generating a virtual oral cavity feature map based on the effective modal data includes: A generative adversarial network (GAN) model is established, and the effective modal data is input into the GAN model. The GAN model decodes the effective modal data to generate a virtual oral cavity feature map, and a discriminator is set to perform adversarial training on the virtual oral cavity feature map.

[0015] In conjunction with the first aspect, in the eighth implementation of the first aspect of this application, a comprehensive analysis of multiple modal data is performed based on the evaluation model, including: Based on the range of feature values ​​within the comprehensive feature vector, the comprehensive feature vector is converted into a symbolic feature vector. The numerical range of multiple modal data corresponding to different severity levels of each oral disease is set. The symbolic feature vector is matched with each numerical range. If the match is successful, the corresponding oral disease type and severity are obtained. If a match is not found, an evaluation model is built based on a deep learning model. The model is trained using feature vectors labeled with oral disease types, and the age of the elderly person is obtained. The age range is determined to obtain the corresponding age compensation. The age compensation includes saliva flow rate compensation and chewing force attenuation rate compensation. The age compensation is embedded in the symbolic feature vector and input into the evaluation model. The evaluation model analyzes the compensated symbolic feature vector and outputs the possible disease types and risk probabilities.

[0016] Secondly, this application provides an oral health behavior monitoring system for the elderly based on multi-source data fusion, the system comprising: The preliminary identification module is used to acquire multimodal data of multiple elderly people over a historical period. The multimodal data includes lifestyle behavior data, oral image data, and time series data of multiple physiological indicators. Based on the oral image data, the oral condition of the elderly people is preliminarily identified and a health index is obtained. The adjustment module is used to obtain the age range of the elderly, and to compensate and adjust the health index based on different age ranges. If the adjusted health index is less than a first threshold, the multimodal data is spatiotemporally aligned and correlated. The fusion module is used to determine whether there are missing modes in the modal data of multiple sub-time periods within the historical time period, and to perform fusion processing on the same modal data within the historical time period based on the determination result to obtain a preliminary fusion vector. The preliminary fusion vectors of different modalities are fused to obtain a comprehensive feature vector. The assessment module is used to establish an assessment model. The comprehensive feature vector is input into the assessment model, which performs comprehensive analysis on multiple modal data, outputs the risk probability and disease type of oral diseases that elderly users will experience in the future, sets behavioral adjustment strategies for each disease type, and guides the elderly to improve their oral health.

[0017] Compared with the prior art, the beneficial effects of the present invention are at least as follows: This application comprehensively assesses the oral health status of the elderly by integrating time-series data of lifestyle behavior, oral imaging, and physiological indicators. It adjusts the health index according to age range to ensure accurate assessment based on the characteristics of different age groups, avoiding over- or under-management of health. Through spatiotemporal alignment and modal data fusion, the system effectively addresses data gaps and ensures information integrity, further improving the accuracy of the analysis. After comprehensively analyzing multimodal data, the assessment model can predict the risk and type of oral diseases in the elderly, providing personalized behavioral adjustment strategies for each disease type to help them take appropriate health interventions, thereby improving their oral health status and reducing the probability of oral diseases. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram of an embodiment of the oral health behavior monitoring method for the elderly based on multi-source data fusion in this application. Figure 2 This is a schematic diagram illustrating the different age ranges of teeth in the embodiments of this application; Figure 3 This is a schematic diagram of one embodiment of the oral health behavior monitoring system for the elderly based on multi-source data fusion in this application. Detailed Implementation

[0020] This application provides a method and system for monitoring oral health behavior in the elderly based on multi-source data fusion. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data used can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0021] For ease of understanding, the specific process of the embodiments of this application is described below. Please refer to [link / reference]. Figure 1 One embodiment of the method for monitoring oral health behaviors of the elderly based on multi-source data fusion in this application includes: Step S1: Obtain multimodal data of multiple elderly people over a historical period. The multimodal data includes lifestyle data, oral image data, and time series data of multiple physiological indicators. Based on the oral image data, the oral condition of the elderly people is initially identified and a health index is obtained.

[0022] Specifically, most existing oral health monitoring systems rely on a single data source, such as oral images or physiological indicators. This single-data-source monitoring method cannot comprehensively reflect the oral health status of the elderly. Moreover, the elderly may not be able to provide complete data due to physical inconvenience or cognitive decline. For example, oral images may be missing or behavioral data may be incomplete. Furthermore, it is impossible to provide personalized health management recommendations based on the individual differences of the elderly, resulting in limited effectiveness of intervention measures.

[0023] Therefore, to address the aforementioned issues and better realize the monitoring of oral health behaviors in the elderly, enabling behavioral intervention before the onset of diseases, this application proposes a method for monitoring oral health behaviors in the elderly. First, multimodal data from multiple elderly individuals over historical time periods is acquired. This multimodal data includes lifestyle behavior data, oral image data, and time-series data of multiple physiological indicators. Lifestyle behavior data is obtained using devices such as smart toothbrushes, smart bracelets, and mobile applications, capturing brushing frequency and time, as well as smoking and drinking frequencies. Oral image data is acquired using devices such as oral endoscopes or smartphone cameras, including panoramic images of the oral cavity and tooth surface images. Time-series data of multiple physiological indicators, including saliva flow rate, saliva pH, and chewing force, are acquired using devices such as saliva analyzers and chewing force analyzers. By collecting lifestyle behavior data, oral image data, and physiological indicator data, the oral health status of the elderly can be comprehensively assessed, avoiding the limitations of a single data source.

[0024] Computer vision technology is used to analyze oral image data to identify the condition of teeth, gums and periodontal tissues. Key features are extracted from the images, and a preliminary oral health index is calculated based on the extracted features and physiological indicators. The higher the health index, the healthier the elderly user's oral cavity is; otherwise, there is a greater risk of disease. The specific implementation steps will be explained later.

[0025] Step S2: Obtain the age range of the elderly, and adjust the health index based on different age ranges. If the adjusted health index is less than the first threshold, perform spatiotemporal alignment and correlation on the multimodal data.

[0026] Specifically, the oral health of the elderly naturally declines with age. Therefore, it is necessary to adjust the health index according to different age ranges. The specific adjustment process will be explained later. Through age compensation, the oral health of the elderly can be more accurately assessed, avoiding misjudgments caused by age factors. If the adjusted health index is greater than the first threshold, the elderly person's oral health is considered good and continued observation is maintained. If the adjusted health index is less than the first threshold, the elderly person's oral health is considered to require further attention, and the next step is to align the data of different modalities in time and space to ensure data consistency and comparability. The specific alignment method will be explained later. Through spatiotemporal alignment, the consistency of data of different modalities in time and space is ensured, improving the integrity and comparability of the data.

[0027] Step S3: Determine whether there are missing modalities in the modal data of multiple sub-time periods within the historical time period. Based on the determination result, perform fusion processing on the same modal data within the historical time period to obtain a preliminary fusion vector. Fuse the preliminary fusion vectors of different modalities to obtain a comprehensive feature vector.

[0028] Specifically, a historical time period, such as one month, is divided into multiple sub-time periods, such as one day. Multimodal data from elderly users is collected once a day. However, due to physical limitations or cognitive decline, elderly users may not be able to provide complete data. For each sub-time period, it is checked for missing modal data. Based on the presence of missing data, data of the same modality are fused to generate a preliminary fusion vector. The specific fusion method will be explained later. The preliminary fusion vectors of different modalities are then fused to generate a comprehensive feature vector, for example, a comprehensive feature vector = {preliminary fusion vector of physiological indicators, preliminary fusion vector of behavioral data, and preliminary fusion vector of image data}.

[0029] Step S4: Establish an assessment model. Input the comprehensive feature vector into the assessment model. The assessment model performs comprehensive analysis on multiple modal data, outputs the risk probability and disease type of oral diseases that elderly users will develop in the future, sets behavioral adjustment strategies for each disease type, and guides the elderly to improve their oral health.

[0030] Specifically, an evaluation model is built based on a deep learning model. The evaluation model is trained using feature vectors labeled with oral disease types. The evaluation model is used to predict the probability of elderly users developing different types of oral diseases (such as dental caries, periodontal disease, oral cancer, etc.) in the future. The specific evaluation process will be explained in detail later.

[0031] After outputting the disease type and risk probability, the system develops behavioral adjustment strategies based on different disease types, including dietary recommendations, oral hygiene habits, regular check-ups, and dental treatment. For example, if the prediction results show a high risk of periodontal disease, the system recommends that patients have regular oral check-ups, improve their brushing methods, or increase oral care. Based on different ages, oral conditions, and lifestyle habits, the system can provide elderly people with customized health behavior adjustment strategies to improve the effectiveness of health management.

[0032] In one specific embodiment, the preliminary identification of the oral health status of elderly individuals and the acquisition of a health index specifically includes the following steps: Each oral image is assigned a first label, which indicates whether the corresponding elderly user has oral problems and the type of problems. A target processing model is selected based on the first label. The target processing model performs image processing on the oral image data to obtain the first image data. A recognition model is created based on a machine learning model. The first label, the first image data, and the oral image data are used as training data for the recognition model. The recognition model identifies the oral condition of the elderly user, including whether there are oral problems and the probability of having oral problems. The probability of having oral problems is converted into a health index and output.

[0033] Specifically, each oral image is labeled based on the presence and type of oral problems. Image 1: The first label is "healthy," indicating no oral problems were found. Image 2: The first label is "dental caries," indicating that the image shows signs of dental caries. Based on the characteristics of different oral problems, the most suitable image processing method and corresponding target processing model are selected. For example, grayscale processing may help detect dental caries, but edge detection may be more effective for detecting periodontal disease. The specific process for selecting the target processing model will be explained later.

[0034] The selected target processing model is used to process oral image data to obtain the first image data. Image processing includes grayscale conversion, edge detection, and color conversion. These image processing techniques typically make the features in the image more prominent, thereby helping the model to better identify meaningful oral problem features during the learning process. A recognition model is created based on a neural network model, using both processed and unprocessed images as training data. The model can learn features at different levels. Once the model is trained, it can be used to analyze new oral image data. The recognition model will determine the oral condition of the elderly based on the image data. For example, the model may identify the presence of cavities or swollen gums in the image and give the corresponding risk probability. The risk probability can be converted into a health index. For example, if the risk probability of cavities is 80%, the health index may be 20 (i.e., poor health condition, high probability of oral problems). In this way, potential oral health problems in the elderly, such as cavities and periodontal disease, can be identified early, avoiding the deterioration of oral diseases due to ignoring early symptoms.

[0035] In one specific embodiment, the target selection processing model based on the first label specifically includes the following steps: Oral image data is input into multiple different machine learning models. Each machine learning model uses different image processing methods to process the first image data and outputs a corresponding feature vector. The image processing methods include grayscale processing, edge detection processing, and color conversion processing. The matching degree between the feature vector output by each machine learning model and the corresponding first label is calculated. The machine learning model with a matching degree greater than a first threshold is defined as the target processing model.

[0036] Specifically, suppose there is a set of oral image data that shows different types of oral problems, such as tooth decay, periodontal disease, and overall oral health. This oral image data is then input into multiple machine learning models, each employing different image processing techniques to extract different types of features from the oral images: Grayscale conversion: Converting color images to grayscale removes color information and highlights morphological features, which is especially important for models that need to analyze tooth morphology or surface features, such as the edges of decayed teeth; Edge detection: Detecting the edges of teeth in the image helps the model identify the contours and defects of teeth, which is very effective for detecting oral problems; Color conversion: In some cases, changes in tooth color, such as tooth decay or bleeding gums, can serve as an important feature for judging oral health. Therefore, color conversion can highlight differences in tooth color, helping the model better identify problems. Each machine learning model generates a feature vector that summarizes the key features of the input image. The matching degree between each model's generated feature vector and the first label is calculated using cosine similarity or Euclidean distance. The matching degrees of all models are compared. If a model's matching degree exceeds a first threshold (e.g., 80%), it is considered to have a good match with the label and can effectively identify oral problems. The model with the highest matching degree is selected as the target processing model; that is, this model will be used to process and analyze subsequent oral images.

[0037] In one specific embodiment, adjusting the health index based on different age ranges includes the following steps: The elderly users are divided into multiple age ranges. User information and the number of various dental categories are obtained for each age range. Dental categories include abnormal teeth, filled teeth, missing teeth and healthy teeth. The average number of each dental category in each age range is calculated, and a compensation coefficient is set for each dental category. The health index is adjusted based on the compensation coefficient.

[0038] Specifically, the elderly users are divided into several age ranges, such as 60-69 years old, 60-69 years old, and over 80 years old. User information for each age range is retrieved from the system, including the number of users, and information on the number of tooth categories for each age range is collected. Tooth categories include healthy teeth (teeth without obvious problems), abnormal teeth (teeth with problems such as cavities or periodontitis), filled teeth (teeth that have undergone filling treatment), and missing teeth (teeth that have been lost or extracted). Based on actual data, the average number of each tooth category within each age range is calculated, such as... Figure 2 The diagram shows the average number of tooth categories across different age ranges. Compensation coefficients are set for different tooth categories, reflecting the degree of influence of each tooth category on the health index. The compensation coefficients for abnormal teeth, filled teeth, missing teeth, and healthy teeth are -0.1, -0.2, -0.3, and 1, respectively.

[0039] For each elderly person, the similarity between the number of their tooth categories and the average number for their corresponding age range is calculated. The tooth category with the highest similarity is selected, and the health index is adjusted according to the compensation coefficient corresponding to that tooth category. For example, if the highest similarity is for abnormal teeth, the original health index, assumed to be 20, is adjusted based on the first formula: Adjusted Health Index = Original Health Index + Compensation Coefficient * Original Health Index. Substituting these values, we get the adjusted health index = 20 - 0.1 * 20 = 18. By adjusting the health index through the compensation coefficient, multiple tooth categories such as healthy teeth, abnormal teeth, filled teeth, and missing teeth are comprehensively considered, enabling a more comprehensive assessment of the oral health status of the elderly and avoiding the limitations of a single indicator.

[0040] In one specific embodiment, spatiotemporal alignment and association of multimodal data specifically includes the following steps: The historical time period is divided into multiple sub-time periods. A second label is set for the multimodal data of each sub-time period. The second label is used to indicate whether there is a missing data in the sub-time period and the missing modality type. Each type of modality data corresponding to each sub-time period is uniformly converted into a triple. The triple includes a timestamp, feature value and feature type. Based on the triple, the multimodal data is spatiotemporally aligned and related.

[0041] Specifically, a second label is set for the multimodal data of each sub-time period to indicate whether there are missing data in that sub-time period and the type of missing modality. For example, in sub-time period 1: 2023-10-01, the second label is: {"Lifestyle data": "1", "Oral image data": "0", "Physiological index data": "1"}, where 1 indicates the presence of the modality and 0 indicates the absence of the modality. Each modality data corresponding to each sub-time period is uniformly converted into a triplet, which includes a timestamp, feature value, and feature type. For example: Sub-time period 1: Lifestyle data: ["2023-10-01", 2, "Brushing frequency"]; where 2 indicates brushing teeth twice a day; Physiological index data: ["2023-10-01", 0.4, "Salivary flow rate"], where 0.4 indicates salivary flow rate. Based on the triplet, multimodal data is spatiotemporally aligned and correlated. Spatiotemporally aligning and correlating data from different modalities allows for a more comprehensive assessment of the oral health status of the elderly, avoiding the limitations of single-modal data.

[0042] In one specific embodiment, the fusion processing of data of the same modality within a historical time period based on the judgment result specifically includes the following steps: Based on the second label, determine whether there is a missing modality in each sub-time period. If not, set an extraction model for each modality data to extract the key features of each modality data, output a feature value embedding vector of uniform length, convert the timestamp into a fixed-dimensional time embedding vector, convert the feature type into a fixed-dimensional feature type embedding vector, and define the average value of the feature value embedding vector, time embedding vector and feature type embedding vector of multiple sub-time periods as the initial fusion vector of the same modality data.

[0043] Specifically, assume that there are two sub-time periods within the historical time period: sub-time period 1 and sub-time period 2, and that no missing modalities exist in each sub-time period. Sub-time period 2 consists of the following triples: lifestyle behavior data: ["2023-10-02", 3, "brushing frequency"], oral image data: ["2023-10-02", 1, "number of cavities"], and physiological indicator data: ["2023-10-02", 0.5, "saliva flow rate"].

[0044] Sub-time period 1: For lifestyle data: extract brushing frequency and output a feature value embedding vector of length 1 to obtain [2]; oral image data: extract the number of cavities and output a feature value embedding vector of length 1 to obtain [1]; physiological index data: extract saliva flow rate and output a feature value embedding vector of length 1 to obtain [0.4].

[0045] Using the above method, the life behavior data, oral image data and oral image data of sub-time period 2 are converted into corresponding feature value embedding vectors, namely [3], [1] and [0.5].

[0046] The timestamp is converted into a fixed-dimensional embedding vector using a pre-trained time embedding model. It is assumed that the dimensions of the time embedding vector and the feature type embedding vector are both 1. For example, the time embedding vector of sub-time period 1: 2023-10-01 is converted into a time embedding vector [0.4], and the time embedding vector of sub-time period 2: 2023-10-02 is [0.5].

[0047] Assign a fixed-dimensional embedding vector to each feature type. For example, for lifestyle data [1], add the feature value embedding vector, time embedding vector and feature type embedding vector of sub-time period 1 and sub-time period 2 to the lifestyle data modality respectively, and calculate the average value to obtain the preliminary fusion vector of lifestyle data modality: [(0.4+0.5) / 2,(2+3) / 2,(1+1) / 2]=[0.45, 2.5,1], where 0.45, 2.5,1 represent the average value of time embedding vector, feature value embedding vector and feature type embedding vector of lifestyle data modality respectively.

[0048] In one specific embodiment, the fusion processing of data of the same modality within a historical time period based on the judgment result further includes the following steps: If a missing modality exists, the modality data present in each sub-time period is defined as valid modality data. Based on the second label, it is determined whether the missing modality data is oral image data. If so, a virtual oral feature map is generated based on the valid modality data. The virtual oral feature map is used as the modality data of the oral image data. A corresponding time embedding vector, feature value embedding vector, or feature type embedding vector is generated for each modality data. Different modality weights are set for the modality data labeled with the second label and other modality data in each sub-time period. The average value of the embedding vectors corresponding to all sub-time periods in the historical time period under the same modality data is defined as the preliminary fusion vector of the same modality data.

[0049] Specifically, it is assumed that there are missing modalities in sub-time period 3 and sub-time period 4. Among them, sub-time period 3: 2023-10-03: lifestyle data: [2] (brushing frequency 2 times / day), oral image data is missing, physiological index data: [0.4] (saliva flow rate 0.4 ml / min).

[0050] Sub-time period 3: Oral cavity image data is missing. If the missing modality is oral cavity image data, a virtual oral cavity feature map is generated based on the valid modality data.

[0051] For sub-time period 3, a virtual oral cavity feature map is generated using a generative adversarial network, which will be explained in detail later. Different modal weights are set for the modal data labeled with the second label and other modal data in each sub-time period. For example, since oral cavity image data is lacking in sub-time period 3, the weight of this modal data is set to be lower. For example, the weights of daily behavior data, virtual oral cavity image data and physiological indicator data are 0.4, 0.2 and 0.4, respectively. The average value of each modal data is calculated based on the modal weights. The daily behavior data in sub-time period 3 is [2 * 0.4] = [0.8], the virtual oral cavity image data is [0.8 * 0.3] = [0.24], and the physiological indicator data is [0.4 * 0.3] = [0.12].

[0052] The average value of each modality data is added to the corresponding embedding vector. The time embedding vector, feature value embedding vector, and feature type embedding vector corresponding to the life behavior data of sub-time period 1, sub-time period 2, and sub-time period 3 are added together and the average value is calculated to obtain the preliminary fusion vector of life behavior data. Other modalities are not illustrated.

[0053] In one specific embodiment, generating a virtual oral cavity feature map based on effective modal data specifically includes the following steps: A generative adversarial network (GAN) model is established. The effective modal data is input into the GAN model, which decodes the effective modal data to generate a virtual oral cavity feature map. A discriminator is then set up to perform adversarial training on the virtual oral cavity feature map.

[0054] Specifically, the Generative Adversarial Network (GAN) consists of a generator and a discriminator. The generator is responsible for generating virtual oral cavity feature maps, while the discriminator is responsible for distinguishing between the generated images and real images. The generator is a deep neural network, including an encoder that encodes the input valid modal data into a low-dimensional feature representation, and a decoder that decodes the encoded feature representation into a feature value embedding vector of the virtual oral cavity feature map. The discriminator evaluates the generated virtual oral cavity feature map and compares it with a real oral cavity image, providing adversarial loss. Through adversarial training, the generator continuously optimizes the generated images to make them as close as possible to real images. By generating virtual oral cavity feature maps through the GAN, missing modal data is filled in, ensuring data integrity.

[0055] In one specific embodiment, the comprehensive analysis of multiple modal data based on the evaluation model includes the following steps: Based on the range of eigenvalues ​​within the comprehensive feature vector, the comprehensive feature vector is converted into a symbolic feature vector. The numerical range of multiple modal data corresponding to different severity levels of each oral disease is set. The symbolic feature vector is matched with each numerical range. If the match is successful, the corresponding oral disease type and severity are obtained.

[0056] If a match is not found, an evaluation model is built based on a deep learning model. The model is trained using feature vectors labeled with oral disease types, and the age of the elderly is obtained. The age range is determined to obtain the corresponding age compensation, which includes saliva flow rate compensation and chewing force attenuation rate compensation. The age compensation is embedded into the symbolic feature vector and input into the evaluation model. The evaluation model analyzes the compensated symbolic feature vector and outputs the possible disease types and risk probabilities.

[0057] Specifically, sub-time period 1: lifestyle data: {brushing frequency: medium, sugar intake frequency: medium}, oral imaging data: {number of cavities: medium, degree of gingival inflammation: low}, physiological indicator data: {saliva flow rate: low, saliva pH value: normal}.

[0058] Numerical ranges of multiple modalities were set for different severity levels of each oral disease. For example, high risk of periodontitis: lifestyle data: low brushing frequency or high sugar intake frequency; oral image data: high number of cavities or high degree of gingival inflammation; physiological indicator data: low saliva flow rate or acidic saliva pH. High risk of tooth decay: lifestyle data: high sugar intake frequency; oral image data: high number of cavities; physiological indicator data: acidic saliva pH. The symbolic feature vector was matched with each numerical range. After matching, sub-time period 1 did not meet all the conditions of "high risk of periodontitis" or "high risk of tooth decay". Assuming that the symbolic feature vector did not match successfully, a deep learning model was used for further analysis.

[0059] An evaluation model is built using deep learning models (such as Convolutional Neural Networks (CNNs) or Recurrent Neural Networks (RNNs), and trained using feature vectors labeled with oral disease types. Assuming the elderly person is 75 years old, age-compensated features are obtained: saliva flow rate typically decreases with age; for individuals over 70, the saliva flow rate compensation is 0.1 ml / min, and the chewing force attenuation rate compensation is 2%. The compensated symbolic feature vectors are: Saliva flow rate: 0.4 + 0.1 = 0.5 ml / min (moderate), Chewing force attenuation rate: 18% - 2% = 16% (mild attenuation). The compensated symbolic feature vectors are input into the evaluation model, which analyzes them and outputs the possible disease types and their probabilities. For example, the evaluation model outputs a probability of 85% for periodontitis.

[0060] The above describes a method for monitoring oral health behavior in the elderly based on multi-source data fusion, as described in the embodiments of this application. The following describes a system for monitoring oral health behavior in the elderly based on multi-source data fusion, as described in the embodiments of this application. Please refer to [link / reference]. Figure 3One embodiment of the oral health behavior monitoring system for the elderly based on multi-source data fusion in this application includes: The preliminary identification module is used to acquire multimodal data of multiple elderly people over a historical period. The multimodal data includes daily behavior data, oral image data, and time series data of multiple physiological indicators. Based on the oral image data, the oral condition of the elderly is preliminarily identified and a health index is obtained. The adjustment module is used to obtain the age range of the elderly and adjust the health index based on different age ranges. If the adjusted health index is less than the first threshold, the multimodal data is spatiotemporally aligned and correlated. The fusion module is used to determine whether there are missing modes in the modal data of multiple sub-time periods within a historical time period. Based on the judgment result, the same modal data within the historical time period is fused to obtain a preliminary fusion vector. The preliminary fusion vectors of different modalities are fused to obtain a comprehensive feature vector. The assessment module is used to build an assessment model. The comprehensive feature vector is input into the assessment model, which performs comprehensive analysis on multiple modal data, outputs the risk probability and disease type of oral diseases that elderly users will develop in the future, and sets behavioral adjustment strategies for each disease type to guide the elderly to improve their oral health.

[0061] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0062] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0063] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for monitoring oral health behaviors of the elderly based on multi-source data fusion, characterized in that, The method includes: Step S1: Obtain multimodal data of multiple elderly people over a historical time period. The multimodal data includes lifestyle behavior data, oral image data, and time series data of multiple physiological indicators. Based on the oral image data, the oral condition of the elderly people is initially identified and a health index is obtained. Step S2: Obtain the age range of the elderly, and adjust the health index based on different age ranges. If the adjusted health index is less than the first threshold, perform spatiotemporal alignment association on the multimodal data. Step S3: Determine whether there are missing modalities in the modal data of multiple sub-time periods within the historical time period. Based on the determination result, perform fusion processing on the same modal data within the historical time period to obtain a preliminary fusion vector. Then, fuse the preliminary fusion vectors of different modalities to obtain a comprehensive feature vector. Step S4: Establish an evaluation model. Input the comprehensive feature vector into the evaluation model. The evaluation model performs comprehensive analysis on multiple modal data, outputs the risk probability and disease type of oral diseases that elderly users will experience in the future, sets behavioral adjustment strategies for each disease type, and guides the elderly to improve their oral health.

2. The method according to claim 1, characterized in that, To initially identify the oral health status of older adults and obtain health indices, including: A first label is assigned to each oral image data. The first label indicates whether the corresponding elderly user has oral problems and the type of problems. A target processing model is selected based on the first label. The target processing model performs image processing on the oral image data to obtain first image data. A recognition model is created based on a machine learning model. The first label, the first image data, and the oral image data are used as training data for the recognition model. The recognition model identifies the oral condition of the elderly user. The oral condition includes whether there are oral problems and the probability of having oral problems. The probability of having oral problems is converted into a health index and output.

3. The method according to claim 2, characterized in that, Based on the first label, a target processing model is selected, including: The oral cavity image data is input into multiple different machine learning models. Each machine learning model uses a different image processing method to process the first image data and outputs a corresponding feature vector. The image processing method includes grayscale processing, edge detection processing, and color conversion processing. The matching degree between the feature vector output by each machine learning model and the corresponding first label is calculated. The machine learning model corresponding to the matching degree being greater than a first threshold is defined as the target processing model.

4. The method according to claim 1, characterized in that, The health index is adjusted for compensation based on different age ranges, including: The elderly users are divided into multiple age ranges. User information and the number of various dental categories are obtained for each age range. Dental categories include abnormal teeth, filled teeth, missing teeth, and healthy teeth. The average number of each dental category in each age range is calculated, and a compensation coefficient is set for each dental category. The health index is adjusted based on the compensation coefficient.

5. The method according to claim 1, characterized in that, Spatiotemporal alignment and correlation of the multimodal data includes: The historical time period is divided into multiple sub-time periods. A second label is set for the multimodal data of each sub-time period. The second label is used to mark whether the sub-time period is missing and the missing modality type. Each type of modality data corresponding to each sub-time period is uniformly converted into a triple. The triple includes a timestamp, a feature value, and a feature type. The multimodal data is spatiotemporally aligned based on the triple and related associations are performed.

6. The method according to claim 5, characterized in that, Based on the judgment result, the data of the same modality within the historical time period are fused, including: Based on the second label, it is determined whether there is a missing modality in each sub-time period. If not, an extraction model is set for each modality data to extract the key features of each modality data, output a feature value embedding vector of uniform length, and convert the timestamp into a time embedding vector of fixed dimension, and the feature type into a feature type embedding vector of fixed dimension. The average value of the feature value embedding vector, the time embedding vector and the feature type embedding vector of multiple sub-time periods is defined as the preliminary fusion vector of the same modality data.

7. The method according to claim 6, characterized in that, Based on the judgment result, the data of the same modality within the historical time period are fused, which also includes: If a missing modality exists, the modality data present in each sub-time period is defined as valid modality data. Based on the second label, it is determined whether the missing modality data is oral image data. If so, a virtual oral feature map is generated based on the valid modality data. The virtual oral feature map is used as the modality data of the oral image data. A corresponding time embedding vector, feature value embedding vector, or feature type embedding vector is generated for each modality data. Different modality weights are set for the modality data labeled with the second label and other modality data in each sub-time period. The average value of the embedding vectors corresponding to all sub-time periods in the historical time period under the same modality data is defined as the preliminary fusion vector of the same modality data.

8. The method according to claim 7, characterized in that, Generating a virtual oral cavity feature map based on the effective modal data includes: A generative adversarial network (GAN) model is established, and the effective modal data is input into the GAN model. The GAN model decodes the effective modal data to generate a virtual oral cavity feature map, and a discriminator is set to perform adversarial training on the virtual oral cavity feature map.

9. The method according to claim 1, characterized in that, Based on the aforementioned evaluation model, a comprehensive analysis of multiple modal data is performed, including: Based on the range of feature values ​​within the comprehensive feature vector, the comprehensive feature vector is converted into a symbolic feature vector. The numerical range of multiple modal data corresponding to different severity levels of each oral disease is set. The symbolic feature vector is matched with each numerical range. If the match is successful, the corresponding oral disease type and severity are obtained. If a match is not found, an evaluation model is built based on a deep learning model. The model is trained using feature vectors labeled with oral disease types, and the age of the elderly person is obtained. The age range is determined to obtain the corresponding age compensation. The age compensation includes saliva flow rate compensation and chewing force attenuation rate compensation. The age compensation is embedded in the symbolic feature vector and input into the evaluation model. The evaluation model analyzes the compensated symbolic feature vector and outputs the possible disease types and risk probabilities.

10. A multi-source data fusion-based oral health behavior monitoring system for the elderly, used to implement the multi-source data fusion-based oral health behavior monitoring method for the elderly as described in any one of claims 1-9, characterized in that, The system includes: The preliminary identification module is used to acquire multimodal data of multiple elderly people over a historical period. The multimodal data includes lifestyle behavior data, oral image data, and time series data of multiple physiological indicators. Based on the oral image data, the oral condition of the elderly people is preliminarily identified and a health index is obtained. The adjustment module is used to obtain the age range of the elderly, and to compensate and adjust the health index based on different age ranges. If the adjusted health index is less than a first threshold, the multimodal data is spatiotemporally aligned and correlated. The fusion module is used to determine whether there are missing modes in the modal data of multiple sub-time periods within the historical time period. Based on the determination result, the same modal data within the historical time period is fused to obtain a preliminary fusion vector. The preliminary fusion vectors of different modalities are fused to obtain a comprehensive feature vector. The assessment module is used to establish an assessment model. The comprehensive feature vector is input into the assessment model, which performs comprehensive analysis on multiple modal data, outputs the risk probability and disease type of oral diseases that elderly users will experience in the future, and sets behavioral adjustment strategies for each disease type to guide the elderly to improve their oral health.

Citation Information

Patent Citations

  • Research model reflecting oral health and model construction method thereof

    CN119673426A

  • Oral hypofunction self-detection method and system for old people

    CN120220932A

Cited By

  • Oral disease monitoring method and device for elderly patients and storage medium

    CN122398183A