Multi-modal diagnosis method for dental diseases and electronic equipment

Through multimodal diagnostic methods, combined with CT images and disease text information, the diagnosis of dental diseases is used using Mask R-CNN and SAM models, which solves the problem that traditional diagnosis relies on doctors' experience and achieves more efficient and accurate diagnostic results.

CN120495300AActive Publication Date: 2025-08-15PEKING UNIV SCHOOL OF STOMATOLOGY +1
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510985071.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-08-15
Estimated Expiration
2045-07-17

AI Technical Summary

Technical Problem

Traditional dental disease diagnosis methods rely on doctors' experience, have low accuracy, are prone to misdiagnosis and misdiagnosis, and have limited ability to detect early subtle lesions, so it is impossible to comprehensively consider the patient's medical history and systemic health status.

Method used

Multimodal diagnostic method is adopted, combining CT images and disease text information, and using Mask R-CNN and SAM models for pathological information extraction and fusion, multimodal model is constructed for diagnosis, and comprehensive analysis is performed by combining image and text information.

Benefits of technology

It improves the accuracy and efficiency of dental disease diagnosis, reduces the rate of misdiagnosis, and can consider patient information more comprehensively, reduce resource consumption and doctor burden, and adapt to stable diagnosis in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495300A_ABST
    Figure CN120495300A_ABST
Patent Text Reader

Abstract

The invention provides a multi-mode diagnosis method for dental diseases and electronic equipment. The multi-modal diagnosis method for the dental diseases comprises the following steps: obtaining a patient CT image data set and an illness state text information data set; inputting the CT image data of the patient into the trained Mask R-CNN model to obtain first CT pathological information; inputting the CT image data of the patient into the trained SAM model to obtain second CT pathological information; performing information fusion on the first CT pathological information and the corresponding second CT pathological information to obtain CT pathological information, and obtaining a CT pathological information data set based on the CT pathological information; pre-training the constructed multi-modal model based on the patient CT image data set, the illness state text information data set and the CT pathological information data set to obtain a trained multi-modal model; and diagnosing the dental diseases based on the trained multi-modal model. The purposes of improving diagnosis and treatment efficiency and accuracy and avoiding misdiagnosis are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical technology, and in particular relates to a multimodal diagnostic method and electronic equipment for dental diseases. Background Art

[0002] In the medical field, accurate diagnosis of dental diseases is crucial for patient treatment and prognosis. Traditional dental diagnosis relies primarily on the physician's subjective judgment. Doctors visually inspect the patient's oral cavity and perform simple instrumental examinations, such as probing the tooth surface with a probe, to make a preliminary assessment of conditions such as caries and periodontitis. X-rays are also a commonly used auxiliary diagnostic tool, allowing doctors to examine the internal structure of teeth, such as the root and pulp cavity, and are particularly helpful in diagnosing conditions such as apical periodontitis and impacted teeth.

[0003] However, the accuracy of traditional dental disease diagnosis methods is highly dependent on the doctor's experience, professional level and condition. There are differences in judgment among different doctors, and doctors are prone to misdiagnosis and missed diagnosis due to fatigue and negligence. At the same time, the naked eye and traditional X-rays have limited ability to detect early and subtle lesions (such as early dentin caries and minor pulp inflammation). When they are discovered, the condition is often already serious, delaying the best time for treatment. At the same time, traditional methods focus on the local condition of the teeth and make insufficient use of the patient's overall condition information such as medical history and general health status, which are crucial for diagnosis and treatment plan formulation. Summary of the Invention

[0004] In response to the problems existing in the prior art, the present invention provides a multimodal diagnostic method and electronic equipment for dental diseases, which at least partially solve the problem of low diagnostic efficiency existing in the prior art.

[0005] In a first aspect, embodiments of the present disclosure provide a multimodal diagnostic method for dental diseases, comprising: Based on the acquired patient CT image data and the corresponding condition text information data, a patient CT image dataset and a condition text information dataset are obtained; The patient's CT image data is input into the trained Mask R-CNN model to obtain the first CT pathology information; Inputting the patient's CT image data into the trained SAM model to obtain second CT pathology information, where the second CT pathology information corresponds one-to-one with the first CT pathology information; fusing the first CT pathology information and the corresponding second CT pathology information to obtain CT pathology information, and obtaining a CT pathology information dataset based on the CT pathology information; Pre-training the constructed multimodal model based on the patient CT image dataset, the condition text information dataset, and the CT pathology information dataset to obtain a trained multimodal model; Diagnose dental diseases based on the trained multimodal model.

[0006] In a second aspect, an embodiment of the present disclosure further provides an electronic device, the electronic device comprising: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the multimodal diagnosis method for dental diseases described in any one of the first aspects.

[0007] The present invention provides a multimodal diagnostic method and electronic device for dental diseases. This multimodal diagnostic method uses a trained multimodal model to diagnose patients using CT scans, replacing existing manual diagnosis methods and thereby improving diagnostic and treatment efficiency. The multimodal model training utilizes a dataset of text information about the condition and a dataset of CT pathology information derived from the Mask R-CNN and SAM models. The Mask R-CNN and SAM models provide more accurate and comprehensive information, thereby improving diagnostic and treatment accuracy and avoiding misdiagnoses. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The above and other objects, features and advantages of the present disclosure will become more apparent through a more detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings, wherein like reference numerals generally represent like components throughout the exemplary embodiments of the present disclosure.

[0009] Figure 1 This is a flow chart of a multimodal diagnostic method for dental diseases provided in the first embodiment of the present disclosure; Figure 2 A flowchart of training a Mask R-CNN model using a labeled CT image dataset in the multimodal diagnosis method for dental diseases provided in the first embodiment of the present disclosure; Figure 3 A flowchart of loading and fine-tuning the SAM model in the multimodal diagnostic method for dental diseases provided in the first embodiment of the present disclosure; Figure 4 A flowchart of obtaining a trained multimodal model in the multimodal diagnosis method for dental diseases provided in the first embodiment of the present disclosure; Figure 5 This is a flow chart of a multimodal diagnostic method for dental diseases provided in the second embodiment of the present disclosure; Figure 6 This is a functional block diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0010] The embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.

[0011] It should be clear that the following embodiments of the present disclosure are described through specific concrete examples, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the contents disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that the following embodiments and features in the embodiments can be combined with each other in the absence of conflict. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present disclosure.

[0012] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on this disclosure, it should be understood by those skilled in the art that an aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement an apparatus and / or practice a method. In addition, other structures and / or functionalities other than one or more of the aspects described herein can be used to implement this apparatus and / or practice this method.

[0013] It should also be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present disclosure. The illustrations only show components related to the present disclosure and are not drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component can be changed at will, and the component layout type may also be more complicated.

[0014] Additionally, in the following description, specific details are provided to provide a thorough understanding of the examples. However, one skilled in the art will appreciate that the aspects described can be practiced without these specific details.

[0015] Example 1: For ease of understanding, Figure 1 As shown, this embodiment discloses a multimodal diagnostic method for dental diseases, comprising: Step S101: obtaining a patient CT image dataset and a patient condition text information dataset based on the acquired patient CT image data and the corresponding condition text information data; Dental CT images of patients are collected using specialized CT scanning equipment. These images can include images of patients of different ages and genders from multiple medical institutions. They clearly depict the internal structure of teeth and surrounding tissues, providing intuitive imaging support for subsequent pathology analysis. Furthermore, textual information on the condition is collected, including the chief complaint, current medical history, and past medical history of patients of different ages and genders from multiple medical institutions. This textual information contains key information such as the disease progression and symptoms, which aids in image understanding and diagnosis. After the CT image dataset is collected, it can be preprocessed to improve image quality and usability. Common preprocessing operations include adjusting image contrast and brightness to make image features more clearly discernible; unifying image formats to facilitate subsequent processing by trained multimodal models; and normalizing images to map pixel values to a specific range, enabling the trained multimodal model to more stably learn image features. The collected CT pathology information dataset includes the location coordinates, size, morphological characteristics, and disease classification labels of dental lesions.

[0016] Step S102: Inputting the patient's CT image data into the trained Mask R-CNN model to obtain first CT pathology information; inputting the patient's CT image data into the trained SAM model to obtain second CT pathology information, where the second CT pathology information corresponds one-to-one with the first CT pathology information; fusing the first CT pathology information and the corresponding second CT pathology information to obtain CT pathology information, and obtaining a CT pathology information dataset based on the CT pathology information; Optionally, Mask R-CNN model training includes: Based on the Mask R-CNN framework, build a neural network structure for image segmentation; The acquired training data is input into the neural network structure, the learning rate is set, the optimizer is selected, and the neural network structure is trained based on the loss function obtained by combining the classification cross entropy loss function, the regression L1 / L2 loss function, and the mask segmentation loss function until the neural network structure converges.

[0017] Optional, SAM model training, including: Obtaining pre-trained weights of the SAM model; The acquired training data was input into the SAM model after loading the pre-trained weights. The weights of the general feature extraction layer selected at the bottom layer were frozen. The parameters of the high-level semantic segmentation-related layers were adjusted. Combined with the prior segmentation knowledge of the tooth structure, the learning rate was adjusted, and a lightweight optimizer was used for SAM model training.

[0018] When training the Mask R-CNN model and the SAM model, the first step is to construct training data. The training data is to label the CT image data in the CT image dataset of some patients to obtain the labeled CT image dataset, and then construct the training data based on the labeled CT image dataset. The Mask R-CNN model is a powerful instance segmentation model. Trained using a dataset of annotated CT images, the trained multimodal model can learn the characteristics and boundaries of different lesions (such as caries and periapical periodontitis). The trained multimodal model continuously compares predicted data with annotated CT image data and adjusts its parameters, gradually acquiring the ability to accurately segment lesions. The Segment Anything (SAM) model inherently possesses strong general segmentation capabilities. In dental diagnostic scenarios, it is optimized by loading its pretrained weights and then fine-tuning it to the specific characteristics of dental images. During fine-tuning, the weights of some underlying general feature extraction layers are frozen, and only the parameters of the higher-level semantic segmentation layers are adjusted. This, combined with prior knowledge from the dental field, allows the fine-tuned SAM model to focus on learning the characteristics of dental lesions.

[0019] The preprocessed CT images were input into the trained Mask R-CNN model and the fine-tuned SAM model, respectively. The Mask R-CNN model locates and segments the lesion areas in the CT image based on previously learned knowledge, outputting information such as the lesion's location, size, and shape. The SAM model, with its fine-tuned ability to capture microscopic dental lesions, further supplements and refines the lesion information. Combining the outputs of the Mask R-CNN model and the fine-tuned SAM model creates a CT pathology information dataset. This dataset records the specific conditions of the lesions in each CT image in detail, including the lesion's location coordinates, size, morphological characteristics, and disease classification labels, providing important image feature basis for subsequent multimodal diagnosis.

[0020] The integration of the output results of the Mask R-CNN model and the fine-tuned SAM model can be achieved in the following ways: The first step is to align the outputs of the two trained models, including spatial alignment and slice alignment. Spatial alignment ensures that the lesion regions output by the two trained models for the same CT image correspond to each other in spatial position. Slice alignment is applied to multi-layer CT images, ensuring that the outputs of the two trained models on each slice match each other. The outputs of the two trained models can be sorted and aligned according to the sequence number of the slices, so that the information about the lesions from the two trained models can be accurately associated on the slices with the same sequence number.

[0021] The spatial position alignment formula is: , in, is the original space position coordinate, is the position coordinate after transformation, is the radial basis function, is the control point, The parameters to be estimated, a and b are the data sets to be aligned, and a and b are matrices.

[0022] Slice correspondence is determined by calculating the similarity between slices. For example, using metrics such as cosine similarity or Euclidean distance, the point pairs with the highest similarity between the two slices are found as corresponding points.

[0023] The second step is to fuse the suspected lesion areas detected by the two trained models. Multi-data voting or weighted fusion methods can be used. For suspected lesion areas detected by the two trained models, if there are overlapping or similar parts, a majority vote can be used to determine the final lesion area, or different weights can be assigned based on the performance of the two trained models in detecting different types of lesions. For example, if the Mask R-CNN model performs better in detecting large lesion areas, while the SAM model has an advantage in capturing subtle lesions, the output results of the Mask R-CNN model can be given a higher weight for large lesions during integration, while the output results of the SAM model can be given a higher weight for subtle lesions. The output results of the two trained models are then linearly combined according to the weights to obtain the final lesion area.

[0024] The third step is to splice the feature information about the lesion output by the two trained models. The Mask R-CNN model may output geometric features such as the location, size, and shape of the lesion, and the SAM model may extract features such as the texture and grayscale of the lesion. These features are spliced into a longer feature vector as the comprehensive feature information of the lesion. The features output by the two trained models are then screened and fused to remove redundant features and retain the most representative features. Feature selection algorithms, such as correlation analysis and principal component analysis, can be used to evaluate the importance of each feature to lesion classification or diagnosis, select features with higher importance for fusion, and then combine these screened features to obtain a more concise and effective feature representation.

[0025] The fourth step is to verify and correct the outputs of the two trained models. This can be done by having a professional dentist manually verify the integrated CT pathology dataset, or by comparing it with known gold standard data (such as pathology examination results). The verification results can then be compared with known gold standard data (such as pathology examination results).

[0026] Step S103: pre-training the constructed multimodal model based on the patient CT image dataset, the condition text information dataset, and the CT pathology information dataset to obtain a trained multimodal model; First, the constructed multimodal model is pre-trained using a CT pathology information dataset, a patient CT image dataset, and a condition text information dataset. During the pre-training process, the multimodal model learns how to extract useful features from data of different modalities and attempts to fuse these features. For example, the multimodal model needs to learn how to associate lesion features in CT images with information such as symptom descriptions and medical history in text information. Then, the pre-trained multimodal model is fine-tuned, and the parameters of the multimodal model are adjusted based on the evaluation results of the validation set so that the trained multimodal model can better adapt to the task of dental disease diagnosis and improve the accuracy and reliability of diagnosis. The final trained multimodal model can comprehensively consider image and text information to provide more comprehensive and accurate results for the diagnosis of dental diseases.

[0027] Step S104: diagnose dental diseases based on the trained multimodal model.

[0028] The patient's dental CT images and text information related to the condition are input into the trained multimodal model to output the dental disease diagnosis results.

[0029] The trained multimodal model integrates the input dental CT images and text information, analyzing the condition from different perspectives. It considers the morphology, structure, and pathological characteristics of the teeth in the image, as well as the patient's symptom description and medical history in the text, ultimately outputting a diagnosis of the dental disease, providing a basis for determining the patient's condition.

[0030] In other embodiments of the present invention, a one-click upload function can be set for uploading patient dental CT images and text information about their medical condition, avoiding tedious file selection and formatting processes. When entering medical condition text, preset common symptom options are provided, and grassroots personnel only need to check to complete information entry, reducing the workload and error rate of manual input. The diagnostic process and results are displayed in a visual manner. For example, the lesion area in the CT image is presented in the form of charts, image annotations, etc., and the characteristics of the lesion and the possible type of disease are explained in plain language. For the diagnostic results, progress bars, color markings, etc. can be used to intuitively display the severity and possibility of the disease, so that grassroots personnel can quickly understand and grasp key information.

[0031] The multimodal diagnosis method for dental diseases provided in the first embodiment of the present invention uses the Mask R-CNN model and the SAM model to accurately segment the lesion area to obtain CT pathology information, and then combines the text information of the disease condition to allow the trained multimodal model to comprehensively judge the disease condition from multiple dimensions, reduce the misdiagnosis rate, make the diagnosis result more consistent with the standard pathological diagnosis, realize automated data processing and intelligent diagnosis output, generate structured diagnosis reports in a short time, reduce the time consumption of single case diagnosis, and increase the number of medical institutions receiving patients. The trained and fine-tuned multimodal model can cope with complex situations such as poor image quality and unclear patient medical history. The multimodal model after multimodal training combines text and image information to ensure stable and accurate diagnosis in different scenarios. At the same time, the technical solution provided by the embodiment of the present invention is easy to operate, lowers the threshold of grassroots diagnostic technology, and promotes the sinking of high-quality diagnostic technology. At the same time, it reduces the burden on doctors, reduces resource consumption, rationally allocates computing power, and optimizes the allocation of medical resources.

[0032] The following describes in detail the implementation of each step in the multimodal diagnosis method for dental diseases provided in Example 1.

[0033] Figure 2 As shown in FIG, a flowchart for implementing the training of a Mask R-CNN model using a labeled CT image dataset in the multimodal diagnosis method for dental diseases provided in the first embodiment of the present invention is provided. That is, in step S102, the training of the Mask R-CNN model using the labeled CT image dataset may include the following steps: Step S201: Based on the open source Mask R-CNN framework, a neural network structure for dental CT image segmentation is constructed, including a backbone network, a region proposal network (RPN), a ROI pooling layer, a fully connected layer, and a mask generation branch.

[0034] Mask R-CNN is a deep learning model used in object detection and instance segmentation. It's available in a variety of open-source frameworks, such as those based on PyTorch or TensorFlow. Choosing the right open-source framework can save development time. We used the open-source Mask R-CNN framework to build a neural network architecture for dental CT image segmentation.

[0035] The backbone network forms the foundation of Mask R-CNN. Its primary function is to extract features from the input dental CT images. Common backbone networks include ResNet and VGG. Taking ResNet as an example, through a series of convolutional layers and residual blocks, the input image is gradually converted into feature maps with varying scales and semantic information. These feature maps contain various features of the teeth and lesions in the dental CT images, providing the foundation for subsequent object detection and segmentation tasks.

[0036] The RPN is responsible for generating candidate regions (also called proposal boxes) that may contain objects on the feature map output by the backbone network. A sliding window is scanned across the feature map, predicting for each window position whether it contains an object and the corresponding bounding box coordinates. The introduction of the RPN enables Mask R-CNN to efficiently locate possible object regions in an image, reducing the computational complexity of subsequent processing.

[0037] The ROI pooling layer maps the candidate regions of varying sizes generated by the RPN onto a fixed-size feature map. Because the candidate regions generated by the RPN vary in size, and the subsequent fully connected layers and mask generation branches require fixed-size inputs, the ROI pooling layer is required for this purpose. The ROI pooling layer divides the candidate regions into several small grids and pools the features within each grid to produce a fixed-size feature vector.

[0038] The fully connected layer receives the fixed-size feature vector output by the ROI pooling layer and performs further feature extraction and classification. The fully connected layer consists of multiple neurons, which map the input feature vector to the output category score through weighted connections. In dental CT image segmentation tasks, the fully connected layer can be used to determine the type of tooth or lesion within the candidate region.

[0039] The mask generation branch generates a segmentation mask for each candidate region. It receives the feature map output by the ROI pooling layer and, through a series of convolutional layers and upsampling operations, gradually restores the spatial resolution of the feature map. It ultimately generates a binary segmentation mask of the same size as the candidate region, which is used to accurately segment the object's outline.

[0040] Step S202: Input the labeled CT image dataset into the constructed neural network structure in batches, set an appropriate learning rate, select SGD or Adam optimizer, and combine classification cross entropy loss, regression L1 / L2 loss and mask segmentation loss as loss functions for training until the constructed neural network structure converges.

[0041] In this step, the labeled CT image dataset is grouped into batches of a certain size (e.g., 4, 8, 16, etc.). The benefit of batch processing is that it can process multiple samples simultaneously in each iteration, improving training efficiency and enhancing the generalization ability of the Mask R-CNN model. In each training iteration, a batch of samples is randomly selected from the dataset and input into the constructed neural network.

[0042] The learning rate is a hyperparameter that controls the step size for model parameter updates. An appropriate learning rate is crucial for training the Mask R-CNN model. If the learning rate is too large, the Mask R-CNN model may skip the optimal solution, resulting in unstable training or even non-convergence. If the learning rate is too small, the Mask R-CNN model will train very slowly. An appropriate learning rate can usually be found by experimenting with different learning rate values or employing a learning rate scheduling strategy (such as learning rate decay).

[0043] You can choose between the stochastic gradient descent (SGD) and adaptive moment estimation (Adam) optimizers. SGD is a basic optimization algorithm that calculates the gradient of the loss function with respect to the parameters of a trained multimodal model and then updates the parameters in the opposite direction of the gradient. Adam is an optimization algorithm with an adaptive learning rate. It combines the concepts of momentum and adaptive learning rate, automatically adjusting the learning rate based on the historical gradient information of the parameters, resulting in faster convergence in many cases.

[0044] Mask R-CNN is a model for object detection and instance segmentation. Its loss function is multi-task and consists of the following three parts: The categorical cross entropy loss function is used to predict the category of each candidate region.

[0045] Typically, an L1 or L2 loss function is used to precisely adjust the bounding box position of the candidate region.

[0046] A pixel-wise cross entropy loss function is used to predict the segmentation mask for each object.

[0047] Combining the above three loss functions, we get the total loss function: , Among them, λ1 and λ2 are weight coefficients for balancing the losses of different tasks, which are used to adjust the relative importance of classification, regression and segmentation tasks in the total loss. is the classification cross entropy loss function, is the L1 or L2 loss function, is the pixel-level cross entropy loss function.

[0048] Categorical cross entropy loss: This is used to measure the accuracy of the Mask R-CNN model’s prediction of target categories. In the dental CT image segmentation task, it helps the Mask R-CNN model correctly classify different types of teeth and lesions.

[0049] Regression L1 / L2 loss: This is used to measure the accuracy of the Mask R-CNN model's prediction of the object's bounding box coordinates. The L1 loss is the sum of the absolute differences between the predicted and true values, while the L2 loss is the sum of the squares of the differences between the predicted and true values. Regression loss helps the Mask R-CNN model locate the object more accurately.

[0050] Mask segmentation loss: This measures the difference between the segmentation mask generated by the Mask R-CNN model and the ground-truth segmentation mask. Common mask segmentation loss functions include binary cross-entropy loss, which encourages the Mask R-CNN model to generate more accurate segmentation masks. In each training iteration, the Mask R-CNN model performs a forward propagation calculation based on the input batch of samples to obtain a prediction result. The total loss function is then calculated based on the prediction result and the ground-truth annotation. Next, the selected optimizer is used to update the parameters of the Mask R-CNN model based on the gradient of the loss function. This process is repeated until the loss function value of the Mask R-CNN model no longer decreases significantly, indicating that convergence has been achieved. At this point, the parameters of the Mask R-CNN model have been adjusted to a relatively optimal state and can be used for subsequent dental CT image segmentation tasks.

[0051] For example, the stochastic gradient descent (SGD) optimization algorithm is used, the initial learning rate is set to 0.001, the training batch size is 8, and the number of training rounds is 50 to achieve effective training of the Mask R-CNN model.

[0052] like Figure 3 FIG. 1 is a flowchart illustrating the implementation of loading and fine-tuning the Segment Anything (SAM) model in the multimodal diagnostic method for dental diseases provided in the first embodiment of the present invention. That is, loading and fine-tuning the Segment Anything (SAM) model in step S102 may include the following steps: Step S301: Obtain the pre-training weights of the SAM model.

[0053] This step is to obtain the pre-trained weights of the SAM model in preparation for subsequent fine-tuning of the model. This can be achieved in the following ways: the SAM model will be publicly available on open source platforms such as GitHub. Search for its official repository on these platforms. Follow the repository instructions to use command-line tools or directly download the pre-trained weight file saved in a format such as .pth from the link. Store it in a specified local directory. Verify the downloaded file based on the file hash value provided by the repository to ensure that the file is complete and usable.

[0054] Step S302: Input the annotated CT image dataset into the SAM model after loading the pre-trained weights, freeze some of the weights of the underlying general feature extraction layer, fine-tune the parameters of the high-level semantic segmentation-related layers, combine the prior segmentation knowledge of the tooth structure, adjust the learning rate, and use a lightweight optimizer for fine-tuning training.

[0055] In this step, you can use the relevant functions of a deep learning framework (such as PyTorch) in your code to load the previously downloaded pre-trained weights into the SAM model. The underlying general feature extraction layers of the SAM model typically learn common image features, such as edges and textures, which are universal across different tasks. To speed up training and reduce the risk of overfitting, you can choose to freeze some of these underlying weights so that they are not updated during training. Higher-level semantic segmentation layers are responsible for learning more specific features relevant to the target task. For dental CT image segmentation tasks, these layers need to learn features such as tooth morphology and lesion area characteristics. During training, only the parameters of these layers are updated. Since these layers have a relatively small number of parameters, the computational overhead of fine-tuning is also relatively small. Furthermore, tooth structure exhibits certain regularities and prior knowledge, such as tooth shape, position, and arrangement. This prior knowledge can be incorporated into the training process. For example, additional constraints can be added to the loss function to make the SAM model's predictions more consistent with prior knowledge of tooth structure. Alternatively, during the data preprocessing phase, the annotated CT images can be augmented based on prior knowledge, such as mirror enhancement based on tooth symmetry. The learning rate during the fine-tuning phase is typically smaller than that during the pre-training phase. Because the trained multimodal model has been pre-trained on a large dataset, its parameters are already close to optimal and only require minor adjustments. An appropriate learning rate can be found by experimenting with different learning rates or employing learning rate scheduling strategies (such as learning rate decay). Lightweight optimizers are a class of algorithms that optimize the computational resource and memory usage during model training, aiming to improve training efficiency, reduce memory usage, and maintain good model performance. Lightweight optimizers (such as AdamW) are relatively lightweight in terms of computational resource and memory usage, making them suitable for use during the fine-tuning phase. In each training iteration, the annotated CT image dataset is input into the SAM model, a loss function (such as cross-entropy loss or Dice loss) is calculated, and the optimizer is then used to update the trainable parameters based on the gradient of the loss function. Through the above methods, we can The SAM model is fine-tuned to make it better suited for the task of dental disease diagnosis.

[0056] For example: freeze 70% of the general feature extraction layer weights at the bottom of the SAM model, fine-tune the parameters of the high-level semantic segmentation-related layers, use the cross-entropy loss function, and perform 30 rounds of fine-tuning training on the dental CT image dataset to obtain the fine-tuned SAM model.

[0057] like Figure 4 As shown in FIG. 1 , a flowchart for obtaining a trained multimodal model in the multimodal diagnosis method for dental diseases provided in the first embodiment of the present invention, that is, step S103 may include the following steps: Step S401: pre-process the patient CT image dataset, the condition text information dataset, and the CT pathology information dataset according to the multimodal model input requirements.

[0058] In this step, preprocessing can include data type and format unification, data cleaning and missing value processing, and data enhancement. Data from different sources may have different data types and formats, and they need to be converted into a unified format that can be processed by the multimodal model. After that, the data set is checked for missing values, outliers, or erroneous data, and corresponding processing is performed. In order to increase the diversity of the data and the generalization ability of the multimodal model, data enhancement operations can be performed on CT images and text information. For CT images, geometric transformations such as rotation, flipping, and scaling can be performed, or the brightness and contrast of the image can be adjusted; for text information, operations such as synonym replacement and sentence reorganization can be performed.

[0059] Step S402: Process the preprocessed CT image dataset to obtain an image block dataset, and process the CT pathology information dataset and the text information dataset to obtain a text block dataset. The preprocessed CT image is divided into multiple small blocks, each of which has a fixed size. The purpose of this is to localize the image information so that the multimodal model can better capture the detailed features in the image. The CT image can be segmented using a sliding window method, and the window size and step size can be adjusted according to the specific task and multimodal model requirements. During the segmentation process, it is necessary to ensure that each image block contains sufficient information and can cover the entire CT image. For the CT pathology information dataset and the text information dataset, they are segmented into multiple text blocks according to certain rules. Segmentation can be performed based on sentences, paragraphs, or specific keywords, and each text block should have relatively independent semantic information. During the segmentation process, it is necessary to maintain the coherence and logic of the text to avoid excessive fragmentation of the segmented text blocks.

[0060] Step S403: splice the image block dataset and the text block dataset into a multimodal input sequence.

[0061] Because the image block dataset and text block dataset have different dimensions and feature representations, they need to be aligned to ensure that the data from each modality is consistent in terms of dimensions and time steps. Data processing methods such as padding and truncation can be used to ensure that each image block and text block has the same length or dimensions. The aligned image block dataset and text block dataset are concatenated in a specific order to form a multimodal input sequence. Alternating or sequential concatenation methods can be used.

[0062] Splicing the image block dataset and the text block dataset, including: Map the image block features of the image block dataset and the text block features of the text block dataset to a unified dimension to obtain image block feature projection and text block feature projection respectively; Map features of different modalities to a unified dimension D: Image patch feature projection: , , ; Text block feature projection: , , ; in, and are the original image and text block sequence respectively, and is the feature after projection, represents the set of real numbers.

[0063] Adjust the time steps of image block feature projection and text block feature projection to a unified time step; Time step alignment, adjusting the sequence length to a uniform time step T by interpolation or pooling: Image block time step adjustment: or ; Text block time step adjustment: or ; Indicates that interpolation (such as linear interpolation) adjusts the sequence length to T, Indicates that pooling (such as average pooling) compresses or expands the sequence length to T.

[0064] The image block feature projection and text block feature projection at the same time step are concatenated based on the alignment objective function.

[0065] Final alignment result The aligned features are consistent in dimension D and time step T: The alignment objective function includes constraining the feature similarity after alignment through a loss function.

[0066] The loss function constrains the feature similarity after alignment, and the contrast loss formula is as follows: , in, and are the image and text features at the t-th time step respectively.

[0067] Step S404: Based on the SWiFT training framework, load the pre-trained multimodal model weights and perform fine-tuning training using the multimodal input sequence.

[0068] In this step, first, configure the operating environment and parameters of the SWiFT efficient distributed training framework. SWiFT is an efficient distributed training framework that can fully utilize the computing resources of multiple computing devices (such as GPUs) to accelerate the training process of multimodal models. Secondly, load the weights of the pre-trained multimodal model to enable the multimodal model to have initial feature extraction and semantic understanding capabilities. Finally, use the spliced multimodal input sequence to fine-tune the multimodal model after loading the weights. Only update the parameters of the layers related to specific tasks. Select optimization algorithms such as SGD and Adam, and set appropriate hyperparameters such as learning rate and batch size to complete the fine-tuning.

[0069] Step S405: Set an objective function that is adapted to the medical diagnosis task, adjust the multimodal model hyperparameters, and obtain a trained multimodal model.

[0070] The hyperparameters of a multimodal model include training hyperparameters, architecture hyperparameters, optimizer hyperparameters, regularization hyperparameters, and data processing hyperparameters.

[0071] Training hyperparameters, including learning rate, batch size, number of training epochs, and number of gradient accumulation steps.

[0072] Architecture hyperparameters, including the number and dimensions of Transformer layers, attention mechanism parameters, and multimodal fusion methods.

[0073] Optimizer hyperparameters, including optimizer type, weight decay, and momentum.

[0074] Regularization hyperparameters, including the dropout rate.

[0075] Data processing hyperparameters, including input image resolution and audio feature parameters.

[0076] In this step, based on the characteristics and requirements of the dental disease diagnosis task, combined with indicators such as diagnostic accuracy, recall rate, F1 score, etc., cross entropy loss, mean square error loss, Dice loss, etc. are selected to construct the objective function to measure the performance of the trained multimodal model. Grid search, random search, Bayesian optimization and other methods are used to adjust hyperparameters such as learning rate, batch size, and number of training rounds. During the process, the performance of the trained multimodal model is evaluated with the help of the validation set to avoid overfitting. After setting the objective function and hyperparameters, the trained multimodal model is continuously trained. The validation set is used for evaluation regularly during training to monitor changes in performance indicators such as accuracy and recall rate. Finally, the multimodal model with the best performance is selected based on the validation set results.

[0077] Example 2: like Figure 5 FIG. 1 is a flow chart of a multimodal diagnostic method for dental diseases provided in a second embodiment of the present invention. The multimodal diagnostic method for dental diseases includes the following steps: Step S501: Acquire a patient's dental CT image dataset and a disease-related text information dataset, and pre-process the CT image dataset.

[0078] Step S502: Use the labeled CT image dataset to train the Mask R-CNN model, and load and fine-tune the SegmentAnything (SAM) model.

[0079] Step S503: Input the preprocessed CT image dataset into the trained Mask R-CNN model and the fine-tuned SAM model respectively, and output a CT pathology information dataset.

[0080] Step S504: pre-train the multimodal model using the CT pathology information dataset, the annotated CT image dataset, and the disease-related text information dataset, and fine-tune the pre-trained multimodal model to obtain a trained multimodal model.

[0081] Step S505: Use the intersection-over-union ratio and the Dice coefficient to evaluate the segmentation task of the trained multimodal model; use the accuracy, recall rate, F1 score combined with the doctor's subjective evaluation to evaluate the multimodal diagnosis task of the trained multimodal model.

[0082] The purpose of this step is to comprehensively evaluate the performance of the trained multimodal model on segmentation and multimodal diagnosis tasks. The segmentation task is evaluated using the Intersection over Union (IoU) and Dice coefficient metrics. The IoU is the ratio of the intersection and union of the predicted segmented regions to the ground-truth segmented regions. It intuitively reflects the degree of overlap between the predicted and ground-truth regions. An IoU value closer to 1 indicates better segmentation. The Dice coefficient measures the similarity between the predicted and ground-truth segments. Similarly, a value closer to 1 indicates more accurate segmentation. These two metrics can be used to quantify the accuracy of the trained multimodal model on tasks such as segmenting dental pathology. Precision, recall, and F1 score are used to quantitatively evaluate the performance of the multimodal diagnosis task. Precision refers to the proportion of correctly predicted samples by the trained multimodal model to the total number of samples; recall refers to the proportion of correctly predicted positive samples by the trained multimodal model to the actual number of positive samples. The F1 score is the harmonic mean of precision and recall, comprehensively considering the performance of both. At the same time, combined with the doctors' subjective evaluation, the doctors judge the rationality of the diagnostic interpretation output by the trained multimodal model based on their professional knowledge and clinical experience, and evaluate the diagnostic ability of the trained multimodal model from a more professional and practical application perspective.

[0083] Step S506: Obtain content that needs to be optimized for the trained multimodal model in different scenarios based on the evaluation results of the segmentation task and the evaluation results of the multimodal diagnosis task.

[0084] Based on the evaluation results of the segmentation and multimodal diagnosis tasks in step S505, analyze the performance of the trained multimodal model in different scenarios (such as disease type, image quality, and text information completeness) to identify weaknesses and areas for improvement in the trained multimodal model. For example, if the trained multimodal model has low accuracy in diagnosing certain complex disease types, or if the segmentation task's Intersection over Union (IoU) value is unsatisfactory in conditions of poor image quality, then these areas correspond to areas that require optimization. Through careful analysis of the evaluation results, it is possible to identify areas where the trained multimodal model needs further improvement, providing specific guidance for subsequent optimization.

[0085] Step S507: Optimize the trained multimodal model according to the content that needs to be optimized to obtain an optimized multimodal model.

[0086] According to the content that needs to be optimized determined in step S506, the trained multimodal model is optimized in a targeted manner. For example: if it is found that the trained multimodal model performs poorly on certain specific types of cases, the relevant labeled data can be expanded in a targeted manner to increase the learning opportunities of the trained multimodal model for these types of cases; the data enhancement strategy can be improved to generate more diverse training data and improve the generalization ability of the trained multimodal model; for inaccurate segmentation tasks, the relevant parameters of the Mask R-CNN model or the SAM model can be adjusted, such as the network structure, the convolution kernel size, etc.; for the problem of multimodal diagnostic bias, the text feature extraction method, the weight of modal fusion, etc. can be adjusted so that the trained multimodal model can better comprehensively utilize image and text information for diagnosis; or the training hyperparameters, such as the learning rate, the batch size, etc., can be reset, a more appropriate optimization algorithm can be adopted, or the number of training rounds can be increased so that the trained multimodal model can more fully learn the features in the data and continuously improve its performance to achieve better diagnostic results.

[0087] Step S508: Input the patient's dental CT image and text information related to the disease into the optimized multimodal model, and output the dental disease diagnosis result.

[0088] The multimodal diagnostic method for dental diseases provided in Example 2 of the present invention uses the intersection-over-union ratio and the Dice coefficient to evaluate the segmentation task of the trained multimodal model, and uses the accuracy, recall rate, F1 score combined with the doctor's subjective evaluation to evaluate the multimodal diagnostic task. Based on the above evaluation results, the specific content that needs to be optimized in the trained multimodal model in different scenarios is clarified, and the trained multimodal model is targeted and improved according to the optimization content obtained by analysis. In this way, the performance of the trained multimodal model can be scientifically evaluated, problems can be accurately located and continuously optimized, the accuracy and reliability of diagnosis can be improved, misdiagnosis and missed diagnosis can be reduced, the shortcomings of the trained multimodal model in different clinical scenarios can be discovered, and targeted improvements can be made to enable it to adapt to complex situations, improve generalization ability, avoid ineffective treatment and waste of resources, reduce the burden on doctors, and improve the work efficiency of medical personnel and the service capabilities of institutions.

[0089] For detailed description of this embodiment, please refer to the corresponding description in the aforementioned embodiments, which will not be repeated here.

[0090] The multimodal diagnostic method for dental diseases disclosed in this embodiment has the following advantages: First, it integrates multiple advanced models and multivariate data to improve the accuracy of lesion positioning, strengthen the ability to distinguish complex pathological conditions, combine multi-dimensional information for comprehensive judgment, improve the consistency of diagnostic results with gold standard pathological diagnosis, reduce misdiagnosis rate, and provide a reliable basis for treatment.

[0091] Second, it can realize automated data processing and intelligent diagnostic output, generate structured diagnostic reports in a short time, shorten the time required for single case diagnosis, improve diagnosis and treatment efficiency, and increase the number of patients received by medical institutions.

[0092] Third, each model has undergone rigorous training and can cope with complex clinical scenarios such as fluctuations in image quality and unclear patient expressions. Multi-dimensional verification ensures stable and accurate diagnosis and enhances clinical practicality.

[0093] Fourth, relying on a friendly interface, grassroots personnel can obtain diagnostic suggestions by following the instructions, lowering the technical threshold, promoting the penetration of high-quality diagnostic technology, and improving grassroots prevention and control capabilities.

[0094] Fifth, reduce the burden on doctors, alleviate the pressure of talent shortage, reduce resource consumption, use open source frameworks to rationally allocate computing power, optimize medical resource allocation, and improve the benefit-output ratio.

[0095] The electronic device disclosed in this embodiment includes a memory and a processor. The memory is used to store non-transitory computer-readable instructions. Specifically, the memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc.

[0096] The processor can be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device to perform desired functions. In one embodiment of the present disclosure, the processor is used to execute the computer-readable instructions stored in the memory, causing the electronic device to perform all or part of the steps of the multimodal diagnostic method for dental diseases described in the various embodiments of the present disclosure.

[0097] Those skilled in the art should understand that in order to solve the technical problem of how to obtain a good user experience, this embodiment may also include well-known structures such as a communication bus and an interface, and these well-known structures should also be included in the scope of protection of this disclosure.

[0098] like Figure 6 The present invention provides a schematic structural diagram of an electronic device according to an embodiment of the present invention, which is suitable for implementing the electronic device according to an embodiment of the present invention. Figure 6 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0099] like Figure 6As shown, an electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes based on a program stored in a read-only memory (ROM) or a program loaded from a storage device into a random access memory (RAM). The RAM also stores various programs and data required for the operation of the electronic device. The processing device, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.

[0100] Typically, the following devices can be connected to the I / O interface: input devices such as sensors or visual information acquisition devices; output devices such as display screens; storage devices such as tapes and hard disks; and communication devices. The communication device allows the electronic device to communicate with other devices (such as edge computing devices) wirelessly or by wire to exchange data. Figure 6 The electronic device is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.

[0101] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processing device, all or part of the steps of the multimodal diagnostic method for dental diseases of an embodiment of the present disclosure are performed.

[0102] For detailed description of this embodiment, please refer to the corresponding description in the aforementioned embodiments, which will not be repeated here.

[0103] The computer-readable storage medium disclosed in this embodiment stores non-transitory computer-readable instructions, which, when executed by a processor, execute all or part of the steps of the multimodal dental disease diagnosis method of each embodiment of the present disclosure.

[0104] The above-mentioned computer-readable storage media include, but are not limited to, optical storage media (e.g., CD-ROMs and DVDs), magneto-optical storage media (e.g., MOs), magnetic storage media (e.g., magnetic tapes or mobile hard disks), media with built-in rewritable non-volatile memory (e.g., memory cards), and media with built-in ROM (e.g., ROM cartridges).

[0105] For detailed description of this embodiment, please refer to the corresponding description in the aforementioned embodiments, which will not be repeated here.

[0106] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be construed as necessarily possessed by each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.

[0107] In the present disclosure, relational terms such as first and second, etc. are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. The block diagrams of the devices, devices, equipment, and systems involved in the present disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "including," "comprising," "having," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.

[0108] Additionally, as used herein, "or" used in a list of items beginning with "at least one" indicates a separate list, so that, for example, a list of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not mean that the example described is preferred or better than other examples.

[0109] It should also be noted that in the system and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.

[0110] Various changes, substitutions, and modifications may be made to the technology described herein without departing from the teachings defined by the appended claims. Moreover, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of things, means, methods, and actions described above. Currently existing or later developed processes, machines, manufactures, compositions of things, means, methods, or actions that perform substantially the same function or achieve substantially the same results as the corresponding aspects described herein may be utilized. Accordingly, the appended claims include within their scope such processes, machines, manufactures, compositions of things, means, methods, or actions.

[0111] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0112] The above description has been provided for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A multimodal diagnostic method for dental diseases, characterized in that: include: Based on the acquired patient CT image data and the corresponding condition text information data, a patient CT image dataset and a condition text information dataset are obtained; The patient's CT image data is input into the trained Mask R-CNN model to obtain the first CT pathology information; Inputting the patient's CT image data into the trained SAM model to obtain second CT pathology information, where the second CT pathology information corresponds one-to-one with the first CT pathology information; fusing the first CT pathology information and the corresponding second CT pathology information to obtain CT pathology information, and obtaining a CT pathology information dataset based on the CT pathology information; Pre-training the constructed multimodal model based on the patient CT image dataset, the condition text information dataset, and the CT pathology information dataset to obtain a trained multimodal model; Diagnose dental diseases based on the trained multimodal model.

2. The multimodal diagnostic method for dental diseases according to claim 1, characterized in that: Mask R-CNN model training includes: Based on the Mask R-CNN framework, build a neural network structure for image segmentation; The acquired training data is input into the neural network structure, the learning rate is set, the optimizer is selected, and the neural network structure is trained based on the loss function obtained by combining the classification cross entropy loss function, the regression L1 / L2 loss function, and the mask segmentation loss function until the neural network structure converges.

3. The multimodal diagnostic method for dental diseases according to claim 1, characterized in that: SAM model training, including: Obtaining pre-trained weights of the SAM model; The acquired training data was input into the SAM model after loading the pre-trained weights. The weights of the general feature extraction layer selected at the bottom layer were frozen. The parameters of the high-level semantic segmentation-related layers were adjusted. Combined with the prior segmentation knowledge of the tooth structure, the learning rate was adjusted, and a lightweight optimizer was used for SAM model training.

4. The multimodal diagnostic method for dental diseases according to claim 1, wherein: The CT pathology information dataset includes the location coordinates, size, morphological characteristics and disease classification labels of dental lesions.

5. The multimodal diagnostic method for dental diseases according to claim 1, characterized in that: The multimodal model constructed is pre-trained based on the patient CT image dataset, the condition text information dataset, and the CT pathology information dataset to obtain a trained multimodal model, including: An image block dataset is obtained based on the patient CT image dataset, and a text block dataset is obtained based on the CT pathology information dataset and the condition text information dataset; splicing the image block dataset and the text block dataset into a multimodal input sequence; Fine-tuning the constructed multimodal model based on the multimodal input sequence; The hyperparameters of the fine-tuned multimodal model after training are adjusted based on the objective function to obtain the trained multimodal model.

6. The multimodal diagnostic method for dental diseases according to claim 5, characterized in that: The step of splicing the image block dataset and the text block dataset comprises: Map the image block features of the image block dataset and the text block features of the text block dataset to a unified dimension to obtain image block feature projection and text block feature projection respectively; Adjust the time steps of image block feature projection and text block feature projection to a unified time step; The image block feature projection and text block feature projection at the same time step are concatenated based on the alignment objective function.

7. The multimodal diagnostic method for dental diseases according to claim 6, characterized in that: The alignment objective function includes constraining the feature similarity after alignment through a loss function.

8. The multimodal diagnostic method for dental diseases according to claim 1, characterized in that: After the step of diagnosing dental diseases based on the trained multimodal model, the method further includes: Evaluate the trained multimodal model on segmentation tasks and evaluate the trained multimodal model on multimodal diagnosis tasks; Based on the evaluation results of the segmentation task and the multimodal diagnosis task, the content that needs to be optimized for the trained multimodal model in different scenarios is obtained; The trained multimodal model is optimized according to the content that needs to be optimized.

9. The multimodal diagnostic method for dental diseases according to claim 8, characterized in that: Segmentation task evaluation parameters, including intersection-over-union and / or Dice coefficient; Modality diagnosis task evaluation parameters, including precision, recall, and / or F1 score.

10. An electronic device, comprising: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the multimodal diagnosis method for dental diseases according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Mask R-CNN-based gastric cancer auxiliary diagnosis system

    CN115511864A

  • Dental CBCT three-dimensional tooth rapid labeling method based on SAM large model

    CN119228995A

  • AI auxiliary decision support system for skin disease diagnosis

    CN119323574A

  • Few-sample medical image segmentation method combining CNN and SAM

    CN119741315A

  • Hip joint impact syndrome intelligent diagnosis method based on multi-modal data fusion

    CN120089335A