Multimodal diagnostic method for dental diseases and electronic device
By combining CT images and textual information about the patient's condition with a multimodal diagnostic approach, and using Mask R-CNN and SAM models to extract and fuse pathological information, a multimodal model is constructed for the diagnosis of dental diseases. This solves the problem of traditional diagnosis relying on doctors' experience and achieves more efficient and accurate diagnostic results.
Patent Information
- Application Number
- CN202510985071.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-07-17
AI Technical Summary
Traditional methods of diagnosing dental diseases rely on doctors' experience, have low accuracy, are prone to misdiagnosis and missed diagnosis, and have limited ability to detect early and subtle lesions, and cannot comprehensively consider the patient's medical history and overall health status.
A multimodal diagnostic approach was adopted, combining CT images and textual information of the patient's condition. Mask R-CNN and SAM models were used to extract and fuse pathological information, and a multimodal model was constructed for diagnosis. The trained multimodal model was then used to comprehensively analyze the image and textual information.
It improves the accuracy and efficiency of dental disease diagnosis, reduces the misdiagnosis rate, allows for a more comprehensive consideration of patient information, reduces resource consumption, and simplifies the primary care diagnostic process.
Smart Images

Figure CN120495300B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of medical treatment, and in particular relates to a multi-modal diagnosis method for dental diseases and an electronic device. BACKGROUND
[0002] In the medical field, accurate diagnosis of dental diseases is crucial for the treatment and prognosis of patients. Traditional diagnosis of dental diseases mainly relies on the subjective judgment of doctors. Doctors observe the inside of the patient's mouth by naked eye, combined with simple instrument examination, such as using a probe to probe the surface of the teeth, to preliminarily judge whether there are caries, periodontitis and other diseases. At the same time, X-ray film is a commonly used auxiliary diagnostic tool, which can help doctors view the internal structure of the teeth, such as the root, pulp cavity and other conditions, which is helpful for the diagnosis of periapical periodontitis, embedded teeth and other diseases.
[0003] However, the accuracy of traditional dental disease diagnosis method is highly dependent on the experience, professional level and state of the doctor, and there are differences in the judgment of different doctors, and the doctor is easy to misdiagnose and miss diagnosis due to fatigue and negligence, at the same time, naked eye and traditional X-ray film have limited detection ability for early and subtle lesions (such as early dentin caries, small pulp inflammation), and the disease is often serious when discovered, which delays the best treatment opportunity, at the same time, the traditional method focuses on the local situation of the teeth, and the overall disease information of the patient's medical history, general health status and other information is not fully utilized, which is crucial for diagnosis and treatment plan. SUMMARY
[0004] In view of the problems existing in the prior art, the present application provides a multi-modal diagnosis method for dental diseases and an electronic device, which at least partially solves the problem of low diagnosis efficiency in the prior art.
[0005] In a first aspect, the present disclosure provides a multi-modal diagnosis method for dental diseases, comprising:
[0006] Based on the acquired patient CT image data and corresponding disease text information data, a patient CT image data set and a disease text information data set are obtained;
[0007] The patient CT image data is input into the trained Mask R-CNN model to obtain first CT pathological information;
[0008] The patient CT image data is input into the trained SAM model to obtain second CT pathological information, which corresponds one-to-one with the first CT pathological information;
[0009] The first CT pathological information and the corresponding second CT pathological information are fused to obtain CT pathological information, and a CT pathological information data set is obtained based on the CT pathological information;
[0010] Pre-train the multi-modal model based on the CT image data set of the patient, the disease text information data set and the CT pathological information data set to obtain the trained multi-modal model;
[0011] Diagnose the dental disease based on the trained multi-modal model.
[0012] In a second aspect, the present disclosure also provides an electronic device, which comprises:
[0013] at least one processor; and,
[0014] a memory in communication with the at least one processor; wherein,
[0015] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the multi-modal diagnosis method of the dental disease according to any one of the first aspect.
[0016] The present application provides a multi-modal diagnosis method of dental disease and an electronic device. The multi-modal diagnosis method of dental disease is used to diagnose the CT of the patient by training the multi-modal model, replacing the existing artificial diagnosis, so as to improve the diagnosis and treatment efficiency. In the training of the multi-modal model, the disease text information data set and the CT pathological information data set obtained based on the Mask R-CNN model and the SAM model are used. The information obtained by the Mask R-CNN model and the SAM model is more accurate and comprehensive, so as to improve the diagnosis and treatment accuracy and avoid misdiagnosis. BRIEF DESCRIPTION OF DRAWINGS
[0017] The above and other objects, features and advantages of the present disclosure will become more apparent from the following detailed description taken in conjunction with the accompanying drawings, in which like reference characters refer to the like elements throughout the different views. The following detailed description is presented primarily for the purpose of enabling a person skilled in the art to make and use the disclosure and the best mode and / or embodiments to support the claims.
[0018] Figure 1 A flowchart of the multi-modal diagnosis method of dental disease provided by the first embodiment of the present disclosure;
[0019] Figure 2 A flowchart of training the Mask R-CNN model with the labeled CT image data set in the multi-modal diagnosis method of dental disease provided by the first embodiment of the present disclosure;
[0020] Figure 3 A flowchart of loading and fine-tuning the SAM model in the multi-modal diagnosis method of dental disease provided by the first embodiment of the present disclosure;
[0021] Figure 4A flowchart of a process of obtaining a trained multi-modal model in the multi-modal diagnosis method of dental diseases provided by Embodiment One of the present disclosure is provided.
[0022] Figure 5 A flowchart of the multi-modal diagnosis method of dental diseases provided by Embodiment Two of the present disclosure is provided.
[0023] Figure 6 A principle block diagram of an electronic device provided by the present disclosure is provided. DETAILED DESCRIPTION
[0024] Embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.
[0025] It should be apparent that the following describes embodiments of this disclosure by way of specific examples, and that one of ordinary skill in the art will readily understand other advantages and benefits of the present disclosure from this disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, and are not all the embodiments. The present disclosure can also be implemented or applied by other different specific embodiments, and the details in the specification can be modified or changed based on different views and applications without departing from the spirit of the present disclosure. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict. Based on the embodiments in the present disclosure, all other embodiments obtained by one of ordinary skill in the art without creative labor are within the scope of protection of the present disclosure.
[0026] It should be noted that the various aspects of the embodiments described below are within the scope of the appended claims. It should be apparent that the aspects described herein can be embodied in a wide variety of forms and that any specific structure and / or function described herein is merely illustrative. Based on the present disclosure, one of ordinary skill in the art will contemplate other aspects in which the aspects described herein can be implemented independently of and without the independent method of practicing the present disclosure. For example, an apparatus and / or method can be implemented using any number of the aspects set forth herein. In addition, this apparatus and / or method can be implemented using other structures and / or functionalities in addition to or other than one or more of the aspects set forth herein.
[0027] It should also be noted that the diagrams in the following embodiments are only schematically illustrating the basic concepts of the present disclosure, and only show the components related to the present disclosure, not the number, shape and size of the components when actually implemented. The actual implementation of each component can be a random change in shape, number and proportion, and the layout of the components can also be more complex.
[0028] Also in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, one skilled in the relevant art will appreciate that the described aspects can be practiced without these specific details.
[0029] Embodiment one:
[0030] For ease of understanding, as Figure 1 shown, the embodiment discloses a multi-modal diagnosis method for dental diseases, comprising:
[0031] Step S101: based on the acquired patient CT image data and corresponding disease text information data, obtaining a patient CT image data set and a disease text information data set;
[0032] Collect the patient's dental CT image through professional CT scanning equipment. These images can include dental CT images of patients of different ages and different genders from multiple medical institutions, which can clearly present the internal structure of the teeth and their surrounding tissues, providing intuitive image basis for subsequent lesion analysis. On the other hand, collect disease text information, which can be chief complaints, present illness history, past medical history, etc. of patients of different ages and different genders from multiple medical institutions. These text contents contain key information such as disease development process and symptom manifestation, which can assist in understanding and diagnosing images. After collecting the CT image data set, the CT image data set can be preprocessed. The purpose of preprocessing the CT image data set is to improve the image quality and usability. Common preprocessing operations include adjusting the contrast and brightness of the image to make the features in the image clearer and more identifiable; unifying the format of the image to facilitate subsequent processing of the trained multi-modal model; normalizing the image to map the image pixel value to a specific range, so that the trained multi-modal model can more stably learn the image features. The collected CT pathological information data set includes the position coordinates, size, morphological features and disease classification labels of the dental lesions.
[0033] Step S102: input the patient CT image data into the trained Mask R-CNN model to obtain first CT pathological information; input the patient CT image data into the trained SAM model to obtain second CT pathological information, which corresponds one-to-one with the first CT pathological information; fuse the first CT pathological information and the corresponding second CT pathological information to obtain CT pathological information, and obtain a CT pathological information data set based on the CT pathological information;
[0034] Optionally, the Mask R-CNN model training comprises:
[0035] Based on the Mask R-CNN framework, a neural network structure for image segmentation is built.
[0036] The obtained training data is input into the neural network structure, a learning rate is set, an optimizer is selected, and the neural network structure is trained based on a loss function obtained by combining a classification cross-entropy loss function, a regression L1 / L2 loss function, and a mask segmentation loss function until the neural network structure converges.
[0037] Optionally, the SAM model training comprises:
[0038] Pre-training weights of the SAM model are obtained.
[0039] The obtained training data is input into the SAM model loaded with the pre-training weights, the weights of the selected general feature extraction layer at the bottom are frozen, the parameters of the high-level semantic segmentation related layer are adjusted, the prior segmentation knowledge of the dental structure is combined, the learning rate is adjusted, and the SAM model is trained by using a lightweight optimizer.
[0040] When the Mask R-CNN model and the SAM model are trained, first, training data is constructed, the training data is obtained by labeling CT image data in a part of the CT image data set of the patient, and a labeled CT image data set is obtained, and the training data is constructed based on the labeled CT image data set.
[0041] The Mask R-CNN model is a powerful instance segmentation model, which is trained by using the labeled CT image data set. The trained multi-modal model can learn the features and boundary information of different lesion regions (such as dental caries and periapical periodontitis). The trained multi-modal model adjusts its own parameters by continuously comparing the predicted result data and the labeled CT image data, so as to gradually have the ability to accurately segment the lesion region. The Segment Anything (SAM) model itself has strong general segmentation ability. In the dental diagnosis scene, the pre-training weights of the SAM model are loaded, and then the SAM model is fine-tuned according to the characteristics of the dental image. During the fine-tuning process, the weights of part of the general feature extraction layer at the bottom can be frozen, and only the parameters of the high-level semantic segmentation related layer are adjusted. At the same time, the prior knowledge in the dental field is combined, so that the fine-tuned SAM model is more focused on the learning of the dental lesion features.
[0042] The pre-processed CT images are input into the trained Mask R-CNN model and the fine-tuned SAM model respectively. The Mask R-CNN model will locate and segment the lesion area in the CT image according to the knowledge learned before, and output the position, size, shape and other information of the lesion; the SAM model will further supplement and refine the lesion information by virtue of its ability to capture dental micro-lesions after fine-tuning. The output results of the Mask R-CNN model and the fine-tuned SAM model are integrated to form a CT pathology information dataset, which records the specific conditions of the lesions in each CT image in detail, including the position coordinates, size dimensions, morphological features and disease classification labels of the lesions, providing important image feature basis for subsequent multi-modal diagnosis.
[0043] The output results of the Mask R-CNN model and the fine-tuned SAM model can be integrated in the following ways:
[0044] First, the output results of the two trained models are aligned in data, including spatial position alignment and slice correspondence. Spatial position alignment can ensure that the lesion areas output by the two trained models for the same CT image correspond to each other in spatial position, and slice correspondence is used for multi-layer CT images to match the output results of the two trained models on each layer slice. The output of the two trained models can be sorted and corresponding according to the slice number, so that the information about the lesions of the two trained models can be accurately associated on the same slice.
[0045] The spatial position alignment formula is:
[0046] ,
[0047] wherein, is the original spatial position coordinate, is the transformed position coordinate, is the radial basis function, is the control point, the parameters to be estimated, a and b are the data sets to be aligned, and a and b are matrices.
[0048] Slice correspondence is to determine the correspondence by calculating the similarity between slices. For example, using cosine similarity or Euclidean distance and other measurement methods, find the point pair with the highest similarity in the two slices as the corresponding point.
[0049] Second step, the suspected lesion area detected by the two trained models is fused, and a multi-data voting method or a weighted fusion method can be used. If there is an overlapping or similar part in the suspected lesion area detected by the two trained models, a majority voting method can be used to determine the final lesion area, or different weights can be assigned according to the performance of the two trained models in detecting different types of lesions. For example, if the Mask R-CNN model performs better in detecting large lesion areas, and the SAM model is more advantageous in capturing subtle lesions, when integrating, the output result of the Mask R-CNN model can be given a higher weight for large lesions, and the output result of the SAM model can be given a higher weight for subtle lesions. Then the output results of the two trained models are linearly combined according to the weights to obtain the final lesion area.
[0050] Third step, the feature information about the lesion output by the two trained models is spliced. The Mask R-CNN model may output the position, size, shape and other geometric features of the lesion, and the SAM model may extract the texture, gray scale and other features of the lesion. These features are spliced into a longer feature vector as the comprehensive feature information of the lesion. Then the features output by the two trained models are screened and fused to remove redundant features and retain the most representative features. Feature selection algorithms such as correlation analysis and principal component analysis can be used to evaluate the importance of each feature to the classification or diagnosis of the lesion, and the features with higher importance are selected for fusion. Then the screened features are combined to obtain a more concise and effective feature representation.
[0051] Fourth step, the output results of the two trained models are verified and corrected. Professional dentists can manually verify the integrated CT pathological information dataset, or compare it with known gold standard data (such as pathological examination results) for verification. According to the verification result, the known gold standard data (such as pathological examination results) are compared and verified.
[0052] Step S103: Pre-training the constructed multi-modal model based on the patient CT image dataset, the disease text information dataset and the CT pathological information dataset to obtain a trained multi-modal model;
[0053] First, the constructed multi-modal model is pre-trained using the CT pathology information dataset, the patient CT image dataset, and the disease text information dataset. During the pre-training process, the multi-modal model learns how to extract useful features from data in different modalities and attempts to fuse these features. For example, the multi-modal model needs to learn how to associate lesion features in CT images with symptom descriptions, medical history, and other information in text information. Then, the pre-trained multi-modal model is fine-tuned, and according to the evaluation results of the validation set, the parameters of the multi-modal model are adjusted so that the multi-modal model can better adapt to the task of dental disease diagnosis and improve the accuracy and reliability of diagnosis. The final trained multi-modal model can comprehensively consider image and text information to provide more comprehensive and accurate results for dental disease diagnosis.
[0054] Step S104: Diagnose dental diseases based on the trained multi-modal model.
[0055] The patient's dental CT image and disease text information are input into the trained multi-modal model, and the dental disease diagnosis result is output.
[0056] The trained multi-modal model integrates the input dental CT image and text information and analyzes the disease from different dimensions. By comprehensively considering the morphology, structure, and lesion characteristics of the teeth in the image and the patient's symptom description, medical history, and other content in the text information, the final diagnosis result about the dental disease is output, providing a basis for the patient's disease judgment.
[0057] In some embodiments of the present application, one-key upload function can be set for uploading patient dental CT images and disease text information, avoiding the cumbersome file selection and format setting process. When inputting disease text, preset common symptom options are provided, and grassroots personnel only need to check them to complete information input, reducing the workload and error rate of manual input. The diagnosis process and result are displayed in a visual manner. For example, the lesion area in the CT image is presented in the form of charts and image annotations, and the characteristics of the lesion and possible disease types are explained in simple and easy-to-understand words. For the diagnosis result, progress bar, color identification, and other ways can be used to intuitively display the severity and possibility of the disease, so that grassroots personnel can quickly understand and grasp the key information.
[0058] The multi-modal diagnosis method for dental diseases provided by the embodiment one of the present application can accurately segment the lesion area through the Mask R-CNN model and the SAM model to obtain CT pathological information, and then combine the disease text information, so that the trained multi-modal model can comprehensively judge the disease from multiple dimensions, reduce the misdiagnosis rate, make the diagnosis result more consistent with the standard pathological diagnosis, realize automatic data processing and intelligent diagnosis output, generate a structured diagnosis report in a short time, reduce the time consumption of single case diagnosis, improve the reception capacity of medical institutions, and the trained multi-modal model can cope with complex situations such as poor image quality and unclear patient history expression. The multi-modal model trained in combination with text and image information can ensure stable and accurate diagnosis in different scenarios. Meanwhile, the technical solution provided by the embodiment of the present application is simple to operate, reduces the technical threshold of basic diagnosis, and promotes the sinking of high-quality diagnosis technology. At the same time, it reduces the burden of doctors, reduces resource consumption, reasonably allocates computing power, and optimizes the allocation of medical resources.
[0059] The implementation of each step in the multi-modal diagnosis method for dental diseases provided by the embodiment one will be described in detail below.
[0060] Figure 2 As shown in the implementation flowchart of the multi-modal diagnosis method for dental diseases provided by the embodiment one of the present application, the training of the Mask R-CNN model with the labeled CT image dataset can include the following steps:
[0061] Step S201, based on the open source Mask R-CNN framework, a neural network structure for dental CT image segmentation is built, including a backbone network, a region proposal network (RPN), an ROI pooling layer, a fully connected layer, and a mask generation branch.
[0062] Mask R-CNN is a deep learning model applied in the field of target detection and instance segmentation, and there are many open source frameworks, such as versions based on PyTorch or TensorFlow. Selecting a suitable open source framework can save development time. Based on the open source Mask R-CNN framework, a neural network structure for dental CT image segmentation is built.
[0063] The backbone network is the basis of Mask R-CNN, and its main function is to extract features from the input dental CT image. Common backbone networks include ResNet, VGG, etc. Taking ResNet as an example, through a series of convolutional layers, residual blocks and other structures, the input image is gradually converted into feature maps with different scales and semantic information. These feature maps contain various features of teeth, lesion areas, etc. in dental CT images, providing a basis for subsequent target detection and segmentation tasks.
[0064] The RPN is responsible for generating candidate regions (also known as proposal boxes) that may contain targets on the feature map output by the backbone network. By sliding windows on the feature map, it predicts whether each window position contains a target and the corresponding bounding box coordinates. The introduction of RPN enables Mask R-CNN to efficiently locate possible target regions in the image, reducing the computational load of subsequent processing.
[0065] The role of the ROI pooling layer is to map the different sizes of candidate regions generated by the RPN to a fixed size feature map. Since the candidate regions generated by the RPN are of different sizes, and the subsequent fully connected layer and mask generation branch require fixed size inputs, the ROI pooling layer is needed for processing. The ROI pooling layer divides the candidate region into several small grids and performs pooling operations on the features within each grid to obtain a fixed size feature vector.
[0066] The fully connected layer receives the fixed size feature vector output by the ROI pooling layer and performs further feature extraction and classification. The fully connected layer consists of multiple neurons that map the input feature vector to the output class score through weight connections. In the dental CT image segmentation task, the fully connected layer can be used to determine which type of tooth or lesion the target in the candidate region belongs to.
[0067] The mask generation branch is used to generate a corresponding segmentation mask for each candidate region. It receives the feature map output by the ROI pooling layer and gradually restores the spatial resolution of the feature map through a series of convolutional layers and upsampling operations. Finally, it generates a binary segmentation mask of the same size as the candidate region, which is used to accurately segment the target outline.
[0068] Step S202: Input the labeled CT image dataset into the built neural network structure in batch form, set appropriate learning rate, select SGD or Adam optimizer, combine classification cross-entropy loss, regression L1 / L2 loss and mask segmentation loss as loss function for training until the built neural network structure converges.
[0069] In this step, the labeled CT image dataset is grouped according to a certain batch size (such as 4, 8, 16, etc.). The advantage of batch processing is that multiple samples can be processed simultaneously in each iteration, improving training efficiency and helping the generalization ability of the Mask R-CNN model. In each training iteration, a batch of samples is randomly selected from the dataset and input into the built neural network.
[0070] The learning rate is a hyperparameter that controls the step size of the model parameter updates. A suitable learning rate is crucial for the training of the Mask R-CNN model. If the learning rate is too large, the Mask R-CNN model may skip the optimal solution, leading to unstable training or even failure to converge. If the learning rate is too small, the training speed of the Mask R-CNN model will be very slow. Generally, you can find a suitable learning rate by experimenting with different learning rate values or using a learning rate scheduling strategy (such as learning rate decay).
[0071] You can choose the Stochastic Gradient Descent (SGD) or Adaptive Moment Estimation (Adam) optimizer. SGD is a basic optimization algorithm that calculates the gradient of the loss function with respect to the trained multi-modal model parameters, and then updates the parameters in the opposite direction of the gradient. Adam is an adaptive learning rate optimization algorithm that combines the ideas of momentum and adaptive learning rate, and can automatically adjust the learning rate according to the historical gradient information of the parameters, which can converge faster in many cases.
[0072] Mask R-CNN is a model for object detection and instance segmentation, and its loss function is multi-task, including the following three parts:
[0073] The classification cross-entropy loss function is used to predict the class of each candidate region.
[0074] The L1 or L2 loss function is usually used to fine-tune the position of the bounding box of the candidate region.
[0075] The pixel-level cross-entropy loss function is used to predict the segmentation mask of each target.
[0076] Combine the above three parts of the loss function to get the total loss function:
[0077] ,
[0078] where λ1 and λ2 are the weight coefficients that balance the loss of different tasks, used to adjust the relative importance of classification, regression and segmentation tasks in the total loss. The classification cross-entropy loss function is The L1 or L2 loss function is The pixel-level cross-entropy loss function is
[0079] Classification cross-entropy loss: used to measure the accuracy of the Mask R-CNN model in predicting the class of the target. In the dental CT image segmentation task, it can help the Mask R-CNN model correctly classify different types of teeth and lesions.
[0080] Regression L1 / L2 loss: used to measure the accuracy of the Mask R-CNN model in predicting the coordinates of the target bounding box. L1 loss is the sum of the absolute values of the differences between the predicted values and the true values, and L2 loss is the sum of the squares of the differences between the predicted values and the true values. Regression loss can help the Mask R-CNN model more accurately locate the position of the target.
[0081] Mask segmentation loss: used to measure the difference between the segmentation mask generated by the Mask R-CNN model and the true segmentation mask. Common mask segmentation loss functions include binary cross-entropy loss, etc., which can encourage the Mask R-CNN model to generate more accurate segmentation masks. In each training iteration, the Mask R-CNN model performs forward propagation calculations based on the input batch of samples to obtain the predicted results. Then, the total loss function value is calculated based on the predicted results and the true labels. Then, using the selected optimizer, update the parameters of the Mask R-CNN model according to the gradient of the loss function. Repeat this process until the loss function value of the Mask R-CNN model no longer decreases significantly, i.e., it reaches a state of convergence. At this point, the parameters of the Mask R-CNN model have been adjusted to a relatively optimal state and can be used for subsequent dental CT image segmentation tasks.
[0082] For example: using the stochastic gradient descent (SGD) optimization algorithm, setting the initial learning rate to 0.001, the training batch size to 8, and the training number of rounds to 50, to effectively train the Mask R-CNN model.
[0083] As shown in Figure 3 , is the implementation flowchart of loading and fine-tuning the Segment Anything (SAM) model in the multi-modal diagnosis method of the dental disease provided by the embodiment one of the present application, that is, loading and fine-tuning the Segment Anything (SAM) model in step S102 can include the following steps:
[0084] Step S301, obtaining the pre-training weight of the SAM model.
[0085] This step is to obtain the pre-training weight of the SAM model, which prepares for subsequent model fine-tuning, which can be implemented in the following way: the SAM model will be disclosed on open source platforms such as GitHub, search its official repository on these platforms, follow the repository instructions, download the pre-training weight file saved in.pth format using the command line tool or directly from the link, and store it in the local specified directory. According to the file hash value provided by the repository, verify the downloaded file to ensure that the file is complete and usable.
[0086] Step S302, input the labeled CT image dataset into the SAM model loaded with pre-trained weights, freeze part of the bottom layer general feature extraction layer weights, fine-tune the high layer semantic segmentation related layer parameters, combine the prior segmentation knowledge of tooth structure, adjust the learning rate, and use a lightweight optimizer for fine-tuning training.
[0087] In this step, the pre-trained weights downloaded previously can be loaded into the SAM model using the relevant functions of the deep learning framework (such as PyTorch) in the code. The bottom layer general feature extraction layer of the SAM model usually learns some general image features such as edges, textures, etc. These features have certain universality in different tasks. To speed up the training and reduce the risk of overfitting, part of the bottom layer weights can be frozen so that they do not update during training. The high layer semantic segmentation related layer is responsible for learning more specific features related to the target task. For dental CT image segmentation tasks, these layers need to learn the morphology of teeth, the features of lesion areas, etc. Only the parameters of these layers are updated during training. Since the number of parameters in these layers is relatively small, the amount of computation for fine-tuning is also relatively small. At the same time, the tooth structure has certain regularity and prior knowledge, such as the shape, position, and arrangement of teeth. These prior knowledge can be incorporated into the training process. For example, an additional constraint term can be added to the loss function to make the SAM model's prediction results more consistent with the prior knowledge of tooth structure; or in the data preprocessing stage, some enhancement operations based on prior knowledge can be performed on the labeled CT image, such as mirror enhancement based on the symmetry of teeth. The learning rate in the fine-tuning stage is usually smaller than in the pre-training stage. Because the trained multi-modal model has been pre-trained on a large-scale dataset, the parameters are close to the optimal value, and only a small amount of adjustment is needed. Different learning rate values can be tested, or a learning rate scheduling strategy (such as learning rate decay) can be used to find the appropriate learning rate. Lightweight optimizers are a class of algorithms optimized for computational resource consumption and memory usage during model training, aiming to improve training efficiency, reduce memory usage, and maintain good model performance. Lightweight optimizers (such as AdamW) have relatively small computational resource consumption and memory usage, making them suitable for use in the fine-tuning stage. In each training iteration, the labeled CT image dataset is input into the SAM model, the loss function (such as cross-entropy loss, Dice loss, etc.) is calculated, and then the optimizer updates the trainable parameters according to the gradient of the loss function. Through the above methods, the SAM model can be fine-tuned on the dental CT image dataset
[0088] The above fine-tunes the SAM model to better adapt to the task of dental disease diagnosis.
[0089] For example: freeze the general feature extraction layer weight of the bottom 70% of the SAM model, fine-tune the high-level semantic segmentation related layer parameters, use the cross-entropy loss function, and perform 30 rounds of fine-tuning training on the dental CT image dataset to obtain the fine-tuned SAM model.
[0090] As shown in Figure 4 The implementation flowchart of obtaining the trained multi-modal model in the multi-modal diagnosis method for dental diseases provided by the first embodiment of the application, that is, step S103 can include the following steps:
[0091] Step S401, pre-process the patient CT image dataset, disease text information dataset and CT pathology information dataset according to the multi-modal model input requirements.
[0092] In this step, the pre-processing can include data type and format unification, data cleaning and missing value processing, and data enhancement processing. Different sources of data can have different data types and formats, which need to be converted into a unified format that can be processed by the multi-modal model. Then check whether there are missing values, outliers or error data in the dataset, and perform corresponding processing. In order to increase the diversity of data and the generalization ability of the multi-modal model, data enhancement operations can be performed on the CT image and the text information. For the CT image, geometric transformations such as rotation, flipping and scaling can be performed, or the brightness and contrast of the image can be adjusted. For text information, synonym replacement, sentence reorganization and other operations can be performed.
[0093] Step S402, process the pre-processed CT image dataset to obtain an image block dataset, and process the CT pathology information dataset and the text information dataset to obtain a text block dataset. The pre-processed CT image is divided into multiple small blocks, each small block has a fixed size. The purpose of this is to localize the image information for better capturing of detailed features in the image by the multi-modal model. The sliding window method can be used to divide the CT image, and the size and step length of the window can be adjusted according to the specific task and the requirements of the multi-modal model. During the division process, it is necessary to ensure that each image block contains sufficient information and can cover the entire CT image. For the CT pathology information dataset and the text information dataset, they are divided into multiple text blocks according to certain rules. The division can be based on sentences, paragraphs or specific keywords, and each text block should have relatively independent semantic information. During the division process, attention should be paid to maintaining the coherence and logicality of the text to avoid excessive fragmentation of the divided text blocks.
[0094] Step S403, splice the image block dataset and the text block dataset into a multi-modal input sequence.
[0095] Since the image block dataset and the text block dataset have different dimensions and feature representations, they need to be aligned so that the modal data remains consistent in dimension and time step. Methods such as padding, truncation, etc. can be used to process the data to ensure that each image block and text block has the same length or dimension. The aligned image block dataset and text block dataset are spliced into a multi-modal input sequence in a certain order. Alternating splicing, sequential splicing, etc. can be used.
[0096] Splicing the image block dataset and the text block dataset includes:
[0097] Mapping the image block features of the image block dataset and the text block features of the text block dataset to a unified dimension to obtain image block feature projections and text block feature projections, respectively;
[0098] Mapping the features of different modalities to a unified dimension D:
[0099] Image block feature projection: , , ;
[0100] Text block feature projection: , , ;
[0101] wherein, and are the original image and text block sequences, and are the projected features, denotes the set of real numbers.
[0102] Adjusting the image block feature projection and the text block feature projection time steps to a unified time step;
[0103] Time step alignment, adjusting the sequence length to a unified time step T by interpolation or pooling:
[0104] Image block time step adjustment: or ;
[0105] Text block time step adjustment: or ;
[0106] denotes interpolation (such as linear interpolation) to adjust the sequence length to T, denotes pooling (such as average pooling) to compress or expand the sequence length to T.
[0107] Splicing the image block feature projection and the text block feature projection of the unified time step based on the alignment objective function.
[0108] Final alignment result
[0109] Aligned features are consistent in dimension D and time step T:
[0110] Aligning objective function, including constraining similarity of aligned features by loss function.
[0111] In the step of constraining similarity of aligned features by loss function, the contrastive loss formula is as follows:
[0112] ,
[0113] Wherein, and are the image and text features of the t-th time step.
[0114] Step S404, based on the SWiFT training framework, load the pre-trained multi-modal model weight, and fine-tune training using the multi-modal input sequence.
[0115] In this step, first, configure the running environment and parameters of the SWiFT efficient distributed training framework. SWiFT is an efficient distributed training framework that can fully utilize the computing resources of multiple computing devices (such as GPUs) to accelerate the training process of multi-modal models. Secondly, load the pre-trained weights of the multi-modal model, so that the multi-modal model has initial feature extraction and semantic understanding capabilities. Finally, fine-tune the multi-modal model loaded with weights using the spliced multi-modal input sequence, only update the parameters of the layers related to the specific task, select optimization algorithms such as SGD, Adam, etc., and set appropriate learning rate, batch size, etc. Hyperparameters to complete fine-tuning.
[0116] Step S405, set the target function adapted to the medical diagnosis task, adjust the multi-modal model hyperparameters, and obtain the trained multi-modal model.
[0117] The hyperparameters of the multi-modal model include training hyperparameters, architecture hyperparameters, optimizer hyperparameters, regularization hyperparameters, and data processing hyperparameters.
[0118] Training hyperparameters include learning rate, batch size, training rounds, and gradient accumulation steps.
[0119] Architecture hyperparameters include the number of Transformer layers and dimensions, attention mechanism parameters, and multi-modal fusion methods.
[0120] Optimizer hyperparameters include optimizer type, weight decay, and momentum.
[0121] Regularization hyperparameters include dropout rate.
[0122] Data processing hyperparameters, including input image resolution and audio feature parameters.
[0123] In this step, according to the characteristics and requirements of the dental disease diagnosis task, combined with the indexes such as diagnosis accuracy, recall rate and F1 score, the objective function is constructed by selecting cross-entropy loss, mean square error loss and Dice loss, so as to measure the performance of the trained multi-modal model. Grid search, random search and Bayesian optimization are used to adjust the learning rate, batch size, training round and other hyperparameters. During the process, the performance of the trained multi-modal model is evaluated by the validation set to avoid overfitting. After setting the objective function and hyperparameters, the trained multi-modal model is continuously trained, and the performance indicators such as accuracy and recall rate are monitored during the training. Finally, the multi-modal model with the best performance is selected according to the validation set results.
[0124] Embodiment two
[0125] As Figure 5 shown is a flowchart of a multi-modal diagnosis method for dental diseases provided by Embodiment Two of the present application. The multi-modal diagnosis method for dental diseases comprises the following steps:
[0126] Step S501, collect the patient's dental CT image data set and the disease-related text information data set, and preprocess the CT image data set.
[0127] Step S502, train the Mask R-CNN model with the labeled CT image data set, load and fine-tune the SegmentAnything (SAM) model.
[0128] Step S503, input the preprocessed CT image data set into the trained Mask R-CNN model and the fine-tuned SAM model respectively, and output the CT pathological information data set.
[0129] Step S504, pre-train the multi-modal model with the CT pathological information data set, the labeled CT image data set and the disease-related text information data set, and fine-tune the pre-trained multi-modal model to obtain the trained multi-modal model.
[0130] Step S505, evaluate the segmentation task of the trained multi-modal model using the intersection over union and Dice coefficient; evaluate the multi-modal diagnosis task of the trained multi-modal model using the accuracy, recall rate, F1 score and subjective evaluation of doctors.
[0131] The purpose of this step is to comprehensively evaluate the performance of the trained multi-modal model on the segmentation task and the multi-modal diagnosis task. The intersection over union (IoU) and Dice coefficient are used to evaluate the segmentation task. The intersection over union is the ratio of the intersection area of the predicted segmentation region and the real segmentation region to the union area, which can intuitively reflect the overlap between the predicted region and the real region. The closer the IoU value is to 1, the better the segmentation effect. The Dice coefficient measures the similarity between the predicted segmentation region and the real segmentation region. Similarly, the closer the value is to 1, the more accurate the segmentation result. Through these two indicators, the accuracy of the trained multi-modal model in tasks such as segmenting tooth pathological sites can be quantified. The accuracy, recall rate and F1 score are used to quantitatively evaluate the multi-modal diagnosis task. The accuracy rate is the proportion of the number of samples correctly predicted by the trained multi-modal model to the total number of samples. The recall rate is the proportion of the number of positive samples correctly predicted by the trained multi-modal model to the actual number of positive samples. The F1 score is the harmonic mean of accuracy and recall, which considers the performance of both. At the same time, combined with the subjective evaluation of doctors, doctors can judge the rationality of the diagnosis explanation output by the trained multi-modal model based on their professional knowledge and clinical experience, and evaluate the diagnosis ability of the trained multi-modal model from a more professional and practical application perspective.
[0132] Step S506, according to the evaluation results of the segmentation task and the evaluation results of the multi-modal diagnosis task, obtaining the content that needs to be optimized of the trained multi-modal model in different scenarios.
[0133] According to the evaluation results of the segmentation task and the multi-modal diagnosis task in step S505, analyze the performance of the trained multi-modal model in different scenarios (such as different disease types, image quality, text information completeness, etc.), find out the weak links and places that need to be improved of the trained multi-modal model. For example, if the accuracy of the trained multi-modal model is low in the diagnosis of some complex disease types, or the IoU value of the segmentation task is not ideal in the case of poor image quality, then these corresponding parts are the content that needs to be optimized. Through detailed analysis of the evaluation results, it is clear that the trained multi-modal model needs to be further improved in which aspects, providing specific direction for subsequent optimization.
[0134] Step S507, according to the content that needs to be optimized, optimizing the trained multi-modal model to obtain an optimized multi-modal model.
[0135] According to the content determined to be optimized in step S506, the trained multi-modal model is optimized in a targeted manner. For example, if it is found that the trained multi-modal model performs poorly on certain types of cases, the relevant labeled data can be expanded to increase the learning opportunities of the trained multi-modal model for these types of cases; the data augmentation strategy can be improved to generate more diverse training data and improve the generalization ability of the trained multi-modal model; for inaccurate segmentation tasks, the relevant parameters of the Mask R-CNN model or the SAM model can be adjusted, such as network structure, convolution kernel size, etc.; for multi-modal diagnosis bias problems, the text feature extraction method and the weight of modal fusion can be adjusted to enable the trained multi-modal model to better utilize image and text information for diagnosis; or the training hyperparameters such as learning rate, batch size, etc. are reset, a more appropriate optimization algorithm is used, or the number of training rounds is increased, so that the trained multi-modal model can learn the features in the data more fully and continuously improve its performance to achieve better diagnostic results.
[0136] Step S508: inputting the dental CT image of the patient and the text information related to the disease into the optimized multi-modal model to output the dental disease diagnosis result.
[0137] The multi-modal diagnosis method for dental diseases provided in Embodiment Two of the present application can evaluate the segmentation task of the trained multi-modal model by using the intersection over union and the Dice coefficient, evaluate the multi-modal diagnosis task by combining the accuracy, recall rate, F1 score with the subjective evaluation of doctors, and according to the above evaluation results, determine the specific content of the trained multi-modal model that needs to be optimized in different scenarios, and make targeted adjustments and improvements to the trained multi-modal model according to the analyzed optimization content. Thereby, the performance of the trained multi-modal model can be scientifically evaluated, the problems can be accurately located and continuously optimized, the diagnostic accuracy and reliability can be improved, the misdiagnosis and missed diagnosis can be reduced, the short board of the trained multi-modal model in different clinical scenarios can be found, and targeted improvements can be made to adapt to complex situations, improve the generalization ability, avoid ineffective treatment and resource waste, reduce the burden on doctors, and improve the work efficiency of medical personnel and the service ability of institutions.
[0138] The detailed description of the present embodiment can be referred to the corresponding description in the foregoing embodiments, which will not be repeated here.
[0139] The multi-modal diagnosis method for dental diseases disclosed in the present embodiment has the following advantages:
[0140] Firstly, a variety of advanced models and multi-element data are fused to improve the lesion positioning accuracy, strengthen the ability to distinguish complex pathological conditions, combine multi-dimensional information for comprehensive judgment, improve the consistency of the diagnostic results with the gold standard pathological diagnosis, reduce the misdiagnosis rate, and provide a reliable basis for treatment.
[0141] Second, to realize automatic data processing and intelligent diagnosis output, generate structured diagnosis report in a short time, shorten the time consumption of single case diagnosis, improve the diagnosis efficiency, and increase the number of patients treated by medical institutions.
[0142] Third, each model is rigorously trained to cope with complex clinical scenarios such as image quality fluctuations and unclear patient expressions, and to ensure stable and accurate diagnosis through multi-dimensional verification to enhance clinical practicability.
[0143] Fourth, relying on a friendly interface, grassroots personnel can obtain diagnosis suggestions by following the guidelines, reducing the technical threshold, promoting the sinking of high-quality diagnosis technology, and improving the prevention and control capacity of grassroots.
[0144] Fifth, to reduce the burden of doctors, alleviate the pressure of talent shortage, reduce resource consumption, reasonably allocate computing power with the help of open source framework, optimize medical resource allocation, and improve the benefit output ratio.
[0145] The electronic device disclosed in the embodiment includes a memory and a processor. The memory is used to store non-transitory computer readable instructions. Specifically, the memory can include one or more computer program products, which can include various forms of computer readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory, etc. The non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, etc.
[0146] The processor can be a central processing unit (CPU) or other forms of processing units with data processing capability and / or instruction execution capability, and can control other components in the electronic device to perform desired functions. In one embodiment of the present disclosure, the processor is used to run the computer readable instructions stored in the memory, so that the electronic device performs all or part of the steps of the multi-modal diagnosis method for dental diseases of the embodiments of the present disclosure.
[0147] Those skilled in the art should understand that, in order to solve the technical problem of how to obtain a good user experience effect, the present embodiment can also include well-known structures such as communication bus, interface, etc., which should also be included in the protection scope of the present disclosure.
[0148] As Figure 6 A structural schematic diagram of an electronic device provided by an embodiment of the present disclosure is shown. It shows a structural schematic diagram suitable for realizing the electronic device in the embodiment of the present disclosure. Figure 6 The electronic device shown is only an example and should not impose any limitation on the function and use range of the embodiments of the present disclosure.
[0149] As Figure 6As shown, the electronic device can include a processing device (e.g., a central processor, a graphics processor, etc.) that can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) or a program loaded from a storage device into a random access memory (RAM). In the RAM, various programs and data required for the operation of the electronic device are also stored. The processing device, the ROM, and the RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.
[0150] Generally, the following devices can be connected to the I / O interface: input devices including, for example, sensors or visual information collection devices; output devices including, for example, display screens; storage devices including, for example, magnetic tapes, hard disks, etc.; and communication devices. The communication devices can allow the electronic device to communicate wirelessly or wired with other devices (such as edge computing devices) to exchange data. Although Figure 6 The electronic device is shown with various devices, but it is understood that not all of the shown devices are required to be implemented or possessed. More or fewer devices can alternatively be implemented or possessed.
[0151] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication device, or installed from the storage device, or installed from the ROM. When the computer program is executed by the processing device, all or part of the steps of the multi-modal diagnosis method of dental diseases of embodiments of the present disclosure are performed.
[0152] Detailed descriptions related to the present embodiments can refer to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.
[0153] The computer-readable storage medium disclosed in the present embodiments has non-transitory computer-readable instructions stored thereon. When the non-transitory computer-readable instructions are run by a processor, all or part of the steps of the multi-modal diagnosis method of dental diseases of the embodiments of the present disclosure described above are performed.
[0154] The computer-readable storage medium described above includes, but is not limited to, optical storage media (e.g., CD-ROM and DVD), magneto-optical storage media (e.g., MO), magnetic storage media (e.g., magnetic tape or mobile hard disk), media with built-in rewritable non-volatile memory (e.g., memory card), and media with built-in ROM (e.g., ROM cartridge).
[0155] The detailed description of the embodiments hereinabove with reference to the accompanying drawings is only used to illustrate the basic or specific principles of the present disclosure and is not intended to limit the present disclosure. Therefore, the above description should not be used to restrict the present disclosure.
[0156] The above describes the basic principles of the present disclosure in combination with specific embodiments, but it should be noted that the advantages, benefits, effects and the like mentioned in the present disclosure are only examples and are not limiting, and these advantages, benefits, effects and the like cannot be considered as necessary for each embodiment of the present disclosure. In addition, the above specific details of the disclosure are only for the purpose of illustration and understanding, and are not limiting, and the above details do not limit the present disclosure to the above specific details.
[0157] In the present disclosure, the relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. The block diagrams of devices, apparatuses, equipment, systems involved in the present disclosure are only illustrative examples and are not intended to require or imply the connection, arrangement, configuration shown in the block diagram. As those skilled in the art will recognize, these devices, apparatuses, equipment, systems can be connected, arranged, configured in any manner. Words such as "include", "contain", "have" and the like are open-ended words, which mean "including but not limited to", and can be used interchangeably. The words "or" and "and" used herein mean the word "and / or", and can be used interchangeably unless the context clearly indicates otherwise. The word "such as" used herein means the phrase "such as but not limited to", and can be used interchangeably.
[0158] In addition, as used herein, "or" used in the list of items preceded by "at least one of" means a disjunctive list, such that, for example, a list of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). In addition, the word "exemplary" does not mean that the described example is preferred or better than other examples.
[0159] It should also be noted that in the systems and methods of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalents of the present disclosure.
[0160] Various changes, modifications, and alterations to the techniques described herein can be made without departing from the teachings of the attached claims. Moreover, the scope of the claims of this disclosure is not limited to the particular aspects described above. In addition, where a process, machine, manufacture, composition of matter, means, method, or result containing procedural, business, and other steps is described, it is understood that the description is meant to encompass the specific implementation of the steps described, as well as the substitution of equivalent steps, or equivalent steps in the performance order. Accordingly, the attached claims are to be interpreted as embracing the specific aspects and embodiments described herein, as well as future modifications, changes, and alterations of the aspects and embodiments.
[0161] The above description of the disclosed aspects is given for illustrative purposes only and is not intended to limit the scope of the disclosure. The aspects are described in terms of "preferred" embodiments and various modifications, alterations, and permutations of these preferred embodiments. These descriptions are not exhaustive and are intended to provide further examples. Accordingly, other alternatives, modifications, and variations should be considered within the scope of the disclosure. Other objects, advantages, and embodiments can be realized and can be resolved from the teaching of the disclosure, which is described in this disclosure and the annexed drawings.
[0162] The above description has been presented for the purpose of illustration and description. Furthermore, this description is not intended to limit the embodiments of the disclosure to forms disclosed herein. Although various example aspects and embodiments have been discussed above, those of ordinary skill in the art will appreciate a variety of modifications, alternative constructions, permutations, and equivalents.
Claims
1. A multimodal diagnostic method for dental diseases, characterized in that, include: Based on the acquired patient CT image data and corresponding medical condition text information data, a patient CT image dataset and a medical condition text information dataset are obtained. The patient's CT image data is input into the trained Mask R-CNN model to obtain the first CT pathological information; The patient's CT image data is input into the trained SAM model to obtain the second CT pathology information, which corresponds one-to-one with the first CT pathology information. The first CT pathology information and the corresponding second CT pathology information are fused to obtain CT pathology information, and a CT pathology information dataset is obtained based on the CT pathology information. The multimodal model was pre-trained based on patient CT image datasets, medical condition text information datasets, and CT pathology information datasets to obtain the trained multimodal model. Diagnosing dental diseases based on trained multimodal models; The Mask R-CNN model locates and segments lesion areas in CT images, outputting the location, size, and shape of the lesions; The SAM model supplements and refines lesion information; Integrating the outputs of the Mask R-CNN model and the SAM model includes: Data alignment is performed on the outputs of the two trained models, including spatial alignment and slice correspondence. The formula for spatial alignment is: , in, These are the original spatial coordinates. These are the transformed position coordinates. For radial basis functions, As control points, The parameters to be estimated, and a and b are the datasets to be aligned.
2. The multimodal diagnostic method for dental diseases according to claim 1, characterized in that, Training the Mask R-CNN model includes: A neural network structure for image segmentation was built based on the Mask R-CNN framework; The acquired training data is input into the neural network structure, the learning rate is set, the optimizer is selected, and the neural network structure is trained based on the loss function obtained by combining the classification cross-entropy loss function, the regression L1 / L2 loss function, and the mask segmentation loss function until the neural network structure converges.
3. The multimodal diagnostic method for dental diseases according to claim 1, characterized in that, SAM model training includes: Obtain the pre-trained weights of the SAM model; The acquired training data is input into the SAM model after loading the pre-trained weights. The weights of the selected general feature extraction layers at the bottom are frozen. The parameters of the high-level semantic segmentation related layers are adjusted. The learning rate is adjusted by combining the prior segmentation knowledge of the tooth structure. A lightweight optimizer is used to train the SAM model.
4. The multimodal diagnostic method for dental diseases according to claim 1, characterized in that, The CT pathology data set includes the location coordinates, size, morphological features, and disease classification labels of dental lesions.
5. The multimodal diagnostic method for dental diseases according to claim 1, characterized in that, The multimodal model constructed based on patient CT image datasets, medical condition text information datasets, and CT pathology information datasets is pre-trained to obtain the trained multimodal model, which includes: An image block dataset is obtained based on the patient CT image dataset, and a text block dataset is obtained based on the CT pathology information dataset and the disease text information dataset. The image patch dataset and the text patch dataset are concatenated into a multimodal input sequence; The constructed multimodal model is fine-tuned and trained based on the multimodal input sequence. The hyperparameters of the trained multimodal model are fine-tuned based on the objective function to obtain the trained multimodal model.
6. The multimodal diagnostic method for dental diseases according to claim 5, characterized in that, The step of concatenating the image patch dataset and the text patch dataset includes: The image patch features of the image patch dataset and the text patch features of the text patch dataset are mapped to a unified dimension to obtain the image patch feature projection and the text patch feature projection, respectively. Adjust the time steps of image block feature projection and text block feature projection to a unified time step; Based on the alignment objective function, the image block feature projection and text block feature projection at the same time step are stitched together.
7. The multimodal diagnostic method for dental diseases according to claim 6, characterized in that, The alignment objective function includes constraining the similarity of aligned features through a loss function.
8. The multimodal diagnostic method for dental diseases according to claim 1, characterized in that, The step of diagnosing dental diseases based on the trained multimodal model also includes: Evaluate the segmentation task of the trained multimodal model and evaluate the multimodal diagnosis task of the trained multimodal model; Based on the evaluation results of the segmentation task and the multimodal diagnosis task, the content that needs to be optimized in the trained multimodal model under different scenarios is obtained. The trained multimodal model is optimized according to the requirements for optimization.
9. The multimodal diagnostic method for dental diseases according to claim 8, characterized in that, The segmentation task evaluation parameters include the intersection-over-union ratio and / or the Dice coefficient; Modal diagnostic task evaluation parameters include accuracy, recall, and / or F1 score.
10. An electronic device comprising: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed, enables the at least one processor to perform the multimodal diagnostic method for dental diseases according to any one of claims 1-9.
Citation Information
Patent Citations
Few-sample medical image segmentation method combining CNN and SAM
CN119741315A
Hip joint impact syndrome intelligent diagnosis method based on multi-modal data fusion
CN120089335A