Multi-modal information processing method and device, storage medium and program product
By deeply fusing multimodal dental imaging data through multiple machine learning models, the problem of isolated processing of multimodal images was solved, enabling high-precision, real-time dental image analysis, improving the detection rate and diagnostic consistency of early lesions, and meeting the needs of real-time clinical diagnosis.
Patent Information
- Application Number
- CN202511102447.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies process multimodal dental imaging data in isolation, lacking end-to-end collaborative fusion, resulting in insufficient accuracy in identifying cracked teeth and early caries lesions. Fragmented clinical information makes personalized diagnosis difficult, and the complex models cannot meet the needs of real-time diagnosis.
This study employs multiple machine learning models to deeply fuse features from oral images, unstructured text, and structured text. Multimodal feature fusion is achieved through fuzzy membership and Choquet integrals. Combined with texture representation and semantic modeling, this improves diagnostic accuracy and real-time performance.
It significantly improves the intelligence level of dental image analysis, enhances the early lesion detection rate and diagnostic consistency, achieves millisecond-level real-time inference, enhances diagnostic transparency and trust, and reduces deployment and maintenance costs.
Smart Images

Figure CN120977546A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence, and in particular to a multimodal information processing method and apparatus, storage medium and program product. Background Technology
[0002] With an aging population and changing lifestyles, the incidence of dental diseases (such as interproximal caries, cracked teeth, deep caries, periodontitis, and periapical periodontitis) continues to rise. If these conditions are not diagnosed and intervened in a timely and accurate manner, they will lead to irreversible damage to tooth structure (enamel and dentin defects) and periodontal tissues (alveolar bone resorption), affecting chewing function and even triggering systemic diseases. The following imaging techniques are commonly used in clinical practice to aid in diagnosis.
[0003] 1) Panoramic dental X-ray / periapical radiograph: Its advantage lies in observing proximal caries and bone resorption.
[0004] 2) Three-dimensional CBCT (Cone-beam Computer Tomography): It is crucial for determining the orientation of cracked teeth, the extent of periapical lesions, and the three-dimensional structure of bone.
[0005] 3) Intraoral optical scanning: Visually displays the crown morphology, minor caries, and early enamel demineralization. Summary of the Invention
[0006] The inventors have noticed the following problems in the related technology.
[0007] 1) Modal isolation: Studies on multi-focus single modality (using panoramic or CBCT images independently) or local module optimization lack in-depth exploration of end-to-end collaborative fusion of multi-source heterogeneous images such as 2D X-ray (panoramic / periapical), 3D CBCT images and high-resolution intraoral optical scans, which limits the further improvement of its accuracy in identifying lesions such as cracked teeth, early proximal caries and complex caries.
[0008] 2) Fragmented Clinical Information: Actual clinical decision-making requires a comprehensive assessment of imaging findings, patient complaints (such as pain in specific tooth locations or occlusal discomfort), and detailed medical history (previous restorative history, systemic diseases, and medication history) to evaluate disease progression and risk. Existing systems fragment the processing of imaging and clinical information, making it difficult to establish multimodal correlations and provide personalized diagnostic recommendations.
[0009] 3) Difficulty in practical application: Current multimodal research is mostly limited to laboratory prototypes. The models are computationally complex and have high inference latency, which cannot meet the needs of real-time interaction in the clinic or on mobile terminals.
[0010] Accordingly, this disclosure provides a multimodal information processing method that can obtain more accurate detection results by deeply fusing oral image features, unstructured text features corresponding to the patient's chief complaint, and structured text features corresponding to the patient's medical record.
[0011] In a first aspect of this disclosure, a multimodal information processing method is provided, executed by a multimodal information processing device, comprising: processing oral cavity image features and unstructured text features using a first machine learning model to obtain a first feature; processing the oral cavity image features, the unstructured text features, and the structured text features using a second machine learning model to obtain a second feature; processing the unstructured text features and the structured text features using a third machine learning model to obtain a third feature; determining an output feature based on the first feature, the second feature, and the third feature, wherein the output feature includes a first feature map corresponding to structural information, a second feature map corresponding to detail information, and a third feature map corresponding to semantic information; and classifying the first feature map, the second feature map, and the third feature map to obtain a first classification result.
[0012] In some embodiments, the classification process for the first feature map, the second feature map, and the third feature map includes: determining the fuzzy membership degree of the p-th pixel in the i-th feature map, where 1≤i≤3, 1≤p≤P, and P is the total number of pixels; sorting the three fuzzy membership degrees of the p-th pixel in the first feature map, the second feature map, and the third feature map to obtain a feature subset corresponding to the p-th pixel, thereby obtaining P feature subsets; assigning weights to each feature subset in the P feature subsets; integrating each feature subset to obtain a fused feature map; and classifying the fused feature map to obtain the first classification result.
[0013] In some embodiments, the classification process of the fused feature map includes: performing a global average pooling operation on the fused feature map using a pooling layer to obtain a first intermediate feature map; mapping the first intermediate feature map to the logits space using a fully connected layer to obtain a second intermediate feature map; and calculating the probability density of the second intermediate feature map in each category to obtain the first classification result.
[0014] In some embodiments, sorting the three fuzzy membership degrees of the p-th pixel in the first feature map, the second feature map, and the third feature map includes sorting the three fuzzy membership degrees in descending order.
[0015] In some embodiments, assigning weights to each of the P feature subsets includes: assigning weights to each feature subset using the Sugeno λ-measure.
[0016] In some embodiments, the sum of the fuzziness measures of the first feature map, the second feature map, and the third feature map is 1.
[0017] In some embodiments, integrating each feature subset includes performing Choquet fuzzy integration on each feature subset to obtain the fused feature map.
[0018] In some embodiments, the output features further include modality consistency features, and the multimodal information processing method further includes: extracting texture representation features from the modality consistency features; processing the texture representation features using a coarse classifier to obtain a coarse classification result; obtaining semantic representation features based on the texture representation features and the coarse classification result; processing the texture representation features using a filter with multiple frequency responses to obtain a first enhanced feature; performing similarity modeling using the first enhanced feature and the semantic representation features to obtain a second enhanced feature; and performing classification processing on the second enhanced feature to obtain a second classification result.
[0019] In some embodiments, processing the oral cavity image features and unstructured text features using the first machine learning model includes: generating a query vector of the first machine learning model based on the oral cavity image features; generating a key vector and a value vector of the first machine learning model based on the unstructured text features; processing the query vector, key vector, and value vector of the first machine learning model using the channel attention model of the first machine learning model to obtain a first channel attention feature; concatenating the oral cavity image features and the first channel attention feature to obtain a first concatenated feature; and processing the first concatenated feature using the spatial attention model of the first machine learning model to obtain the first feature.
[0020] In some embodiments, processing the oral cavity image features, the unstructured text features, and the structured text features using the second machine learning model includes: processing the oral cavity image features and the unstructured text features using a first sub-model in the second machine learning model to obtain a second channel attention feature; processing the unstructured text features and the structured text features using a second sub-model in the second machine learning model to obtain a third channel attention feature; concatenating the second channel attention feature and the third channel attention feature to obtain a second concatenated feature; and processing the second concatenated feature using the spatial attention model of the second machine learning model to obtain the second feature.
[0021] In some embodiments, processing the oral cavity image features and the unstructured text features using the first sub-model in the second machine learning model includes: generating a query vector of the first sub-model based on the unstructured text features; generating a key vector and a value vector of the first sub-model based on the oral cavity image features; and processing the query vector, key vector, and value vector of the first sub-model using the channel attention model of the first sub-model to obtain the second channel attention features.
[0022] In some embodiments, processing the unstructured text features and the structured text features using the second sub-model in the second machine learning model includes: generating a query vector of the second sub-model based on the unstructured text features; generating a key vector and a value vector of the second sub-model based on the structured text features; and processing the query vector, key vector, and value vector of the second sub-model using the channel attention model of the second sub-model to obtain the third channel attention feature.
[0023] In some embodiments, processing the unstructured text features and the structured text features using a third machine learning model includes: generating a query vector of the third machine learning model based on the structured text features; generating a key vector and a value vector of the third machine learning model based on the unstructured text features; processing the query vector, key vector, and value vector of the third machine learning model using a channel attention model of the third machine learning model to obtain a fourth channel attention feature; concatenating the structured text features and the fourth channel attention feature to obtain a third concatenated feature; and processing the third concatenated feature using a spatial attention model of the third machine learning model to obtain the third feature.
[0024] In some embodiments, determining the output feature based on the first feature, the second feature, and the third feature includes: concatenating the first feature, the second feature, and the third feature to obtain a multimodal feature; and projecting the multimodal feature onto a predetermined feature space to obtain the output feature.
[0025] In some embodiments, features of an oral cavity image are extracted to obtain oral cavity image features; features of unstructured text are extracted to obtain unstructured text features; and features of structured text are extracted to obtain structured text features.
[0026] In a second aspect of this disclosure, a multimodal information processing apparatus is provided, comprising: a memory; and a processor coupled to the memory, the processor being configured to execute instructions stored in the memory to implement the multimodal information processing method as described in any of the above embodiments.
[0027] In a third aspect of this disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions that, when executed by a processor, implement the method as described in any of the above embodiments.
[0028] In a fourth aspect of this disclosure, a computer program product is provided, including computer instructions, wherein the computer instructions, when executed by a processor, implement the method as described in any of the above embodiments.
[0029] Other features and advantages of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0030] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0031] Figure 1 This is a flowchart illustrating a multimodal information processing method according to an embodiment of the present disclosure;
[0032] Figure 2 This is a schematic diagram of a model architecture for processing multimodal information according to an embodiment of the present disclosure;
[0033] Figure 3 This is a schematic diagram of the structure of a first machine learning model according to an embodiment of the present disclosure;
[0034] Figure 4This is a schematic diagram of the structure of a second machine learning model according to an embodiment of the present disclosure;
[0035] Figure 5 This is a schematic diagram of the structure of a third machine learning model according to an embodiment of the present disclosure;
[0036] Figure 6 This is a flowchart illustrating a classification processing method according to an embodiment of the present disclosure;
[0037] Figure 7 This is a flowchart illustrating a classification processing method according to another embodiment of the present disclosure;
[0038] Figure 8 This is a schematic diagram of a model architecture for processing multimodal information according to another embodiment of the present disclosure;
[0039] Figure 9 This is a schematic diagram of the structure of a multimodal information processing device according to an embodiment of the present disclosure. Detailed Implementation
[0040] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.
[0041] Unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of this disclosure.
[0042] At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn according to actual scale.
[0043] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0044] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.
[0045] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0046] Figure 1 This is a schematic flowchart of a multimodal information processing method according to an embodiment of the present disclosure. In some embodiments, the following multimodal information processing method is executed by a multimodal information processing device, including steps 11-15.
[0047] In step 11, the oral cavity image features and unstructured text features are processed using the first machine learning model to obtain the first feature.
[0048] In some embodiments, features of the oral cavity image are extracted to obtain oral cavity image features. Features of unstructured text are extracted to obtain unstructured text features.
[0049] In some embodiments, the step of processing oral cavity image features and unstructured text features using a first machine learning model includes steps S11-S15.
[0050] S11. Generate the query vector for the first machine learning model based on the oral cavity image features.
[0051] For example, if the features of an oral cavity image are The query vector Q of the first machine learning model img As shown in formula (1).
[0052]
[0053] In formula (1), It is a mapping matrix.
[0054] S12. Generate the key vector and value vector of the first machine learning model based on the features of the unstructured text.
[0055] For example, if the unstructured text features are Then the key vector K of the first machine learning model txt As shown in formula (2), the value vector V of the first machine learning model txt As shown in formula (3).
[0056]
[0057]
[0058] In formula (2), Let be the mapping matrix. In formula (3), It is a mapping matrix.
[0059] S13. The query vector, key vector, and value vector of the first machine learning model are processed using the channel attention model of the first machine learning model to obtain the first channel attention features.
[0060] S14. The oral cavity image features and the first channel attention features are concatenated to obtain the first concatenated features.
[0061] S15. The first spliced feature is processed using the spatial attention model of the first machine learning model to obtain the first feature.
[0062] In step 12, the oral cavity image features, unstructured text features, and structured text features are processed using a second machine learning model to obtain the second feature.
[0063] In some embodiments, features of structured text are extracted to obtain structured text features.
[0064] In some embodiments, the step of processing oral image features, unstructured text features, and structured text features using a second machine learning model includes steps S21-S25.
[0065] S21. The oral cavity image features and unstructured text features are processed using the first sub-model in the second machine learning model to obtain the second channel attention features.
[0066] In some embodiments, a query vector for the first sub-model is generated based on unstructured text features, and a key vector and value vector for the first sub-model are generated based on oral cavity image features. Next, the query vector, key vector, and value vector of the first sub-model are processed using the channel attention model of the first sub-model to obtain the second channel attention features.
[0067] For example, if the features of an oral cavity image are Unstructured text features are Then the query vector Q of the first sub-model txt As shown in formula (4), the key vector K of the first sub-model img As shown in formula (5), the value vector V of the first sub-model img As shown in formula (6).
[0068]
[0069]
[0070]
[0071] In formula (4), Let be the mapping matrix. In formula (5), Let be the mapping matrix. In formula (6), It is a mapping matrix.
[0072] S22. The second sub-model in the second machine learning model is used to process the unstructured text features and structured text features to obtain the third channel attention features.
[0073] In some embodiments, a query vector for the second sub-model is generated based on unstructured text features, and a key vector and value vector for the second sub-model are generated based on structured text features. Next, the query vector, key vector, and value vector of the second sub-model are processed using the channel attention model of the second sub-model to obtain third channel attention features.
[0074] For example, if the unstructured text features are Structured text features are Then the query vector Q of the second sub-model txt As shown in formula (7), the key vector K of the second sub-model tab As shown in formula (8), the value vector V of the second sub-model tab As shown in formula (9).
[0075]
[0076]
[0077]
[0078] In formula (7), Let be the mapping matrix. In formula (8), Let be the mapping matrix. In formula (9), It is a mapping matrix.
[0079] S23. The second channel attention feature and the third channel attention feature are concatenated to obtain the second concatenated feature.
[0080] S24. The second splicing feature is processed using the spatial attention model of the second machine learning model to obtain the second feature.
[0081] In step 13, the unstructured text features and structured text features are processed using a third machine learning model to obtain the third feature.
[0082] In some embodiments, the step of processing unstructured text features and structured text features using a third machine learning model includes S31-S35.
[0083] S31. Generate the query vector for the third machine learning model based on the structured text features.
[0084] For example, if the structured text features are The query vector Q of the third machine learning model tabAs shown in formula (10).
[0085]
[0086] In formula (10), It is a mapping matrix.
[0087] S32. Generate the key vector and value vector of the third machine learning model based on the features of the unstructured text.
[0088] For example, if the unstructured text features are Then the key vector K of the third machine learning model txt As shown in formula (11), the value vector V of the first machine learning model txt As shown in formula (12).
[0089]
[0090]
[0091] In formula (11), Let be the mapping matrix. In formula (12), It is a mapping matrix.
[0092] S33. The query vector, key vector, and value vector of the third machine learning model are processed using the channel attention model of the third machine learning model to obtain the fourth channel attention feature.
[0093] S34. Concatenate the structured text features and the fourth channel attention features to obtain the third concatenated feature.
[0094] S35. The spatial attention model of the third machine learning model is used to process the third splicing feature to obtain the third feature.
[0095] In step 14, output features are determined based on the first feature, the second feature, and the third feature, wherein the output features include a first feature map corresponding to structural information, a second feature map corresponding to detail information, and a third feature map corresponding to semantic information.
[0096] In some embodiments, the first feature, the second feature, and the third feature are concatenated to obtain a multimodal feature. Next, the multimodal feature is projected onto a predetermined feature space to obtain the output feature.
[0097] It should be noted that the output features include not only the first feature map corresponding to structural information, the second feature map corresponding to detail information, and the third feature map corresponding to semantic information, but also modality consistency features.
[0098] Figure 2This is a schematic diagram of a model architecture for processing multimodal information according to an embodiment of the present disclosure.
[0099] like Figure 2 As shown, the first machine learning model 21 processes the oral cavity image features 201 and the unstructured text features 202 to obtain the first feature. The second machine learning model 22 processes the oral cavity image features 201, the unstructured text features 202, and the structured text features 203 to obtain the second feature. The third machine learning model 23 processes the unstructured text features 202 and the structured text features 203 to obtain the third feature.
[0100] Next, the first, second, and third features are concatenated in functional layer 24 to obtain multimodal features. Then, the multimodal features are projected onto a predetermined feature space to obtain the output features.
[0101] In some embodiments, the first machine learning model 21 is as follows: Figure 3 As shown.
[0102] exist Figure 3 In this process, a query vector for the first machine learning model is generated based on oral cavity image features 201. Key and value vectors for the first machine learning model are generated based on unstructured text features 202. The query vector, key vector, and value vector of the first machine learning model are processed using the channel attention model to obtain the first channel attention feature. Next, the oral cavity image features 201 and the first channel attention feature are concatenated to obtain the first concatenated feature. The first concatenated feature is processed using the spatial attention model of the first machine learning model, and the result is passed through a linear layer to obtain the first feature.
[0103] In some embodiments, the second machine learning model 22 is as follows Figure 4 As shown.
[0104] like Figure 4 As shown, the first machine learning model 22 includes a first sub-model 221 and a second sub-model 222.
[0105] The query vector of the first sub-model 221 is generated based on the unstructured text features 202, and the key vector and value vector of the first sub-model 221 are generated based on the oral cavity image features 201. Next, the query vector, key vector, and value vector of the first sub-model 221 are processed using the channel attention model of the first sub-model 221 to obtain the second channel attention features.
[0106] The query vector for the second sub-model 222 is generated based on the unstructured text features 202, and the key and value vectors for the second sub-model 222 are generated based on the structured text features 203. The query vector, key vector, and value vector of the second sub-model 222 are then processed using the channel attention model of the second sub-model 222 to obtain the third channel attention features.
[0107] Next, the second-channel attention features and the third-channel attention features are concatenated to obtain the second concatenated feature. The second concatenated feature is then processed using the spatial attention model of the second machine learning model, and the result is passed through a linear layer to obtain the second feature.
[0108] In some embodiments, the third machine learning model 23, such as Figure 5 As shown.
[0109] like Figure 5 As shown, the query vector of the third machine learning model is generated based on the structured text features 203. The key vector and value vector of the third machine learning model are generated based on the unstructured text features 202. The query vector, key vector, and value vector of the third machine learning model are then processed using the channel attention model of the third machine learning model to obtain the fourth channel attention features.
[0110] Next, the structured text features and the fourth channel attention features are concatenated to obtain the third concatenated feature. The spatial attention model of the third machine learning model is used to process the third concatenated feature, and the result is passed through a linear layer to obtain the third feature.
[0111] return Figure 1 In step 15, the first feature map, the second feature map, and the third feature map are classified to obtain the first classification result.
[0112] Figure 6 This is a schematic flowchart of a classification processing method according to an embodiment of the present disclosure. In some embodiments, the following classification processing method is performed by a multimodal information processing device, including steps 61-65.
[0113] In step 61, the fuzzy membership degree of the p-th pixel in the i-th feature map is determined, where 1≤i≤3, 1≤p≤P, and P is the total number of pixels.
[0114] It should be noted that fuzzy membership is used to measure the degree to which a pixel belongs to the current feature map.
[0115] In step 62, the three fuzzy membership degrees of the p-th pixel in the first feature map, the second feature map, and the third feature map are sorted to obtain the feature subset corresponding to the p-th pixel, thus obtaining P feature subsets.
[0116] In some embodiments, the three fuzzy membership degrees are sorted in descending order.
[0117] For example, the fuzzy membership degree of the p-th pixel in the first feature map is The fuzzy membership degree of the p-th pixel in the second feature map is The fuzzy membership degree of the p-th pixel in the third feature map is The feature subset A(p) corresponding to the p-th pixel is shown in formula (13), and and It satisfies formula (14).
[0118]
[0119]
[0120] In step 63, weights are assigned to each of the P feature subsets.
[0121] In some embodiments, weights are assigned to each feature subset using the Sugenoλ-measure.
[0122] It's worth noting that the Sugeno λ-measure, also known as the λ-fuzzy measure, was proposed by Japanese scholar Michio Sugeno. This is a non-additive measure with the parameter λ, widely used in fuzzy mathematics and non-additive measure theory to handle uncertainties where the probability space cannot be covered. It can capture mutually reinforcing or inhibiting relationships between multiple features.
[0123] In some embodiments, the sum of the fuzziness measures of the first feature map, the second feature map, and the third feature map is 1.
[0124] It should be noted here that, to ensure the consistency of fuzzy measures across the entire feature set, the sum of the fuzzy measures is set to 1. Let the first feature map be... Second feature map The third feature map is Then formula (15) holds true.
[0125]
[0126] In formula (15), ζ is the fuzzy measure function.
[0127] In step 64, each feature subset is integrated to obtain a fused feature map.
[0128] In some embodiments, Choquet fuzzy integration is performed on each feature subset to obtain a fused feature map.
[0129] It's important to note that Choquet fuzzy integral is a nonlinear integration method based on fuzzy measures. By integrating a sorted subset of features using Choquet fuzzy integral, nonlinear fusion between features can be achieved. This integration process not only integrates strong response features but also adjusts the weights of redundant or conflicting features through the nonlinear relationship introduced by the measure, effectively enhancing the expressiveness of the features.
[0130] In step 65, the fused feature map is classified to obtain the first classification result.
[0131] In some embodiments, the step of classifying the fused feature map includes S41-S43.
[0132] S41. Use a pooling layer to perform global average pooling on the fused feature map to obtain the first intermediate feature map.
[0133] S42. Use a fully connected layer to map the first intermediate feature map to the logits space to obtain the second intermediate feature map.
[0134] It should be noted that the Logits space refers to the linear space formed by the raw values of the output of the last layer (usually a fully connected layer) of a deep learning model. These values have not undergone normalization processing (such as Softmax or Sigmoid function transformation) and directly reflect the model's confidence scores for different categories.
[0135] S43. Calculate the probability density of the second intermediate feature map for each category to obtain the first classification result.
[0136] For example, the first classification result Cls is shown in formula (16).
[0137]
[0138] In formula (16), FC is a fully connected layer, GAP is a global average pooling operation, and C μ This is the second intermediate feature map. Let N be the index score for the j-th category, and N be the total number of categories.
[0139] The above classification method expresses the dependencies between features using the Sugeno λ-measure and achieves non-additive and non-linear feature fusion using Choquet integral, which effectively enhances the expressiveness of features and the decision-making performance of classification, while improving the model's ability to model complex semantic structures and its interpretability.
[0140] In some embodiments, the output features include, in addition to the first feature map corresponding to structural information, the second feature map corresponding to detail information, and the third feature map corresponding to semantic information, modal consistency features are also included, so as to extract texture identification features using modal consistency features, thereby improving the detection rate of early proximal caries, hidden cracked teeth, and microcaries.
[0141] Figure 7 This is a flowchart illustrating a classification processing method according to another embodiment of the present disclosure. In some embodiments, the following classification processing method is performed by a multimodal information processing device, including steps 71-76.
[0142] In step 71, texture representation features are extracted from modal consistency features.
[0143] It should be noted that texture representation features preserve low-level visual information such as local structure and edge details.
[0144] In step 72, the texture representation features are processed using a coarse classifier to obtain the coarse classification result.
[0145] In step 73, semantic representation features are obtained based on texture representation features and coarse classification results.
[0146] It should be noted that semantic representation features can encode high-level abstract information such as category semantics and context through deep networks.
[0147] In step 74, the texture representation features are processed using a filter with multiple frequency responses to obtain the first enhanced feature.
[0148] It should be noted that by using filters with multiple frequency responses to process the texture representation features, the texture representation features become more discriminative and the interference of invalid or redundant information can be reduced.
[0149] In step 75, similarity modeling is performed using the first enhanced feature and semantic representation feature to obtain the second enhanced feature.
[0150] It should be noted that similarity modeling can generate more expressive feature representations.
[0151] In step 76, the second enhanced features are classified to obtain the second classification result.
[0152] Figure 8 This is a schematic diagram of a model architecture for processing multimodal information according to another embodiment of the present disclosure. Figure 8 and Figure 2 The difference is that, in Figure 8 The embodiment shown also includes a first classification model 81 and a second classification model 82.
[0153] The first classification model 81 is used for execution. Figure 6 The steps described in any of the embodiments express the dependencies between features through the Sugeno λ-measure and achieve non-additive and non-linear feature fusion using Choquet integrals, which effectively enhances the expressive power of features and the decision-making performance of classification, while improving the model's ability to model complex semantic structures and its interpretability.
[0154] The second classification model 82 is used for execution. Figure 7 The steps described in any of the embodiments extract texture identification features by utilizing modal consistency features in order to improve the detection rate of early proximal caries, hidden cracked teeth, and microcaries.
[0155] In the multimodal information processing method provided in the above embodiments of this disclosure, more accurate detection results can be obtained by deeply fusing oral image features, unstructured text features corresponding to the patient's chief complaint, and structured text features corresponding to the patient's medical record.
[0156] Figure 9 This is a schematic diagram of the structure of a multimodal information processing device according to an embodiment of the present disclosure.
[0157] like Figure 9 As shown, the multimodal information processing device 90 can be represented in the form of a general computing device. The multimodal information processing device 90 includes a memory 91, a processor 92, and a bus 93 connecting different system components.
[0158] Memory 91 may include, for example, system memory, non-volatile storage media, etc. System memory may store, for example, an operating system, application programs, a boot loader, and other programs. System memory may include volatile storage media, such as random access memory (RAM) and / or cache memory. Non-volatile storage media may store, for example, at least one program in execution.
[0159] Instructions for a corresponding embodiment of the method. Non-volatile storage media include, but are not limited to, disk storage, optical storage, flash memory, etc.
[0160] The processor 92 can be implemented using a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete hardware components such as discrete gates or transistors. Accordingly, each module, such as the acquisition module, the calculation module, and the adjustment module, can be implemented by the central processing unit (CPU) running instructions in the memory to execute the corresponding steps, or by dedicated circuitry to execute the corresponding steps.
[0161] For example, processor 92 is configured for memory-based instruction execution implementation such as Figure 1 , 6 The method involved in any of the embodiments in 7.
[0162] Bus 93 can use any of the various bus architectures. For example, bus architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MCA) bus, and the Peripheral Component Interconnect (PCI) bus.
[0163] The interfaces 94, 95, and 96 of the multimodal information processing device 90, as well as the memory 91 and processor 92, can be connected via bus 93. Input / output interface 94 provides a connection interface for input / output devices such as displays, mice, and keyboards. Network interface 95 provides a connection interface for various networked devices. Storage interface 96 provides a connection interface for external storage devices such as floppy disks, USB flash drives, and SD cards.
[0164] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus, and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations thereof, can be implemented by computer-readable program instructions.
[0165] These computer-readable program instructions are provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable device to produce a machine, such that execution of the instructions by the processor produces means for implementing the functions specified in one or more boxes of the flowchart and / or block diagram.
[0166] These computer-readable program instructions may also be stored in a computer-readable storage medium. These instructions cause a computer to work in a particular manner to produce an article of manufacture, including instructions that implement the functions specified in one or more boxes in a flowchart and / or block diagram.
[0167] This disclosure may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects.
[0168] This disclosure also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement... Figure 1 , 6 The method involved in any of the embodiments in 7.
[0169] This disclosure also provides a computer program product, including computer instructions, wherein the computer instructions, when executed by a processor, implement as follows: Figure 1 , 6 The method involved in any of the embodiments in 7.
[0170] In some embodiments, the aforementioned machine learning models may include CNN (Convolutional Neural Network) models, U-Net (Convolutional Networks for Biomedical Image Segmentation) models, Mask R-CNN (Mask Region-based Convolutional Neural Network), Transformer, and other models.
[0171] By using the solutions provided in the above embodiments of this disclosure, the level of intelligence in dental image analysis can be effectively improved.
[0172] For example, on a dataset of more than 500 two-dimensional panoramic and periapical radiographs, the detection rate of dental caries (especially proximal caries) increased from about 75% to over 90%.
[0173] For example, in 3D CBCT images, high-precision segmentation of periapical periodontitis and periodontal bone resorption has been achieved, reducing the false negative rate by about 20%.
[0174] By implementing this disclosure, the following beneficial effects can be obtained.
[0175] 1. Significantly improve diagnostic performance: By integrating multimodal imaging and clinical data, the detection rate and diagnostic consistency of early micro-lesions (such as proximal caries, occult cracks, and superficial caries) are greatly improved.
[0176] 2. Enable real-time clinical applications: The end-to-end lightweight architecture ensures millisecond-level real-time inference at the clinic chair and on mobile devices, meeting the needs of immediate clinical decision support.
[0177] 3. Enhance decision-making transparency and trust: Interpretable output (visual membership graphs and confidence reports) makes diagnostic evidence traceable, improving doctors' trust in the results provided by artificial intelligence and the efficiency of clinical decision-making.
[0178] 4. Reduce deployment and maintenance costs: The integrated process and lightweight design simplify deployment and maintenance, lowering the application threshold.
[0179] In some embodiments, the functional units described above may be implemented as general-purpose processors, programmable logic controllers (PLCs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or any suitable combination thereof for performing the functions described herein.
[0180] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0181] The description in this disclosure is provided for illustrative and descriptive purposes only and is not intended to be exhaustive or to limit the disclosure to its forms. Many modifications and variations will be apparent to those skilled in the art. The embodiments were chosen and described in order to better illustrate the principles and practical application of this disclosure and to enable those skilled in the art to understand this disclosure and to design various embodiments with various modifications suitable for a particular purpose.
Claims
1. A multimodal information processing method, executed by a multimodal information processing device, comprising: The first machine learning model is used to process the oral cavity image features and unstructured text features to obtain the first feature; The oral cavity image features, unstructured text features, and structured text features are processed using a second machine learning model to obtain the second feature; The unstructured text features and the structured text features are processed using a third machine learning model to obtain a third feature; Based on the first feature, the second feature, and the third feature, an output feature is determined, wherein the output feature includes a first feature map corresponding to structural information, a second feature map corresponding to detail information, and a third feature map corresponding to semantic information; The first feature map, the second feature map, and the third feature map are classified to obtain a first classification result.
2. The multimodal information processing method according to claim 1, wherein, The classification process for the first feature map, the second feature map, and the third feature map includes: Determine the fuzzy membership degree of the p-th pixel in the i-th feature map, where 1≤i≤3, 1≤p≤P, and P is the total number of pixels; The three fuzzy membership degrees of the p-th pixel in the first feature map, the second feature map, and the third feature map are sorted to obtain the feature subset corresponding to the p-th pixel, thereby obtaining P feature subsets; Assign weights to each of the P feature subsets; Integrate each feature subset to obtain a fused feature map; The fused feature map is then classified to obtain the first classification result.
3. The multimodal information processing method according to claim 2, wherein, The classification process of the fused feature map includes: A global average pooling operation is performed on the fused feature map using a pooling layer to obtain a first intermediate feature map; The first intermediate feature map is mapped to the logits space using a fully connected layer to obtain the second intermediate feature map; Calculate the probability density of the second intermediate feature map for each category to obtain the first classification result.
4. The multimodal information processing method according to claim 2, wherein, The step of sorting the three fuzzy membership degrees of the p-th pixel in the first feature map, the second feature map, and the third feature map includes: The three fuzzy membership degrees are sorted in descending order.
5. The multimodal information processing method according to claim 2, wherein, The process of assigning weights to each feature subset in the P feature subsets includes: Weights are assigned to each feature subset using the Sugeno λ-measure.
6. The multimodal information processing method according to claim 5, wherein, The sum of the fuzzy measures of the first feature map, the second feature map, and the third feature map is 1.
7. The multimodal information processing method according to claim 2, wherein, The integration of each feature subset includes: The Choquet fuzzy integral is performed on each feature subset to obtain the fused feature map.
8. The multimodal information processing method according to claim 1, wherein, The output features also include modality consistency features, and the multimodal information processing method further includes: Extract texture representation features from the modality consistency features; The texture representation features are processed using a coarse classifier to obtain a coarse classification result; Based on the texture representation features and the coarse classification results, semantic representation features are obtained; The texture representation features are processed using a filter with multiple frequency responses to obtain a first enhanced feature; Similarity modeling is performed using the first enhanced feature and the semantic representation feature to obtain the second enhanced feature; The second enhanced feature is then classified to obtain the second classification result.
9. The multimodal information processing method according to claim 1, wherein, The process of using the first machine learning model to process oral cavity image features and unstructured text features includes: Generate a query vector for the first machine learning model based on the oral cavity image features; Generate the key vector and value vector of the first machine learning model based on the unstructured text features; The query vector, key vector, and value vector of the first machine learning model are processed using the channel attention model of the first machine learning model to obtain the first channel attention feature; The oral cavity image features and the first channel attention features are concatenated to obtain the first concatenated feature; The first concatenated feature is obtained by processing the first spliced feature using the spatial attention model of the first machine learning model.
10. The multimodal information processing method according to claim 1, wherein, The process of using the second machine learning model to process the oral cavity image features, the unstructured text features, and the structured text features includes: The oral cavity image features and the unstructured text features are processed using the first sub-model in the second machine learning model to obtain the second channel attention features; The unstructured text features and the structured text features are processed using the second sub-model in the second machine learning model to obtain the third channel attention features; The second channel attention feature and the third channel attention feature are concatenated to obtain the second concatenated feature; The second concatenated feature is processed using the spatial attention model of the second machine learning model to obtain the second feature.
11. The multimodal information processing method according to claim 10, wherein, The process of using the first sub-model in the second machine learning model to process the oral cavity image features and the unstructured text features includes: Generate the query vector for the first sub-model based on the unstructured text features; Generate the key vector and value vector of the first sub-model based on the oral cavity image features; The query vector, key vector, and value vector of the first sub-model are processed using the channel attention model of the first sub-model to obtain the second channel attention feature.
12. The multimodal information processing method according to claim 10, wherein, The process of using the second sub-model in the second machine learning model to process the unstructured text features and the structured text features includes: Generate the query vector for the second sub-model based on the unstructured text features; Generate the key vector and value vector of the second sub-model based on the structured text features; The query vector, key vector, and value vector of the second sub-model are processed using the channel attention model of the second sub-model to obtain the third channel attention feature.
13. The multimodal information processing method according to claim 1, wherein, The process of using a third machine learning model to process the unstructured text features and the structured text features includes: The query vector of the third machine learning model is generated based on the structured text features. Generate the key vector and value vector of the third machine learning model based on the unstructured text features; The query vector, key vector, and value vector of the third machine learning model are processed using the channel attention model of the third machine learning model to obtain the fourth channel attention feature; The structured text features and the fourth channel attention features are concatenated to obtain the third concatenated feature; The third concatenated feature is obtained by processing the spatial attention model of the third machine learning model.
14. The multimodal information processing method according to any one of claims 1-13, wherein, The step of determining the output feature based on the first feature, the second feature, and the third feature includes: The first feature, the second feature, and the third feature are concatenated to obtain a multimodal feature; The multimodal features are projected onto a predetermined feature space to obtain the output features.
15. The multimodal information processing method according to any one of claims 1-13, further comprising: Extract features from the oral cavity image to obtain the oral cavity image features; Extract features from unstructured text to obtain the unstructured text features; Extract features from the structured text to obtain the structured text features.
16. A multimodal information processing device, comprising: Memory; A processor, coupled to a memory, is configured to implement the multimodal information processing method as described in any one of claims 1-15 based on the execution of instructions stored in the memory.
17. A computer-readable storage medium, wherein, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the multimodal information processing method as described in any one of claims 1-15.
18. A computer program product comprising computer instructions, wherein the computer instructions, when executed by a processor, implement the multimodal information processing method as described in any one of claims 1-15.