A skin phenomenon classification method based on multi-modal machine learning
By integrating skin lesion images and clinical data through multimodal machine learning methods, the limitations of single-modality detection and the challenge of FF phenomenon identification in AD and PV detection are solved. This achieves highly accurate and interpretable skin phenomenon classification, supporting the early identification and quantitative prediction of FF phenomenon.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- THE UNIVERSITY OF HONG KONG SHENZHEN HOSPITAL
- Filing Date
- 2026-02-19
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies for detecting AD and PV suffer from limitations such as single-modality limitation, gaps in identifying FF phenomenon, insufficient model generalization ability, and poor interpretability, making it difficult to detect the risk of phenotypic transformation in patients at an early stage.
A multimodal machine learning approach is employed, combining skin lesion images and clinical data. Features are transformed through an image feature extractor and a text encoder, and feature fusion is performed using a cross-attention mechanism and a multi-expert network. Modal contributions are dynamically adjusted to achieve classification of AD, PV, and FF phenomena.
It improves the accuracy of skin phenomenon classification, enhances the ability to identify rare cases of FF, enables quantitative prediction of the probability of FF occurrence, provides interpretable diagnostic evidence, and supports personalized treatment decisions.
Smart Images

Figure CN122135935A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of image processing and artificial intelligence, and in particular to a skin phenomenon classification method based on multimodal machine learning. Background Technology
[0002] In the field of dermatology, atopic dermatitis (AD) and psoriasis vulgaris (PV) are two common chronic inflammatory skin diseases. Traditionally, they were considered to be mutually exclusive in their immunological mechanisms: AD primarily manifests as a Th2 immune response, while PV is mainly characterized by a Th1 / Th17 response. However, recent clinical observations have revealed that some patients exhibit characteristics of both diseases, and others have shown a "flip-flop" (FF) phenomenon—a transition from one disease phenotype to the other—after treatment with biologics. This presents new challenges for medical testing.
[0003] Currently, the detection of Alzheimer's disease (AD) and polymorphic leukemia (PV) mainly relies on clinicians' history taking and skin lesion examination, lacking objective and quantitative auxiliary tools. Existing computer-aided detection methods mainly have the following problems: Single-modal limitations: Existing methods often use a single data source (images only or clinical data only), which cannot comprehensively capture disease characteristics; The gap in identifying the FF phenomenon: There is currently no automatic identification method for the FF phenomenon, making it difficult to detect the risk of phenotypic transformation in patients at an early stage; Insufficient model generalization ability: Existing deep learning models are mostly based on general pre-trained models and lack domain-adaptive optimization for skin lesion images; Poor interpretability: Black-box deep learning models are difficult to provide detection evidence, which limits their clinical application. Summary of the Invention
[0004] This invention provides a skin phenomenon classification method based on multimodal machine learning to solve at least one of the above-mentioned problems.
[0005] In a first aspect, embodiments of the present invention provide a skin phenomenon classification method based on multimodal machine learning, including: S110. Obtain skin lesion images and clinical data of the patients to be classified; S120. The skin lesion images and clinical data are converted into image features and text features respectively using an image feature extractor and a text encoder; S130. The image features and text features are fused using a cross-attention mechanism, so that the two features pay attention to each other's relevant parts. S140. Multiple expert networks are used to process the fused features to learn different symptom patterns. Then, the output features of each expert network are fused using gating weights to obtain the final fused features. S150. Using different machine learning networks, predict the probability that the patient belongs to atopic dermatitis, psoriasis, and Flip-Flop phenomenon based on the final fusion features, thereby achieving skin phenomenon classification.
[0006] In a second aspect, embodiments of the present invention provide an electronic device, the electronic device comprising: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the skin phenomenon classification method based on multimodal machine learning as described in any embodiment.
[0007] Thirdly, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the skin phenomenon classification method based on multimodal machine learning as described in any embodiment.
[0008] In summary, the embodiments of the present invention provide a skin phenomenon classification method based on multimodal machine learning, which can achieve the following beneficial effects: 1. Improve the accuracy of skin phenomenon classification. Existing methods fail to effectively integrate the visual features of skin lesions with the semantic information of clinical history, resulting in information waste. This embodiment proposes a cross-modal attention mechanism for skin lesion images and clinical history data. Through semantic space unification, cross-attention mechanisms, multi-expert learning of different visual modes, and dynamic fusion using gating networks, it achieves deep fusion of image visual features and clinical semantic information, overcoming the limitations of single data sources and improving the accuracy of skin phenomenon classification. Compared to using images alone (accuracy approximately 90%) or clinical data alone (accuracy approximately 85%), the method in this embodiment improves accuracy by 7-12 percentage points.
[0009] 2. Strong ability to identify rare cases. Because FF cases are relatively rare in clinical practice, existing models perform poorly in identifying a few categories, easily leading to missed diagnoses. However, this embodiment, through a class balancing strategy, focus loss, and dynamic class weights, still achieves a 96% F1 score even when FF cases account for only 6.9% of the total sample, providing a powerful tool for the early identification of rare clinical diseases.
[0010] 3. To achieve quantitative prediction of the probability of FF occurrence, providing support for early clinical intervention. Existing technologies lack quantitative prediction of the risk of FF occurrence, leading to incorrect clinical intervention selection. This embodiment proposes, for the first time, a quantitative assessment of the probability of FF occurrence, providing early warning for patient monitoring during biologics intervention and supporting personalized treatment decisions. Attached Figure Description
[0011] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0012] Figure 1 This is a flowchart of a skin phenomenon classification method based on multimodal machine learning provided in an embodiment of the present invention; Figure 2 This is a flowchart of another skin phenomenon classification method based on multimodal machine learning provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the image classification model provided in an embodiment of the present invention; Figure 4 This is a visual comparison chart of the prediction results of the three modalities provided in the embodiments of the present invention; Figure 5 This is the Grad-CAM diagram (attention heatmap) for model attention visualization provided in this embodiment of the invention. Figure 6 This is a cross-attention mechanism visualization heatmap provided in an embodiment of the present invention, wherein... Figure 6 A is the attention confusion matrix, representing the overall mean, AD category, FF category, and PV category, respectively; Figure 6 B represents the model weight allocation result for the test set data; Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0014] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0015] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0016] Figure 1 This is a flowchart illustrating a skin phenomenon classification method based on multimodal machine learning, provided in an embodiment of the present invention. This method is applicable to the detection and classification of three skin diseases or phenomena: Alzheimer's disease (AD), polymorphic dermatitis (PV), and follicular fibrillation (FF), and is executed by an electronic device. Figure 1 As shown, the method specifically includes: S110. Obtain skin lesion images and clinical data of the patients to be classified.
[0017] The skin lesion images here refer to dermoscopic images of the patient's lesion sites. Typically, a patient has multiple skin lesion images, each corresponding to different lesion sites, such as one image of the leg and another of the back.
[0018] Clinical data refers to key clinical features extracted from a patient's medical history. Optional clinical data may include: symptom characteristics (severity of itching, distribution of skin lesions, course of disease, etc.), past medical history (history of allergies, family history, treatment history, etc.), and laboratory indicators (IgE levels, eosinophil count, etc.).
[0019] This embodiment first acquires images of the patient's skin lesions and clinical data. Subsequently, based on these data, the probability of the patient belonging to AD, PV, and FF phenomena is predicted to achieve the purpose of skin phenomenon classification.
[0020] S120. Using an image feature extractor and a text encoder respectively, the skin lesion images and clinical data are converted into image features and text features respectively.
[0021] This step extracts features from both image and clinical modalities. Figure 2 In one specific implementation, the clinical history data can first be transformed into a semantic vector through structured feature engineering. For example, the text features of the clinical data are 21-dimensional. Experienced doctors can perform binary classification on the 21 clinical features and encode them as 0 / 1 values, realizing the conversion from textual information to 0 / 1 encoding (for example, the feature of psoriasis history is 1 if present and 0 if absent; the interval in which PSI is located is encoded as 1, and outside the interval is encoded as 0). Finally, all the clinical feature codes are linked together to form a 21-dimensional vector.
[0022] Then, perform the following operations on each skin lesion image: S1-1. The fine-tuned derm-foundation model is used as an image feature extractor to extract image features from the current skin lesion image. Specifically, the derm-foundation pre-trained model is specifically trained on a large-scale dermatology image dataset, enabling it to capture domain-specific features of skin lesion images. This embodiment fine-tunes the model to capture skin features, rather than using the general YOLO / R-CNN model, which can better extract skin phenomenon features. For example, after processing the current skin lesion image using the fine-tuned derm-foundation model, 6144-dimensional image features can be obtained.
[0023] S1-2, dimensionality reduction is performed on the image features using F-tests, mutual information, and neural networks respectively, and the top N most influential dimensions are selected by voting. As mentioned above, image features typically have a larger dimension than text features. To facilitate subsequent fusion of the two features, this step unifies the image features into the same semantic space as the text features. Optionally, combining... Figure 2 The 6144-dimensional image features of the current skin lesion image can be reduced in dimensionality using three methods: F-test, neural network, and mutual information. The top N most influential features are then selected through a voting process. For example, N can be 300 or 400. The neural network can include a weighted layer, two compression layers, and an output layer. The input to the neural network is the 6144-dimensional image features, and the output is the score of each feature dimension.
[0024] S1-3. Utilize a Multilevel Feature Extraction (MLFE) layer to extract depth features of the first N dimensions of the image at different scales, and combine these depth features at different scales to consider both global and detailed information. Optionally, combine... Figure 2The MLFE module can extract depth features at three different resolution scales, ensuring that no scale information (details + global) is missed. Optionally, the MLFE can include three parallel processing channels, corresponding to high, medium and low resolution channels respectively. Each channel is responsible for feature extraction at one scale, and each channel is connected to a convolutional and pooling layer. After feature alignment, multi-scale fusion is performed.
[0025] After performing the above operations on each skin lesion image, the stitched features of each skin lesion image can be obtained. Then, using a multi-head self-attention mechanism, the stitched features of each skin lesion image are fused through channels to obtain the final image features with the same dimension as the text features. At this point, the image features and clinical features are mapped to a common semantic space, and the feature data at this stage represents all the information of the current patient.
[0026] S130. The image features and text features are fused using the CrossModalAttention mechanism, so that the two features pay attention to the relevant parts of each other.
[0027] Combination Figure 2 This step fuses image and text features within the same semantic space through a cross-attention mechanism, enabling image features to focus on relevant clinical semantic information, and vice versa. Simultaneously, the attention mechanism allows the model to dynamically adjust the contributions of each modality, increasing clinical weights when clinical presentations are typical and image weights when skin lesion morphology is prominent.
[0028] S140. The fused features are processed by multiple expert networks to learn different symptom patterns. Then, the output features of each expert network are fused using gating weights to obtain the final fused features.
[0029] This step employs a Mixture of Experts (MoE) ensemble strategy to further fuse features and improve classification robustness. Each expert network learns specifically for different visual patterns (such as erythema, scaling, mossification, etc.).
[0030] Optionally, the fused features can be processed using an XGBoost expert network based on Focal Loss to address class imbalance; a logistic regression expert network can be used to process the fused features to provide a linear probability baseline; a multilayer perceptron expert network can be used to process the fused features to capture nonlinear feature interactions; a support vector machine expert network can be used to process the fused features and construct a classification hyperplane in a high-dimensional feature space; and a K-nearest neighbor expert network can be used to process the fused features and perform classification based on instance similarity.
[0031] Then, a gating network is used to dynamically assign weights to the output features of each expert network, and the output features of multiple expert networks are weighted and fused to improve classification robustness.
[0032] S150. Using different machine learning networks, predict the probability that the patient belongs to atopic dermatitis, psoriasis, and Flip-Flop phenomenon based on the final fusion features, thereby achieving skin phenomenon classification.
[0033] Combination Figure 2 The final fused features are then input into multiple machine learning networks, including XGBoost, SVM, LightGBM, and R-Forest. Each machine learning network outputs a probability distribution belonging to one of the three categories. These probability distributions are then fused to obtain the final output, assisting doctors in making more accurate judgments.
[0034] Optionally, the final output of the method includes: Disease classification results (AD / PV / FF) and confidence scores: The final disease probability distribution for this patient can be obtained by weighting and summing the four models using the weight allocator, similar to AD, FF, and PV (0.2, 0.7, 0.1). The confidence score calculation process is roughly as follows: the validation set accuracy of each model is used as the basic weight, multiplied by the probability value of that model for the predicted class, and then the weighted probabilities of all models are summed to obtain the final confidence score.
[0035] Quantitative scoring of the probability of FF occurrence (0-100%): It can output a softmax probability distribution, in which the probability value of FF category is used as a quantitative indicator of the risk of the flip phenomenon; a risk threshold is set, and a clinical warning is triggered when the FF probability exceeds the threshold. Key diagnostic criteria are visualized (such as attention heatmaps, ranking of important features, etc., the specific generation method will be explained in subsequent embodiments); Clinical recommendations (such as whether further examination is needed, treatment options, and adjustment suggestions).
[0036] As can be seen, this embodiment makes full use of image data and clinical data, and performs multimodal data fusion through various methods such as semantic space unification, cross-attention mechanism, multi-expert dynamic weighted fusion, and gated weighted fusion, which improves the classification accuracy of the three skin phenomena and provides detailed quantitative results on the probability, risk and confidence of FF phenomenon.
[0037] Furthermore, the identification of the FF phenomenon is both a key focus and a challenge in this embodiment. Specifically, FF (flipping phenomenon) is not a typical "disease type," but rather a dynamic immunophenotypic transition between AD and PV, presenting the following difficulties in its identification: Extreme class imbalance: FF accounted for only 6.9% (87 / 1257) of the total sample, far below the balanced distribution of psoriasis / eczema in the comparative documents.
[0038] Diagnostic complexity: FF has mixed characteristics of AD and PV, requiring the capture of transitional states.
[0039] In view of the special characteristics of FF, this application has taken the following measures in model training: Measure 1: FF Sample Full Participation Strategy. In K-fold cross-validation, AD and PV are randomly split, while FF, due to its scarce sample, participates in the training of each fold. This method is particularly suitable for situations where the sample ratios differ greatly in this embodiment.
[0040] Measure 2: Use the Focal Loss function to enhance the learning of difficult samples (especially FF samples) by adjusting the α and γ parameters.
[0041] Measure 3: Give high sample weights to the FF class. Apply higher weights to the FF class in the loss function to ensure that the model does not ignore the minority class.
[0042] Meanwhile, in the modeling stage, this embodiment sets dynamic thresholds and uncertainty quantification indicators for FF risk; and designs dual uncertainty indicators of prediction entropy and expert disagreement degree to address the high risk of missed diagnosis of FF, thereby jointly improving the accuracy of prediction and identification of FF phenomenon.
[0043] In one specific embodiment, training the skin phenomenon classification model consisting of S110-S150 may include the following steps: Step 1: Obtain sample sets for atopic dermatitis, psoriasis, and the Flip-Flop phenomenon. Each sample set includes skin lesion images and clinical data from the same patient. Optionally, after obtaining the original sample sets, data augmentation and regularization can be performed, i.e., using enhancement methods such as rotation, flipping, and color jittering to improve the model's generalization ability.
[0044] Step 2: Divide the atopic dermatitis sample set and the psoriasis sample set into multiple subsets, ensuring that the number of samples in each subset is balanced with the number of samples in the Flip-Flop phenomenon sample set. In practical applications, the number of samples in the atopic dermatitis sample set and the psoriasis sample set is much larger than that in the Flip-Flop phenomenon sample set (approximately 6.9%). Therefore, this embodiment divides the larger sample sets into subsets, ensuring that the size of the resulting subsets is comparable to that of the Flip-Flop phenomenon sample set.
[0045] Step 3: Combine the Flip-Flop phenomenon sample set with each subset of atopic dermatitis samples and each subset of psoriasis samples to form balanced sample sets; use each balanced sample set to train the skin phenomenon classification model composed of S110-S150.
[0046] Steps one through three above correspond to measure one mentioned earlier.
[0047] Furthermore, to address class imbalance, a focal loss function can be used during training, by adjusting the hyperparameters in the focal loss function. and This strengthens the learning of the complexity of Flip-Flop phenomenon samples (i.e., measure two mentioned above). Simultaneously, class weights are set in the loss function, with higher weights assigned to the FF class to strengthen the learning of the rarity of Flip-Flop phenomenon samples (i.e., measure three mentioned above).
[0048] In addition, a cross-validation strategy can be adopted during training, dividing the dataset into a training set (80%) and a test set (20%); using the Adam optimizer with a cosine annealing learning rate; employing an early stopping mechanism, stopping training when the validation set performance has not improved for several consecutive rounds; and training multiple models with different initialization parameters, which are then integrated through voting or averaging.
[0049] Furthermore, to verify the effectiveness of the method, this embodiment also constructed two other classification models: One type is an image classification model, which classifies three skin phenomena based on images of patient skin lesions. Its basic structure is as follows: Figure 3 As shown, classification and prediction are performed using only image modalities. Meanwhile, Figure 3 The paper also demonstrates a method for constructing a balanced sample set.
[0050] One approach is a clinical feature classification model, which categorizes three skin conditions based on textual features from clinical data. This model extracts textual features in the same way as the multimodal models S110 to S150 described above, but instead inputs these features into a Support Vector Machine (SVM) classifier for classification. The classifier uses a Radial Basis Function (RBF) to capture non-linear relationships and optimizes the hyperparameters (C and gamma) through grid search, outputting a class probability distribution.
[0051] Figure 4The visualization compares the prediction results of the three modalities, from top to bottom: image classification model prediction results, clinical feature classification model prediction results, and multimodal model prediction results. The multimodal fusion model achieved an accuracy of 97.06% on the test set, with a macro-average AUC of 0.9922, and 100% recall for AD and FF, effectively avoiding missed diagnoses. The F1 score for the FF category reached 96%, significantly outperforming the single-modal method and improving the accuracy of classification and prediction.
[0052] In addition to verifying the above-mentioned verifications, this embodiment also provides an interpretable model verification method. In one specific implementation, the following operations can be performed on multiple samples in the test set respectively: First, extract the feature vector f_original (i.e., the image features output by the derm-foundation model) from the lesion image (i.e., the original image) in the current sample.
[0053] Then, the skin lesion image is divided into multiple regions, and different regions are occluded using gray blocks to obtain different occluded images.
[0054] The trained model is then used to extract the image features of each occluded image, i.e., the occluded feature vector f_occluded. The image feature difference between each occluded image and the skin lesion image is calculated as sensitivity = ||f_original - f_occluded||.
[0055] Finally, attention weights for each occluded region are determined based on the differences in each feature. The greater the feature difference, the more important the occluded region, and the higher its attention weight. The attention of each region can form an attention heatmap, which can enhance model interpretability and verify model performance. For example, if a high-attention region in the attention heatmap matches a lesion region in the sample, it can prove the model's validity. Figure 5 As shown (AD, FF, PV from top to bottom). Similarly, the attention heatmap of the image to be classified in S150 is also generated using the above method.
[0056] In another specific implementation, to further enhance the interpretability of the model, the 21-dimensional features in the clinical features can be ranked by importance. Optionally, the coefficients of each clinical feature in the test set for classifying each phenomenon in each machine learning network can be read. Taking the SVM model as an example, the coef_ matrix (3×21, including the coefficients of the 21 features for the three-class classification) can be read from the SVM model.
[0057] Then, the coefficients of the same clinical feature for the same skin condition classification in each machine learning network are averaged to obtain the importance of the same clinical feature to the same skin condition classification. Based on the importance of each clinical feature to each skin condition classification, key clinical features for each skin condition classification are recommended. For example, for each clinical feature, the absolute values of its coefficients in a certain classification across the four machine learning models are averaged to obtain the importance score of that clinical feature to that classification; the importance score is normalized to a percentage, and the clinical features ranked from highest to lowest importance are recommended as key features for that classification to assist doctors in subsequent decision-making.
[0058] Furthermore, the importance of the same clinical feature in the same machine learning network can be obtained by averaging the coefficients of the same clinical feature across different skin phenomena. Based on the importance of each clinical feature in each machine learning network, the decision-making mechanism of each network can be explained. For example, for each clinical feature, the absolute value of its coefficients in the three classifiers of a machine learning network is averaged to obtain an importance score. This importance score is then normalized to a percentage and sorted from highest to lowest importance. This ranking can explain the decision-making mechanism of the machine learning network, improving the interpretability of the model's decisions. Figure 6 As shown in B.
[0059] In summary, this embodiment provides a skin phenomenon classification method based on multimodal machine learning, which can achieve the following beneficial effects: 1. Improve the accuracy of skin phenomenon classification. Existing methods fail to effectively integrate the visual features of skin lesions with the semantic information of clinical history, resulting in information waste. This embodiment proposes a cross-modal attention mechanism for skin lesion images and clinical history data. Through semantic space unification, cross-attention mechanisms, multi-expert learning of different visual modes, and dynamic fusion using gating networks, it achieves deep fusion of image visual features and clinical semantic information, overcoming the limitations of single data sources and improving the accuracy of skin phenomenon classification. Compared to using images alone (accuracy approximately 90%) or clinical data alone (accuracy approximately 85%), the method in this embodiment improves accuracy by 7-12 percentage points.
[0060] 2. Strong ability to identify rare cases. Because FF cases are relatively rare in clinical practice, existing models perform poorly in identifying a few categories, easily leading to missed diagnoses. However, this embodiment, through a class balancing strategy, focus loss, and dynamic class weights, still achieves a 96% F1 score even when FF cases account for only 6.9% of the total sample, providing a powerful tool for the early identification of rare clinical diseases.
[0061] 3. To achieve quantitative prediction of the probability of FF occurrence, providing support for early clinical intervention. Existing technologies lack quantitative prediction of the risk of FF occurrence, leading to incorrect clinical intervention selection. This embodiment proposes, for the first time, a quantitative assessment of the probability of FF occurrence, providing early warning for patient monitoring during biologics intervention and supporting personalized treatment decisions.
[0062] 5. An interpretable classification model is provided, offering a reliable basis for clinical decision-making. This embodiment clearly displays the skin lesion areas and key clinical features that the model focuses on through methods such as attention weight visualization and key feature ranking output, providing doctors with a basis for decision-making and enhancing clinical trust.
[0063] It should be noted that all user data involved in this application is information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0064] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention, such as... Figure 7 As shown, the device includes a processor 60, a memory 61, an input device 62, and an output device 63; the number of processors 60 in the device can be one or more. Figure 7 Taking a processor 60 as an example; the processor 60, memory 61, input device 62, and output device 63 in the device can be connected via a bus or other means. Figure 7 Taking the example of a connection between China and Israel via a bus.
[0065] The memory 61, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the skin phenomenon classification method based on multimodal machine learning in this embodiment of the invention. The processor 60 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 61, thereby implementing the aforementioned skin phenomenon classification method based on multimodal machine learning.
[0066] The memory 61 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a given function; the data storage area may store data created based on terminal usage. Furthermore, the memory 61 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory, or other non-volatile solid-state storage device. In some instances, the memory 61 may further include memory remotely located relative to the processor 60, which can be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0067] Input device 62 can be used to receive input digital or character information, and to generate key signal inputs related to user settings and function control of the device. Output device 63 may include display devices such as a display screen.
[0068] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the skin phenomenon classification method based on multimodal machine learning of any embodiment.
[0069] The computer storage medium of this invention can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0070] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0071] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0072] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages—such as Java, Smalltalk, and C++—as well as conventional procedural programming languages—such as C or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0073] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.
Claims
1. A skin phenomenon classification method based on multimodal machine learning, characterized in that, include: S110. Obtain skin lesion images and clinical data of the patients to be classified; S120. The skin lesion images and clinical data are converted into image features and text features respectively using an image feature extractor and a text encoder; S130. The image features and text features are fused using a cross-attention mechanism, so that the two features pay attention to each other's relevant parts. S140. Multiple expert networks are used to process the fused features to learn different symptom patterns. Then, the output features of each expert network are fused using gating weights to obtain the final fused features. S150. Using different machine learning networks, predict the probability that the patient belongs to atopic dermatitis, psoriasis, and Flip-Flop phenomenon based on the final fusion features, thereby achieving skin phenomenon classification.
2. The method according to claim 1, characterized in that, The skin lesion images are multiple, each corresponding to a different lesion site on the patient; Accordingly, S120 includes: Perform the following operations on each skin lesion image: S1-1. Use the finely tuned derm-foundation model as an image feature extractor to extract the image features of the current skin lesion image; S1-2. The image features are reduced in dimensionality using F-test, mutual information and neural network respectively, and the most influential top N dimensions are selected by voting. S1-3. Using a multi-scale feature extraction layer, the depth features of the first N dimensions of the image are extracted at different scales, and the depth features at different scales are stitched together to take into account both global and detailed information. By using a multi-head self-attention mechanism to fuse the stitched features of each skin lesion image, the final image features with the same dimension as the text features are obtained.
3. The method according to claim 1, characterized in that, The process involves using multiple expert networks to process the fused features to learn different symptom patterns, including: The fused features are processed using an XGBoost expert network based on Focal Loss to address the class imbalance problem. The fused features are processed using a logistic regression expert network to provide a linear probability baseline; Multilayer perceptron expert networks are used to process the fused features in order to capture nonlinear feature interactions; The fused features are processed using a support vector machine expert network to construct a classification hyperplane in a high-dimensional feature space; The fused features are processed using a K-nearest neighbor expert network, and classification is performed based on instance similarity.
4. The method according to claim 1, characterized in that, Prior to S120, it also included: The atopic dermatitis sample set and the psoriasis sample set were divided into multiple subsets to balance the sample size of each subset with that of the Flip-Flop phenomenon sample set. The Flip-Flop phenomenon sample set is then combined with each atopic dermatitis sample subset and each psoriasis sample subset to form a balanced sample set. The skin phenomenon classification model consisting of S110-S150 was trained using each balanced sample set.
5. The method according to claim 1, characterized in that, Prior to S120, it also included: The skin phenomenon classification model consisting of S110-S150 was trained using the Focal Loss function. During training: By adjusting the hyperparameters in the Focal Loss function, the learning of the complexity of Flip-Flop phenomenon samples is enhanced; By increasing the loss weights for classifying the Flip-Flop phenomenon, we can enhance the learning of the rarity of Flip-Flop phenomenon samples.
6. The method according to claim 1, characterized in that, Prior to S120, it also included: The skin phenomenon classification model consisting of S110-S150 was trained using the sample set; The skin lesion images of each sample were divided into multiple regions, and different regions were occluded to obtain different occluded images; The trained model is used to extract image features from the lesion image and each occluded image, and the differences in image features between each occluded image and the lesion image are calculated. The attention weight of each occluded region is determined based on the differences in image features. The greater the difference in image features, the higher the attention weight of the occluded region. An attention heatmap is formed from these attention weights. The model's effectiveness is verified by comparing the high-attention regions in the attention thermal image with the lesion regions in the sample.
7. The method according to claim 1, characterized in that, The clinical data includes various clinical characteristics; Correspondingly, prior to S120, it also included: The skin phenomenon classification model consisting of S110-S150 was trained using the training set; The importance of the same clinical feature in classifying the same skin condition is obtained by averaging the coefficients of the same clinical feature in various machine learning networks in the test set for classifying the same skin condition. Based on the importance of each clinical feature to the classification of each skin condition, the following are recommended key clinical features for each skin condition classification.
8. The method according to claim 1, characterized in that, The clinical data includes various clinical characteristics; Correspondingly, prior to S120, it also included: The skin phenomenon classification model consisting of S110-S150 was trained using the training set; The importance of the same clinical feature in the same machine learning network is obtained by averaging the coefficients of the same clinical feature for classifying various skin phenomena in the same test set. Based on the importance of each clinical feature in each machine learning network, explain the decision-making mechanism of each machine learning network.
9. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the skin phenomenon classification method based on multimodal machine learning as described in any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the skin phenomenon classification method based on multimodal machine learning as described in any one of claims 1-8.