Essence-Interpretability-Based Multimodal Data Fusion Method for Discriminating Dermatofibrosarcoma Protuberans

The integration of multi-modal data fusion and a transparent prototype network classifier improves the accuracy and transparency of AI-based skin leiomyosarcoma diagnosis by leveraging clinical and pathological images, overcoming data limitations and imbalances.

CN120088250BActive Publication Date: 2025-07-15NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510560832.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-15
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

In the prior art, the imbalance of single modal data and data samples leads to limited diagnostic accuracy of skin augmentation tumors and dermal fibromas, and the lack of transparency in artificial intelligence diagnosis, making it difficult to meet the real-time decision-making needs of doctors.

Method used

Clinical images and pathological images are collected, feature encoding and matrix fusion are performed to generate new modal data, diffusion probability model is constructed, and essentially interpretable prototype network classifiers are established through forward diffusion and reverse denoising training, and the classifier is trained using reconstruction loss, orthogonal loss and classification loss to obtain the identification results of skin adenomas.

Benefits of technology

Through multimodal data fusion and essential interpretable prototype network classifiers, the accuracy and real-time interpretability of skin augmentation tumor diagnosis are improved, the reliability and rationality of the diagnostic process are enhanced, and each prototype vector captures different feature information, supporting doctors' intuitive judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088250B_ABST
    Figure CN120088250B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of medical artificial intelligence technology, and specifically relates to a method for identifying dermatofibrosarcoma protuberans based on essentially interpretable multi-modal data fusion. The method includes: collecting detection data, performing feature encoding and matrix fusion to generate new modal data; constructing a diffusion probability model based on the new modal data, and training the diffusion probability model through forward diffusion and reverse denoising; obtaining synthetic samples by using the trained diffusion probability model to form a category training data set; establishing an essentially interpretable prototype network classifier, training the essentially interpretable prototype network classifier based on the category training data set, and obtaining the identification result of dermatofibrosarcoma protuberans through the trained essentially interpretable prototype network classifier; enabling doctors to more intuitively judge the proximity of a sample to be tested to any typical pathological pattern during the decision-making process, and then identifying the pathological category of dermatofibrosarcoma protuberans and the normal category of dermatofibroma, thereby improving the rationality of the identification result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical artificial intelligence, and particularly to a method for differentiating dermatofibrosarcoma protuberans based on essentially interpretable multi-modal data fusion. Background Art

[0002] As a highly malignant and fast-spreading disease, early diagnosis and treatment of dermatofibrosarcoma protuberans are crucial for improving the survival rate of patients and reducing the risk of death; while dermatofibroma is benign in most cases, but it often has similarities in appearance with dermatofibrosarcoma protuberans, making early differentiation complex and difficult; traditional diagnostic methods mainly rely on doctors' personal experience and limited observation means, and it is difficult to capture the subtle pathological differences between the two, facing great challenges in early diagnosis.

[0003] With the rapid development of artificial intelligence technology, it has shown great potential in automatically extracting potential features in lesion images, making up for the deficiencies of traditional diagnostic methods to a certain extent, but still facing some problems. First, single-modal data often cannot comprehensively display the unique features of dermatofibroma and dermatofibrosarcoma protuberans, limiting the accuracy of diagnosis; second, due to the relatively small number of case samples of dermatofibrosarcoma protuberans and the imbalance of data distribution, it leads to limitations in the classification accuracy of the classification model; and currently, most artificial intelligence-assisted diagnostic technologies lack transparency and have the black box problem, and the interpretable methods of artificial intelligence have lagged development and lack real-time interpretability, making the credibility of the model not high and difficult to meet the needs of doctors for real-time and transparency in clinical decision-making. Summary of the Invention

[0004] In order to solve the technical problems that the existing solutions use single-modal data, and the small number of data samples and uneven distribution lead to limited model classification, and the interpretable methods of artificial intelligence cannot meet the real-time decision-making needs of doctors, the purpose of the present invention is to provide a method for differentiating dermatofibrosarcoma protuberans based on essentially interpretable multi-modal data fusion, and the specific technical solutions adopted are as follows:

[0005] Collect detection data, and perform feature encoding and matrix fusion to generate new modal data;

[0006] Based on the new modal data, construct a diffusion probability model, and train the diffusion probability model through forward diffusion and reverse denoising;

[0007] Use the trained diffusion probability model to obtain synthetic samples to form a category training data set;

[0008] Build an intrinsically interpretable prototype network classifier, train the intrinsically interpretable prototype network classifier based on the class training dataset, and obtain the discrimination result of dermatofibrosarcoma protuberans through the trained intrinsically interpretable prototype network classifier.

[0009] Preferably, collect detection data, perform feature encoding and matrix fusion to generate new modality data, including:

[0010] The detection data includes clinical images and pathological images;

[0011] Based on the clinical images and pathological images, perform feature encoding through an encoder in sequence, and perform feature splicing through matrix fusion to generate a joint representation, denoted as new modality data.

[0012] Preferably, build a diffusion probability model based on the new modality data, and train the diffusion probability model through forward diffusion and reverse denoising, including:

[0013] Add noise to the data elements of the new modality data in sequence in the forward direction to generate intermediate states, and the corresponding calculation formula is:

[0014]

[0015] where, represents the condition satisfied by the th data element to generate the intermediate state ; represents the intermediate state generated by the th data element in the new modality data; represents the th data element of the noise variance; represents the identity matrix; represents the total number of the new modality data;

[0016] Determine the objective function of the reverse denoising training process, and the corresponding calculation formula is:

[0017]

[0018] where, represents the objective function; represents the first data element in the new modality data, that is, the original data; represents Gaussian noise;

[0019] Train the intermediate state based on the objective function to obtain the trained diffusion probability model, and the corresponding calculation formula is:

[0020]

[0021] where, represents from the first data element to the The proportion of the cumulative retained original signal of each data element; Indicates the time step index from the 1st data element to the th data element.

[0022] Preferably, a trained diffusion probability model is used to obtain synthetic samples to form a category training data set, including:

[0023] Extract a noise vector through the forward diffusion process, and use the trained diffusion probability model to combine with the noise vector for reverse diffusion to iteratively generate synthetic samples;

[0024] Integrate the synthetic samples to obtain a dermatofibrosarcoma protuberans sample, and form a category training data set with the real samples.

[0025] Preferably, an intrinsically interpretable prototype network classifier is established, and the intrinsically interpretable prototype network classifier is trained based on the category training data set. The discrimination result of dermatofibrosarcoma protuberans is obtained through the trained intrinsically interpretable prototype network classifier, including:

[0026] The intrinsically interpretable prototype network classifier includes an autoencoder and a prototype classification network;

[0027] Map the category training data set to a low-dimensional latent feature space based on the autoencoder and reconstruct it, and calculate the reconstruction loss;

[0028] Set several prototype vectors in advance according to the prototype classification network, and impose an orthogonality constraint on the prototype vectors of each category to calculate the orthogonality loss;

[0029] Obtain the original score of each category of the target classification through the prototype classification network, and determine the prediction probability of each category to calculate the classification loss;

[0030] Determine the total loss function according to the reconstruction loss, orthogonality loss and classification loss, and train the intrinsically interpretable prototype network classifier based on the total loss function;

[0031] Obtain a sample to be tested, and input the sample to be tested into the trained intrinsically interpretable prototype network classifier to obtain the discrimination result of dermatofibrosarcoma protuberans.

[0032] Preferably, calculate the reconstruction loss, and the corresponding calculation formula is:

[0033]

[0034] Among them, Indicates the reconstruction loss; Indicates the input category training data set; Indicates the reconstructed category training data set.

[0035] Preferably, calculate the orthogonal loss, and the corresponding calculation formula is:

[0036]

[0037] where, represents the orthogonal loss; represents the th category; represents the total number of categories; represents the th category of the matrix composed of prototype vectors; represents 's identity matrix; represents the number of prototype vectors; represents the Frobenius norm.

[0038] Preferably, calculate the classification loss, and the corresponding calculation formula is:

[0039]

[0040] where, represents the classification loss; represents the th total number of categories corresponding to the prototype vector; represents the th prototype vector; represents the th category of the target classification; represents the total number of categories of the target classification; represents the th prototype vector on the th category of the true label; represents the th prototype vector on the th category of the predicted probability.

[0041] Preferably, determine the total loss function according to the reconstruction loss, orthogonal loss and classification loss, and the corresponding calculation formula is:

[0042]

[0043] where, represents the total loss function; , , respectively represent the classification loss, reconstruction loss and orthogonal loss; , , respectively represent the balance parameters corresponding to the classification loss, reconstruction loss and orthogonal loss.

[0044] Preferably, input the sample to be tested into the trained intrinsically interpretable prototype network classifier to obtain the discrimination result of dermatofibrosarcoma protuberans, including:

[0045] Input the sample to be tested into the trained intrinsically interpretable prototype network classifier, and output the prediction probability of each category;

[0046] Arrange the prediction probabilities in descending order, and identify the sample in the category corresponding to the maximum prediction probability as dermatofibrosarcoma protuberans, and the samples in the remaining categories as dermatofibroma.

[0047] The present invention has the following beneficial effects:

[0048] Perform multi-modal data fusion on clinical images and pathological images, expand the display of data features, and solve the limitations caused by single-modal data; and analyze based on clinical images and pathological images at the same time, avoiding the problem of uneven distribution caused by few samples, and preventing the influence of sample problems on the classification results of the classifier; analyze the similarity between the sample to be tested and the prototype vector in the feature space based on the trained intrinsically interpretable prototype network classifier, effectively increasing the diversity of skin lesion prototypes, ensuring that each different skin prototype can capture and represent different feature information, that is, introducing reconstruction loss and orthogonal loss, improving the semantic expression ability of the classifier, so that each prototype vector can correspond to a typical lesion feature pattern; obtain the original scores of each category of the category training data set through the prototype classification network, determine the prediction probability of each category, that is, calculate the classification loss, combine the reconstruction loss and the orthogonal loss, while ensuring the classification accuracy of the classifier, significantly enhancing the real-time interpretability of the diagnosis process, improving the overall reliability of the classifier, and effectively improving the diagnosis efficiency of dermatofibrosarcoma protuberans; in addition, determine the feature vector during the calculation of the orthogonal loss, and calculate the distance from the prototype vector, enabling doctors to more intuitively judge the proximity of the sample to be tested to any typical pathological pattern during the decision-making process, and then distinguish the lesion category of dermatofibrosarcoma protuberans and the normal category of dermatofibroma, improving the rationality of the discrimination result. Description of the Drawings

[0049] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative work.

[0050] Figure 1 It is a flowchart of the steps of a method for discriminating dermatofibrosarcoma protuberans based on intrinsically interpretable multi-modal data fusion provided by an embodiment of the present invention;

[0051] Figure 2 The schematic structural diagram of the essential interpretable prototype network classifier of a method for identifying dermatofibrosarcoma protuberans based on essential interpretable multi-modal data fusion provided by an embodiment of the present invention. Specific Embodiments

[0052] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following combines the accompanying drawings and preferred embodiments to specifically describe the specific embodiments, structures, features and effects of a method for identifying dermatofibrosarcoma protuberans based on essential interpretable multi-modal data fusion proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.

[0054] The following specifically describes the specific solution of a method for identifying dermatofibrosarcoma protuberans based on essential interpretable multi-modal data fusion provided by the present invention with reference to the accompanying drawings.

[0055] Please refer to Figure 1 , which shows the flowchart of the steps of a method for identifying dermatofibrosarcoma protuberans based on essential interpretable multi-modal data fusion provided by an embodiment of the present invention. The method includes:

[0056] Step S1: Collect detection data, perform feature encoding and matrix fusion to generate new modal data;

[0057] Step S2: Build a diffusion probability model based on the new modal data, and train the diffusion probability model through forward diffusion and reverse denoising;

[0058] Step S3: Use the trained diffusion probability model to obtain synthetic samples and form a class training data set;

[0059] Step S4: Establish an essential interpretable prototype network classifier, train the essential interpretable prototype network classifier based on the class training data set, and obtain the identification result of dermatofibrosarcoma protuberans through the trained essential interpretable prototype network classifier.

[0060] For better illustration, in this embodiment, the detection data mainly focuses on those images that can be visually recognized as dermatofibrosarcoma protuberans; in practical applications, dermatofibrosarcoma protuberans and dermatofibroma show great similarity in appearance. Although these two skin lesions belong to malignant and benign respectively in nature, their appearance characteristics are easily confused. To improve the treatment efficiency, this application proposes a method for distinguishing malignant dermatofibrosarcoma protuberans, which can extract more subtle features from images to distinguish the two, facilitating the provision of more accurate diagnosis and treatment plans for patients and avoiding the deterioration of the condition.

[0061] Further, in step S1, it includes:

[0062] The detection data includes clinical images and pathological images;

[0063] Based on the clinical images and pathological images, feature encoding is sequentially performed through an encoder, and feature splicing is performed through matrix fusion to generate a joint representation, denoted as new modality data.

[0064] Specifically, the clinical image is denoted as , and the pathological image is denoted as ; the features corresponding to each image are obtained through feature encoding by the encoder, that is, respectively , , where represents converting the pixels of the image into a numerical matrix; and matrix fusion operation is performed through linear transformation after feature splicing to generate a joint representation, denoted as new modality data, that is , where represents splicing two features; represents the weight matrix; represents the bias term; represents the activation function.

[0065] Further, in step S2, it includes:

[0066] Step S21: According to the data elements of the new modality data, noise is added sequentially in the forward direction to generate intermediate states, and the corresponding calculation formula is:

[0067]

[0068] where represents the condition satisfied by the -th data element to generate the intermediate state , represents the intermediate state generated by the -th data element in the new modality data; represents the noise variance of the -th data element; represents the identity matrix; Indicates the total number of new modality data.

[0069] Specifically, the new modality data is composed of multiple data elements, that is, the new modality data includes image pairs composed of multiple groups of pathological images and clinical images; starting from the first data element of the new modality data, noise is added in turn to generate an intermediate state according to the corresponding satisfied conditions, that is, the original data The intermediate state generated after adding noise is , and so on, to obtain the intermediate states generated by all data elements of the new modality data, denoted as .

[0070] Step S22: Determine the objective function of the reverse denoising training process, and the corresponding calculation formula is:

[0071]

[0072] Among them, represents the objective function; represents the first data element in the new modality data, that is, the original data; represents Gaussian noise.

[0073] Step S23: Train the intermediate state based on the objective function to obtain a trained diffusion probability model, and the corresponding calculation formula is:

[0074]

[0075] Among them, represents the proportion of the original signal accumulated from the first data element to the th data element; represents the time step index between the first data element and the th data element.

[0076] Specifically, perform the corresponding reverse operation based on the conditions satisfied by adding noise, that is, perform reverse for to obtain , and trace back the data element from the noise state to the original state, where the original data refers to the data element without added noise, that is, the image data; and train the diffusion probability model by minimizing the simplified objective function to ensure that the model can accurately learn the mapping from the noise state to the original state; among them, represents the proportion of the original signal accumulated from the first data element to the th data element. The larger this number is, the higher the proportion of the original information retained. On the contrary, the smaller the value, the higher the proportion of noise. By training the diffusion probability model, the discrimination accuracy of noise-interfered data is improved, so that while maintaining the stability of the model, the discrimination of malignant dermatofibrosarcoma protuberans is enhanced.

[0077] Further, in step S3, it includes:

[0078] Step S31: Extract a noise vector through the forward diffusion process, and use the trained diffusion probability model combined with the noise vector for reverse diffusion to iteratively generate synthetic samples;

[0079] Step S32: Integrate the synthetic samples to obtain dermatofibrosarcoma protuberans samples, and form a class training data set with real samples.

[0080] Specifically, sample a noise vector from the standard normal distribution based on the forward diffusion process in step 2, denoted as , use the trained diffusion probability model combined with the noise vector for reverse diffusion to iteratively generate synthetic samples, denoted as , and similarly obtain synthetic samples generated based on all noise vectors, and integrate them to obtain synthetic dermatofibrosarcoma protuberans samples , to create high-quality synthetic data and make it closer to the distribution of actual sample data; then, construct a class training data set, denoted as , where represents real samples, that is, samples composed of existing data, including benign skin lesions or normal skin states; represents malignant dermatofibrosarcoma protuberans samples, represents the label of dermatofibrosarcoma protuberans samples; it can be explained that the class training data set is composed of real samples and dermatofibrosarcoma protuberans samples, enabling the subsequent classifier to learn to distinguish between benign and malignant skin lesions, and being able to accurately identify benign dermatofibromas, that is, enhancing the discrimination ability for new modality data and improving the discrimination ability for malignant dermatofibrosarcoma protuberans.

[0081] Further, in step S4, it includes:

[0082] The intrinsically interpretable prototype network classifier includes an autoencoder and a prototype classification network.

[0083] It can be explained that the intrinsically interpretable prototype network classifier is used for the auxiliary diagnosis and classification of dermatofibrosarcoma protuberans for the class training data set formed after fusion and enhancement; among them, the autoencoder includes an encoder and a decoder. The encoder is used to map the input class training data set to a low-dimensional latent feature space, that is, to extract feature vectors; the decoder is then used to remap the latent feature back to the original data space to reconstruct the feature vectors; it can constrain the intrinsically interpretable prototype network classifier to learn semantically clear feature representations, providing a basis for subsequent classification of samples to be measured to be interpretable; the prototype classification network includes three sub-layers: a prototype layer, a fully connected layer, and a Softmax layer, which are used to analyze the feature vectors extracted by the encoder and output predicted class probabilities for classification and discrimination.

[0084] Step S41: Map the category training data set to a low-dimensional latent feature space based on the autoencoder and perform reconstruction, and calculate the reconstruction loss.

[0085] Specifically, the encoder is used to compress the input category training data set into low-dimensional latent features, and the corresponding feature vectors that fully express the original input spectral meaning are extracted, which helps to reduce the complexity of the data and remove redundant information, improve the generalization ability of the essentially interpretable prototype network classifier, and enable it to learn more abstract and representative data elements. Using low-dimensional data is to improve the data calculation and storage efficiency; then it is transmitted to the decoder for reconstruction, and the mean square error is used to calculate the reconstruction loss. If the reconstruction loss is small, it means that the feature vectors extracted by the encoder retain sufficient information, and the information of the data elements in the category training data set is better retained through the reconstruction loss to achieve a better reconstruction effect.

[0086] Further, in step S41, calculate the reconstruction loss, and the corresponding calculation formula is:

[0087]

[0088] where represents the reconstruction loss; represents the input category training data set; represents the reconstructed category training data set.

[0089] Step S42: Preset a number of prototype vectors according to the prototype classification network, apply an orthogonality constraint to each category of prototype vectors, and calculate the orthogonality loss.

[0090] It can be explained that by presetting a number of prototype vectors in the prototype layer according to the prototype classification network, the learning parameters can be obtained; it is explained that each prototype vector is located in the feature space output by the encoder, and the dimension is the same as the output feature of the encoder, and each prototype vector represents a representative feature pattern, that is, a typical expression of a corresponding category feature; in the training process of the essentially interpretable prototype network classifier, the position of the prototype vector is continuously updated using the feature vectors extracted by the encoder, so that the prototype vector can approximate some representative feature regions in the category training data set.

[0091] Specifically, after the feature vectors are extracted by the encoder, the prototype layer compares the feature vectors with each prototype vector one by one and calculates the similarity measure between the two; preferably, in this embodiment, the Euclidean distance is used to measure the degree of difference between the feature vector and the prototype vector, that is, the feature vector in the encoder is defined as , the th prototype vector is , calculate the squared Euclidean distance between the two as ; the smaller the distance, the higher the similarity between the feature vector and the prototype vector, indicating that the feature is closer to the prototype; conversely, the larger the distance, the greater the difference; further determine the distances of all prototype vectors, denoted as , represents the total number of prototype vectors; among them, an orthogonal constraint is imposed on the prototype vectors within each category to prevent multiple prototype vectors from collapsing to the same position; finally, it enters the fully connected layer for processing.

[0092] Further, in step S42, calculate the orthogonal loss, and the corresponding calculation formula is:

[0093]

[0094] where, represents the orthogonal loss; represents the -th category; represents the total number of categories; represents the -th category of the matrix composed of prototype vectors; represents 's identity matrix; represents the number of prototype vectors; represents the Frobenius norm.

[0095] It should be noted that the orthogonal constraint requires different prototype vectors within the same category to be orthogonal to each other, avoiding all pairwise prototype vectors from converging to the class center to increase the diversity of prototype vectors, that is, to expand the diversification of feature patterns and ensure that each prototype vector captures different discriminative features of the category; among them, represents the Frobenius norm, that is, the square root of the sum of the squares of all elements of the matrix.

[0096] Step S43: Obtain the original scores of each category of the target classification through the prototype classification network, determine the prediction probability of each category, and calculate the classification loss.

[0097] It can be explained that the fully connected layer is used to calculate the original scores according to the output of the prototype layer for classification; specifically, the fully connected layer introduces a learnable weight matrix , with a size of , represents the number of prototype vectors, represents the number of categories of the target classification. The fully connected layer calculates the original scores for each category and uses them as a linear combination of the corresponding prototype vectors. The corresponding calculation formula is:

[0098]

[0099] Among them, represents the original score of the th category; represents the number of prototype vectors; represents the element in the th row and th column of the weight matrix, that is, the linear contribution weight of the th prototype vector to the th category; is used to convert the squared Euclidean distance between the prototype vector and the feature vector into a linearly weighted form. Preferably, in this embodiment, taking the negative value or the Gaussian kernel function is adopted to map the distance to a similarity score; represents the th category corresponding bias term.

[0100] It should be noted that when the prototype vector is more similar to the input feature vector, that is, the Euclidean distance between the two vectors is smaller, through appropriate conversion, the score contribution of this prototype vector to the category is greater; according to the training and learning of the weight matrix , the fully connected layer will gradually strengthen the correspondence between the prototype vector and the category of the target classification. Among them, assuming that the th prototype vector is representative for the th category, its weight will be learned as a relatively large positive value on the corresponding th category. On the contrary, it is a relatively small value or negative value on the non-corresponding category, so that the fully connected layer determines that the prototype vector mainly contributes to discriminating its corresponding category, ensuring that each prototype vector affects the judgment of each category to different degrees, so as to automatically adjust the prototype vector to more importantly support the decision of the corresponding category.

[0101] Finally, the Softmax layer normalizes all category scores output by the fully connected layer, that is, , and converts it into a probability distribution for each category; that is, the Softmax function will map the original score of each category to a probability value between 0 and 1, and make the sum of the probabilities of all categories equal to 1. Furthermore, based on the predicted probability that the input feature vector belongs to each category, where the predicted probability of each output category is , and the true label is represented by one-hot encoding as .

[0102] Furthermore, in step S43, the classification loss is calculated, and the corresponding calculation formula is:

[0103] Calculate the classification loss, and the corresponding calculation formula is:

[0104]

[0105] Among them, represents the classification loss; represents the total number of categories corresponding to the th prototype vector; represents the th prototype vector; represents the th category of the target classification; represents the th prototype vector on the th category; represents the th prototype vector on the th category;

[0106] Step S44: Determine the total loss function according to the reconstruction loss, orthogonal loss, and classification loss, and train the intrinsically interpretable prototype network classifier based on the total loss function.

[0107] Furthermore, in step S44, the total loss function is determined according to the reconstruction loss, orthogonal loss, and classification loss, and the corresponding calculation formula is:

[0108]

[0109] Among them, represents the total loss function; , , respectively represent the classification loss, reconstruction loss, and orthogonal loss; , , respectively represent the balance parameters corresponding to the classification loss, reconstruction loss, and orthogonal loss.

[0110] It should be noted that the three losses cooperate with each other to not only improve the classification accuracy of the category training data set, but also keep the feature space where the prototype vectors are located having an interpretable structure. Since the decoder is introduced, the trained prototype vectors are input into the decoder to reconstruct approximate original input data. In this embodiment, an image close to any type of skin lesion is generated based on the provided image data, indicating that each prototype vector corresponds to a typical sample feature pattern that can be intuitively understood by people. It can be understood that after the training of the autoencoder and the prototype classification network, the trained intrinsically interpretable prototype network classifier is obtained, which can improve the real-time interpretability ability while maintaining good classification performance.

[0111] Step S45: Obtain a sample to be tested, and input the sample to be tested into the trained intrinsically interpretable prototype network classifier to obtain the discrimination result of dermatofibrosarcoma protuberans.

[0112] Further, in step S45, inputting the sample to be tested into the trained intrinsically interpretable prototype network classifier to obtain the discrimination result of dermatofibrosarcoma protuberans includes:

[0113] Step S451: Input the sample to be tested into the trained intrinsically interpretable prototype network classifier, and output the prediction probability of each category;

[0114] Step S452: Arrange the prediction probabilities in descending order, identify the sample in the category corresponding to the maximum prediction probability as dermatofibrosarcoma protuberans, and identify the samples in the remaining categories as dermatofibroma.

[0115] It should be noted that the sample to be tested refers to the clinical image and the corresponding pathological image that are visually identified as dermatofibrosarcoma protuberans. The low-dimensional latent features of the sample to be tested are extracted by the encoder , and the Euclidean distance is calculated with each prototype vector in the prototype layer to generate a set of distance metrics ; Subsequently, the fully connected layer linearly weights the Euclidean distances corresponding to each prototype vector by combining the pre-learned weight matrix to obtain the raw score of the category of the target classification. After normalizing the raw score through the Softmax layer, the prediction probability of each corresponding category is output; finally, the intrinsically interpretable prototype network classifier selects the category corresponding to the maximum prediction probability as the classification diagnosis result, that is, the sample in this category is identified as dermatofibrosarcoma protuberans; on the contrary, the sample corresponding to the category with a smaller prediction probability is a normal skin texture change, that is, dermatofibroma.

[0116] Understandably, multi-modal data fusion is performed on clinical images and pathological images to expand the display of data features and solve the limitations caused by single-modal data. Moreover, by analyzing clinical images and pathological images simultaneously, the problem of unbalanced distribution caused by a small number of samples is avoided, preventing the classification results of the classifier from being affected by sample problems. Based on the trained essentially interpretable prototype network classifier, the similarity between the sample to be tested and the prototype vector in the feature space is analyzed, effectively increasing the diversity of skin lesion prototypes and ensuring that each different skin prototype can capture and represent different feature information. That is, the reconstruction loss and the orthogonal loss are introduced to improve the semantic expression ability of the classifier, so that each prototype vector can correspond to a typical lesion feature pattern. The original scores of each category of the category training data set are obtained through the prototype classification network, and the prediction probability of each category is determined, that is, the classification loss is calculated, combined with the reconstruction loss and the orthogonal loss. While ensuring the classification accuracy of the classifier, the real-time interpretability of the diagnosis process is significantly enhanced, the overall reliability of the classifier is improved, and the diagnosis efficiency of dermatofibrosarcoma protuberans is effectively improved. In addition, during the calculation of the orthogonal loss, the feature vector is determined and the distance is calculated with the prototype vector, enabling doctors to more intuitively judge the proximity of the sample to be tested to any typical pathological pattern during the decision-making process, and then distinguish the lesion category of dermatofibrosarcoma protuberans and the normal category of skin fibroma, improving the rationality of the discrimination results.

[0117] It should be noted that the above sequence of embodiments of the present invention is only for description and does not represent the advantages or disadvantages of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or may be advantageous.

[0118] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments.

Claims

1. An intrinsically interpretable multi-modal data fusion method for differentiating dermatofibrosarcoma protuberans, characterized in that, The method includes: Collecting detection data, performing feature encoding and matrix fusion to generate new modality data, including: The detection data includes clinical images and pathological images, where the clinical images are denoted as , and the pathological images are denoted as ; Based on clinical images and pathological images, successively performing feature encoding through an encoder and performing feature splicing through matrix fusion to generate a joint representation, denoted as new modality data, including: Feature encoding is performed through an encoder to obtain the features corresponding to each clinical image and pathological image, namely, respectively , , where represents converting the pixels of the image into a numerical matrix; and matrix fusion operation is performed through linear transformation after feature splicing to generate a joint representation, denoted as new modality data, namely , where represents splicing two features; represents the weight matrix; represents the bias term; represents the activation function; Constructing a diffusion probability model based on the new modality data and training the diffusion probability model through forward diffusion and reverse denoising; Using the trained diffusion probability model to obtain synthetic samples to form a class training data set, including: Extracting noise vectors through the forward diffusion process, and performing reverse diffusion by combining the trained diffusion probability model with the noise vectors to iteratively generate synthetic samples; Integrating the synthetic samples to obtain dermatofibrosarcoma protuberans samples, and forming a class training data set with real samples, including: Similarly, synthetic samples generated based on all noise vectors are obtained and integrated to obtain synthetic dermatofibrosarcoma protuberans samples , and a class training dataset is constructed, denoted as , where represents real samples, including benign skin lesions or normal skin states; represents malignant dermatofibrosarcoma protuberans samples, represents the label of dermatofibrosarcoma protuberans samples; Establishing an intrinsically interpretable prototype network classifier, training the intrinsically interpretable prototype network classifier based on the class training data set, and obtaining the discrimination result of dermatofibrosarcoma protuberans through the trained intrinsically interpretable prototype network classifier, including: The intrinsically interpretable prototype network classifier includes an autoencoder and a prototype classification network; Mapping the class training data set to a low-dimensional latent feature space through the autoencoder and performing reconstruction, and calculating the reconstruction loss; Presetting a number of prototype vectors according to the prototype classification network, imposing an orthogonality constraint on the prototype vectors of each class, and calculating the orthogonality loss; Obtaining the original scores of each class of the target classification through the prototype classification network, determining the prediction probability of each class, and calculating the classification loss; Determining the total loss function according to the reconstruction loss, orthogonality loss and classification loss, and training the intrinsically interpretable prototype network classifier based on the total loss function; Obtaining a sample to be tested, and inputting the sample to be tested into the trained intrinsically interpretable prototype network classifier to obtain the discrimination result of dermatofibrosarcoma protuberans.

2. The method for discriminating dermatofibrosarcoma protuberans based on intrinsically interpretable multimodal data fusion according to claim 1, wherein Constructing a diffusion probability model based on the new modality data and training the diffusion probability model through forward diffusion and reverse denoising, including: Successively adding noise to the data elements of the new modality data in the forward direction to generate intermediate states, and the corresponding calculation formula is: ; Among them, represents the condition satisfied by generating an intermediate state for the th data element, and represents the intermediate state generated by the th data element in the new modal data; represents the noise variance of the th data element; represents the identity matrix; represents the total number of new modal data. Determining the objective function of the reverse denoising training process, and the corresponding calculation formula is: ; Among them, represents the objective function; represents the first data element in the new modality data, i.e., the original data; represents Gaussian noise; Training the intermediate states based on the objective function to obtain the trained diffusion probability model, and the corresponding calculation formula is: ; Among them, represents the proportion of the original signal cumulatively retained from the 1st data element to the th data element; represents the time step index between the 1st data element and the th data element.

3. The method for discriminating dermatofibrosarcoma protuberans based on essentially interpretable multimodal data fusion according to claim 1, wherein Calculating the reconstruction loss, and the corresponding calculation formula is: ; Among them, represents the reconstruction loss; represents the input categorical training dataset; represents the reconstructed categorical training dataset.

4. The method for discriminating dermatofibrosarcoma protuberans based on essentially interpretable multimodal data fusion according to claim 1, wherein Calculating the orthogonality loss, and the corresponding calculation formula is: ; Among them, represents the orthogonal loss; represents the th category; represents the total number of categories; represents the th matrix composed of prototype vectors in the th category; represents the identity matrix of represents the number of prototype vectors; represents the Frobenius norm.

5. The method for discriminating dermatofibrosarcoma protuberans based on essentially interpretable multimodal data fusion according to claim 1, wherein Calculating the classification loss, and the corresponding calculation formula is: ; Among them, represents the classification loss; represents the total number of categories corresponding to the th prototype vector; represents the th prototype vector; represents the th category of the target classification; represents the total number of categories of the target classification; represents the true label of the th prototype vector on the th category; represents the predicted probability of the th prototype vector on the th category.

6. The method for discriminating dermatofibrosarcoma protuberans based on essentially interpretable multimodal data fusion according to claim 1, wherein Determining the total loss function according to the reconstruction loss, orthogonality loss and classification loss, and the corresponding calculation formula is: ; Among them, represents the total loss function; , , represent the classification loss, reconstruction loss, and orthogonal loss respectively; , , represent the balance parameters corresponding to the classification loss, reconstruction loss, and orthogonal loss respectively.

7. The method for identifying dermatofibrosarcoma protuberans based on essentially interpretable multimodal data fusion according to claim 1, wherein Inputting the sample to be tested into the trained intrinsically interpretable prototype network classifier to obtain the discrimination result of dermatofibrosarcoma protuberans, including: Inputting the sample to be tested into the trained intrinsically interpretable prototype network classifier, and outputting the prediction probability of each class; Sorting the prediction probabilities in descending order, and discriminating the samples in the class corresponding to the maximum prediction probability as dermatofibrosarcoma protuberans, and discriminating the samples in the remaining classes as dermatofibroma.

Citation Information

Patent Citations

  • Multi-modal medical image fusion and disease prediction method, computer program and terminal

    CN118967480A

  • Systems and methods for image compositing

    US20250022099A1