A Multi-modal Skin Disease Analysis Method Based on Deep Learning

Through photoacoustic imaging combined with dermatological imaging, a multimodal skin analysis model is constructed, which solves the problem of insufficient imaging depth and resolution in the prior art, and achieves efficient and accurate diagnosis of skin diseases, enhancing the interpretability of the model.

CN120260892BActive Publication Date: 2025-08-05AFFILIATED HOSPITAL OF WEIFANG MEDICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510732510.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-08-05
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

In the diagnosis of dermatology, the depth of optical dermatoscope imaging is limited, the spatial resolution and microvascular display ability of ultrasound imaging are poor, and biochemical information cannot be obtained. Multimodal skin disease auxiliary diagnosis technology is lacking.

Method used

Photoacoustic imaging equipment is used to scan skin diseases, combine the images taken by dermatoscopes, and construct and pre-train the skin analysis model based on multimodal skin data. The convolution and attention convolution blocks can be separated into depth, and the feature map is enhanced by state space enhancement modules, and skin disease analysis is analyzed by combining feature pyramid networks and path aggregation networks.

Benefits of technology

It realizes high-resolution, multivariate data-supported skin disease analysis, improves the accuracy and reliability of diagnosis, enhances the interpretability of the model, and avoids the situation where multimodal data information does not match the analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260892B_ABST
    Figure CN120260892B_ABST
Patent Text Reader

Abstract

The present invention relates to a multimodal skin disease analysis method based on deep learning, and relates to the field of auxiliary diagnosis technology. The present invention uses a photoacoustic imaging device to scan, and a dermatoscope to photograph the affected area of the skin disease to obtain photoacoustic imaging data and skin images of the affected area; aligns the photoacoustic imaging with the skin image, fills the space according to the skin image, and then combines it with the skin image to obtain multimodal skin data; constructs and pre-trains a skin analysis model based on multimodal skin data; the pre-trained skin analysis model analyzes the collected multimodal skin data to obtain skin disease analysis results. The present application constructs multimodal skin data, provides multivariate data support, and ensures that disease analysis is accurate and reliable. The skin analysis model extracts skin image feature maps and photoacoustic imaging feature maps of different scales to cover the association between features of different sizes and skin diseases of different specifications; uses another modality data to strengthen its own modality to better combine the two modality features; and has strong interpretability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of auxiliary diagnosis of skin multimodal data, and in particular to a multimodal skin disease analysis method based on deep learning. Background Art

[0002] Skin diseases are among the most common diseases in humans and are characterized by structural and functional changes in the components of skin tissue. Imaging technology plays an important role in the diagnosis and analysis of skin diseases. As a non-invasive means of observation and diagnosis, imaging technology provides valuable information to clinicians. Traditional optical dermoscopy, such as confocal microscopy and optical coherence tomography, can only observe the morphology of the epidermis and superficial dermis due to imaging depth limitations. Ultrasound imaging can visualize the entire skin structure by penetrating deep layers, but its spatial resolution and ability to display microvasculature are poor, and it is unable to obtain biochemical information related to metabolism.

[0003] Photoacoustic imaging directly measures the optical absorption properties of tissue, thereby facilitating the diagnosis of skin diseases. Photoacoustic imaging combines the advantages of optical and ultrasonic imaging. By irradiating the skin with short pulses of laser light and then receiving optical contrast generated by the different absorption of light by endogenous pigments (hemoglobin, melanin, lipids, collagen, glucose, etc.), and ultrasonic signals generated by the thermal expansion effect of the endogenous pigments that absorb light, photoacoustic imaging can provide clinicians with high-contrast, high-resolution skin morphological, functional, and pathological information. It has been used to image melanoma, café au lait spots, psoriasis, and skin vascular diseases. Despite progress in photoacoustic imaging research, multimodal skin disease auxiliary diagnosis technologies based on the optical and acoustic properties of multi-layered skin tissue are still lacking. Summary of the Invention

[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present invention provides a multimodal skin disease analysis method based on deep learning.

[0005] In a first aspect, the present invention provides a multimodal skin disease analysis method based on deep learning, comprising:

[0006] Use photoacoustic imaging equipment to scan and dermatoscope to photograph the affected area of the skin disease to obtain photoacoustic imaging data and skin images of the affected area;

[0007] Align the photoacoustic imaging with the skin image, fill the space according to the skin image, and then combine it with the skin image to obtain multimodal skin data containing skin image and photoacoustic imaging information;

[0008] Construct and pre-train a skin analysis model based on multimodal skin data; the skin analysis model includes: two depthwise separable convolutions that respectively extract skin imaging and photoacoustic imaging modal features from the multimodal skin data, each of which is followed by a backbone network, each of which contains a five-layer attention convolution block and a feature map encoding unit formed by downsampling; a state-space enhancement module connected to the two backbone networks, wherein the feature maps of any modality output by the last three feature map encoding units of the two backbone networks are enhanced by the state-space enhancement module; the enhanced feature maps of the two modalities after each state-space enhancement module are processed by a two-way feature pyramid network and a path aggregation network, fused through a fully connected layer, and transmitted to the skin disease analysis head and mask prediction head respectively;

[0009] The pre-trained skin analysis model analyzes the collected multimodal skin data to obtain skin disease analysis results.

[0010] Furthermore, aligning the photoacoustic imaging with the skin image, filling the space according to the skin image, and then combining the photoacoustic imaging with the skin image to obtain multimodal skin data containing skin image and photoacoustic imaging information includes:

[0011] Markers are set in advance on the affected area of the skin disease for aligning the data of different modalities, and the markers can provide at least three points for calculating the affine matrix;

[0012] When collecting data on skin disease affected areas, the scanning area of the photoacoustic imaging device and the imaging area of the dermatoscope cover the mark;

[0013] stacking the scanned photoacoustic imaging data and arranging them in order, and then transposing the stacked photoacoustic imaging data so that they are aligned with the space of the skin image;

[0014] After spatial alignment, the pixels of the skin image and the corresponding photoacoustic imaging data are aligned based on the representation of the pre-set markers in the photoacoustic image and the skin image;

[0015] After the photoacoustic imaging data is fully aligned with the skin image, the photoacoustic imaging data is filled according to the resolution of the skin image;

[0016] Skin images and photoacoustic imaging data are normalized to form multimodal skin data.

[0017] Furthermore, after the spatial alignment, aligning the pixels of the skin image with the corresponding photoacoustic imaging data based on the representation of the pre-set markers in the photoacoustic imaging and the skin image includes:

[0018] For skin images, edges are extracted by edge detection; and skin image marker edges are extracted from the extracted edges using marker edge features;

[0019] For photoacoustic imaging data, according to the sensitivity of the laser to the fluorescent agent, the layer corresponding to the corresponding laser imaging data is obtained from the stacked photoacoustic imaging data, the edge is extracted by edge detection in the layer, and the photoacoustic imaging data marking edge is extracted from the extracted edge using the marking edge feature;

[0020] Find the positions of at least three pairs of paired points from the edge of the photoacoustic imaging data mark and the edge of the skin image mark;

[0021] An affine transformation matrix between the spatially aligned skin image and the photoacoustic imaging data is calculated using the pairing points, and the affine transformation matrix is applied to the skin image or the photoacoustic imaging data to achieve complete alignment.

[0022] Furthermore, the attention convolution block includes: a layer of depth-wise separable convolution, the output of the depth-wise separable convolution is copied into two branches, one branch passes through the channel attention and spatial attention, and the other branch is transmitted to the convolution layer in parallel with the channel attention and spatial attention; the outputs of the channel attention and spatial attention are added and combined with the output of the convolution layer, wherein the channel attention includes: adaptive average pooling for extracting channel attention weights, convolution layer, ReLU activation function, convolution layer and Sigmoid function, and two convolution layers are used to compress and restore channel dimensions; the spatial attention includes: convolution layer and Sigmoid function for extracting spatial attention weights.

[0023] Furthermore, each of the state space enhancement modules performs independent patch embedding coding of multiple groups of different patch sizes on the input skin image feature map and photoacoustic imaging feature map of any scale;

[0024] Each state space enhancement module alternately connects the patch embedding codes of the corresponding patch sizes of the different modality feature maps generated by it at the pixel level to obtain a fused embedding;

[0025] Each state-space augmentation module projects each fused embedding using a separate linear layer;

[0026] Each state-space enhancement module projects each fused embedding using an independent linear layer and uses an independent state-space model to model the association between the two modalities based on each fused embedding to enhance its own modality with the data of the other modality;

[0027] The output of each state space model is mapped by a linear layer and then added to the corresponding fusion embedding to obtain a sub-enhanced fusion embedding;

[0028] All sub-enhanced fusion embeddings are aggregated to obtain the enhanced fusion embedding, and the enhanced fusion embedding is reshaped in reverse order in the form of arbitrary modality feature maps and split into enhanced feature maps of two modalities.

[0029] Furthermore, each state space enhancement module alternately connects the patch embedding codes of the corresponding patch sizes of the different modal feature maps generated by it at the pixel level to obtain a fused embedding including:

[0030] Traverse the corresponding patch embedding encoding pixels of different modality feature maps by rows and columns;

[0031] Directly and alternately splice two patches of corresponding specifications in sequence to embed the encoded patches of all even columns in the same row;

[0032] Reverse the pixel index order of all odd-numbered columns in the same row of the two patches with corresponding specifications, and then alternately splice the patch embedding coded pixels in order;

[0033] The concatenation results of the patch embedding coded pixels in each row are stacked in row order to form a fused embedding.

[0034] Furthermore, in the feature pyramid network, the enhanced feature maps of the two modalities of the last three feature map encoding units of the backbone network are respectively adjusted to a fixed value through the horizontally connected 1×1 convolutional layer;

[0035] Skin image enhancement feature map or photoacoustic imaging enhancement feature map The output of the horizontally connected 1×1 convolution layer is used as the skin image enhancement pyramid feature map or photoacoustic imaging enhanced pyramid feature map ;

[0036] Enhance the pyramid feature map of skin image or photoacoustic imaging enhanced pyramid feature map Upsampling to skin image enhancement feature map or photoacoustic imaging enhancement feature map Resolution, and skin image enhancement feature map or photoacoustic imaging enhancement feature map The output features of the horizontally connected 1×1 convolutional layer are added to generate the skin image enhancement pyramid feature map or photoacoustic imaging enhanced pyramid feature map ;

[0037] Enhance the pyramid feature map of skin image or photoacoustic imaging enhanced pyramid feature map Upsampling to skin image enhancement feature map or photoacoustic imaging enhancement feature map Resolution, and skin image enhancement feature map and photoacoustic imaging enhancement feature map The output features of the horizontally connected 1×1 convolutional layer are added to generate the skin image enhancement pyramid feature map or photoacoustic imaging enhanced pyramid feature map .

[0038] Furthermore, the path aggregation network introduces a bottom-up path based on the feature pyramid network, including:

[0039] Enhance the pyramid feature map of skin image or photoacoustic imaging enhanced pyramid feature map Downsampling to skin image enhancement pyramid feature map or photoacoustic imaging enhanced pyramid feature map resolution and enhance the pyramid feature map with skin image or photoacoustic imaging enhanced pyramid feature map Add together to generate skin image enhanced pyramid feature aggregation feature map or photoacoustic imaging enhanced pyramid feature aggregation feature map ;

[0040] Skin image enhancement pyramid feature aggregation feature map or photoacoustic imaging enhanced pyramid feature aggregation feature map Downsampling to skin image enhancement pyramid feature map or photoacoustic imaging enhanced pyramid feature map resolution and enhance the pyramid feature map with skin image or photoacoustic imaging enhanced pyramid feature map Add together to generate skin image enhanced pyramid feature aggregation feature map or photoacoustic imaging enhanced pyramid feature aggregation feature map .

[0041] Furthermore, the skin disease analysis head is implemented based on full connection and is used to predict the type of skin disease; the mask prediction head is implemented based on full connection fusion and is used to predict the multimodal skin data area mask related to the skin disease type.

[0042] Furthermore, the loss function used to train the skin analysis model is the sum of the skin disease classification loss and the multimodal skin data region mask IOU loss related to the skin disease type. The skin analysis model parameters are adjusted with the goal of minimizing the loss function.

[0043] In a second aspect, the present invention provides a multimodal skin disease analysis device based on deep learning, comprising: at least one processing unit, the processing unit being connected to a storage unit via a bus unit, the storage unit storing a computer program, and when the computer program is executed by the processing unit, the multimodal skin disease analysis method based on deep learning is implemented.

[0044] In a third aspect, the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the multimodal skin disease analysis method based on deep learning.

[0045] The above technical solution provided by the embodiment of the present invention has the following advantages compared with the prior art:

[0046] This application uses a photoacoustic imaging device to scan and dermatoscope the affected skin area to obtain photoacoustic imaging data and skin images. The photoacoustic image is aligned with the skin image, filling the space corresponding to the skin image, and then combined with the skin image to obtain multimodal skin data containing both skin image and photoacoustic imaging information. The multimodal skin data combines skin images reflecting the skin surface with photoacoustic images reflecting information from different layers of the skin. This provides multivariate data support for subsequent skin disease analysis, ensuring the accuracy and reliability of the analysis.

[0047] This application constructs and pre-trains a skin analysis model based on multimodal skin data to assist in skin disease analysis. The skin analysis model comprises: two depthwise separable convolutions that extract features from the skin image and photoacoustic imaging modalities in the multimodal skin data, respectively. Each depthwise separable convolution is followed by a backbone network, each backbone network comprising a five-layer attention convolution block and a feature map encoding unit formed by downsampling; a state-space enhancement module connected to the two backbone networks, wherein the feature maps of any modality output by the last three feature map encoding units of the two backbone networks are enhanced by the state-space enhancement module; the enhanced feature maps of the two modalities after each state-space enhancement module are processed by a two-way feature pyramid network and a path aggregation network, fused through a fully connected layer, and transmitted to the skin disease analysis head and mask prediction head, respectively; the pre-trained skin analysis model analyzes the collected multimodal skin data to obtain skin disease analysis results. The skin analysis model uses two backbone networks to extract skin image feature maps and photoacoustic imaging feature maps of different scales, respectively. The different-scale feature maps focus on the relationship between multimodal features of different granularity and skin diseases, thereby covering the association between features of different sizes and skin diseases of different specifications. The state space enhancement module uses skin images to enhance photoacoustic imaging, and uses photoacoustic imaging to enhance skin images, so as to strengthen its own modality with the data of another modality and better combine the characteristics of the two modalities for auxiliary analysis.

[0048] In addition, this application task combines mask prediction and auxiliary analysis of skin diseases, and uses mask prediction related to skin diseases to enhance the interpretability of the model, avoiding the situation where the skin disease analysis is correct, but the multimodal skin data information based on the disease analysis does not match the disease analysis results. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0050] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0051] Figure 1 A flowchart of a multimodal skin disease analysis method based on deep learning provided by an embodiment of the present invention;

[0052] Figure 2 A flowchart for obtaining multimodal skin data including skin images and photoacoustic imaging information provided by an embodiment of the present invention;

[0053] Figure 3 Schematic diagram of the marking and two collection methods provided in an embodiment of the present invention;

[0054] Figure 4 A schematic diagram of the spatial relationship between the scanning position of the acoustic imaging device and the skin image provided by an embodiment of the present invention;

[0055] Figure 5 A schematic diagram of signal data collected at a single scanning point according to an embodiment of the present invention;

[0056] Figure 6 A schematic diagram of a skin analysis model provided by an embodiment of the present invention;

[0057] Figure 7 A schematic diagram of an attention convolution block provided by an embodiment of the present invention;

[0058] Figure 8 A schematic diagram of a state space enhancement module provided by an embodiment of the present invention;

[0059] Figure 9 A flowchart of obtaining a fused embedding by alternately connecting the patch embedding codes of the corresponding patch sizes of the different modal feature maps generated by each state space enhancement module provided in an embodiment of the present invention at the pixel level;

[0060] Figure 10 Schematic diagram of a feature pyramid network and a path aggregation network provided by an embodiment of the present invention;

[0061] Figure 11 A schematic diagram of the model training and testing accuracy provided by an embodiment of the present invention;

[0062] Figure 12 A schematic diagram of a classification confusion matrix provided by an embodiment of the present invention;

[0063] Figure 13 Schematic diagram of a multimodal skin disease analysis device based on deep learning provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0064] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0065] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0066] Example 1

[0067] like Figure 1 As shown, the present invention provides a multimodal skin disease analysis method based on deep learning, comprising:

[0068] Photoacoustic imaging equipment is used to scan and a dermatoscope is used to photograph the affected area of the skin disease to obtain photoacoustic imaging data and skin images of the affected area.

[0069] Align the photoacoustic imaging with the skin image, fill the space according to the skin image, and then combine it with the skin image to obtain multimodal skin data containing skin image and photoacoustic imaging information. Figure 2 As shown, the process includes:

[0070] Markers for aligning data of different modalities are set in advance on the affected area of the skin disease; the drawn marks can provide at least three points for calculating the affine matrix, and the marks can affect the photoacoustic imaging, leaving traces in the photoacoustic imaging data, and can also affect the dermoscopic imaging, leaving traces in the skin image pixels; Figure 3 As shown, Figure 3 The left side shows a photoacoustic imaging scan, and the right side shows a dermatoscope image. A colored fluorescent reagent is used to draw a frame around the skin lesion along a black line as a marker. This marker does not affect the detection of the skin lesion area. The frame-shaped marker is only an example; other shapes are possible.

[0071] When collecting data on skin disease affected areas, the scanning area of the photoacoustic imaging device and the imaging area of the dermatoscope cover the mark.

[0072] The space of the skin image is the area formed by its height and width; Figure 4 As shown in FIG, when the photoacoustic imaging device scans and collects data, the scanning process of the photoacoustic imaging device is actually a process of discrete sampling in the skin image space. The obtained photoacoustic imaging data can be associated with the pixel position of the skin image space. Figure 4 The discrete points in the image are the scanning points of photoacoustic imaging. Figure 4 This is just an example. The discrete point spacing is only for the purpose of illustrating the principle clearly and does not represent the actual scanning spacing. In fact, the scanning direction may not be parallel to the width or height of the skin image.

[0073] The signal data collected by the photoacoustic imaging device at a single scanning point is as follows Figure 5 As shown, it contains multiple photoacoustic channels. Figure 5 Given four photoacoustic channels, after scanning, the signal data from different channels in the resulting photoacoustic imaging data are stacked in the time dimension, arranged in the order of the scan points, and then transposed so that the stacked photoacoustic imaging data aligns with the spatial dimension of the skin image. At this point, the photoacoustic imaging data is not yet fully aligned with the pixels in the skin image space; only the time (depth) dimension of the photoacoustic imaging is parallel to the depth dimension of the skin image.

[0074] After spatial alignment, the pixels of the skin image and the corresponding photoacoustic imaging data are aligned based on the representation of the pre-set markers in the photoacoustic imaging and skin images. This fully aligns the photoacoustic imaging data and the skin image, establishing a connection between the photoacoustic imaging data and the skin image pixels. The specific process includes: for the skin image, extracting edges through edge detection, an example edge detection method is the Canny edge detection algorithm; extracting skin image marker edges from the extracted edges using marker edge features, an example marker edge feature is four corner points; for the photoacoustic imaging data, obtaining the layer where the marker representation is located from the stacked photoacoustic imaging data based on the sensitivity of the laser used to the fluorescent agent, extracting edges through edge detection at this layer, and extracting photoacoustic imaging data marker edges from the extracted edges using marker edge features; finding at least three pairs of paired point positions from the photoacoustic imaging data marker edges and the skin image marker edges; calculating the affine transformation matrix between the spatially aligned skin image and the photoacoustic imaging data using the paired point pairs, and applying the affine transformation matrix to the skin image or photoacoustic imaging data to achieve full alignment.

[0075] Because photoacoustic imaging scanning is essentially a discrete sampling of other modal data within the spatial domain of the skin image, the photoacoustic imaging data is sparser than the skin image in the spatial domain of the aligned data, requiring padding to unify the spatial dimensions of the different modal data. After fully aligning the photoacoustic imaging data with the skin image, padding is achieved through interpolation and filtering based on the skin image resolution.

[0076] Skin images and photoacoustic imaging data are normalized to form multimodal skin data.

[0077] A skin analysis model is constructed and pre-trained. The skin analysis model utilizes cross-modal, multimodal skin data to analyze and detect skin diseases. To train the skin analysis model, doctors or experts annotate the multimodal skin data. This annotation includes assigning skin disease type diagnostic labels based on the multimodal skin data and using a mask to demarcate the region within the multimodal skin data that contains the diagnostic label. The doctors or experts then analyze the data within the region to determine the skin disease type indicated by the corresponding multimodal skin data.

[0078] In the specific implementation process, Figure 6 As shown, the skin analysis model includes:

[0079] Two depth-wise separable convolutions are used to extract the modal features of skin image and photoacoustic imaging from multimodal skin data. The skin image feature maps extracted by the two depth-wise separable convolutions are consistent with the spatial and channel dimensions of the photoacoustic imaging feature maps, which can be expressed as:

[0080] ;

[0081] in, Depthwise convolution From skin imaging The skin image feature map extracted from Depthwise convolution From and photoacoustic imaging data The photoacoustic imaging feature map extracted from and The dimensions are exactly the same.

[0082] The backbone network following each depth-wise separable convolution consists of a 5-layer attention convolution block and a feature map encoding unit formed by downsampling. Each feature map encoding unit reduces the spatial dimension of the feature map of any input modality by half and doubles the channel dimension. , the five feature map encoding units of its backbone network are cascaded encoded:

[0083]

[0084] For photoacoustic imaging feature maps , the five feature map encoding units of its backbone network are cascaded encoded:

[0085]

[0086] in, 、 、 、 and They are skin image feature maps of different scales generated by the five feature map encoding units of the backbone network of the skin image feature map; 、 、 、 and They are the photoacoustic imaging feature maps of different scales generated by the five feature map encoding units of the backbone network of the photoacoustic imaging feature map; Skin image feature map The five feature map encoding units of the backbone network; Photoacoustic imaging characteristic map The five feature map encoding units of the backbone network.

[0087] In the specific implementation process, Figure 7As shown, the attention convolution block includes: a layer of depthwise separable convolution, the output of which is copied into two branches, one of which is passed through channel attention and spatial attention, and the other is sent to a convolution layer in parallel with the channel attention and spatial attention; each feature map encoding unit uses its channel attention and spatial attention to model the channel correlation and spatial correlation of the skin image feature map or photoacoustic imaging feature map of the corresponding scale it processes; the output of the channel attention and spatial attention is added and combined with the output of the convolution layer; and the space output of the attention convolution block is reduced by downsampling. The channel attention includes: adaptive average pooling for extracting channel attention weights, a convolution layer, a ReLU activation function, a convolution layer, and a Sigmoid function. The two convolution layers are used to compress and restore the channel dimension; the spatial attention includes a convolution layer and a Sigmoid function for extracting spatial attention weights.

[0088] The outputs of the last three feature map encoding units of the two backbone networks are connected to the state space enhancement modules, and the output feature maps are enhanced by the state space enhancement modules. That is, to enhance the feature maps of any modality at three scales, three state space enhancement modules are set.

[0089] A state space enhancement module based on skin image feature maps with dimensions H / 8, W / 8, C / 8 and photoacoustic imaging characteristic maps Generate skin image enhancement feature map of corresponding scale and photoacoustic imaging enhancement feature map , expressed as:

[0090] ;

[0091] A state space enhancement module is based on the skin image feature map with dimensions H / 16, W / 16, C / 16 and photoacoustic imaging characteristic maps Generate skin image enhancement feature map of corresponding scale and photoacoustic imaging enhancement feature map , expressed as:

[0092] ;

[0093] A state space enhancement module is based on the skin image feature map with dimensions H / 32, W / 32, C / 32 and photoacoustic imaging characteristic maps Generate skin image enhancement feature map of corresponding scale and photoacoustic imaging enhancement feature map , expressed as:

[0094] ;

[0095] in, Represents state space augmentation modules at three different scales.

[0096] In the specific implementation process, Figure 8 As shown, each state space enhancement module performs multiple sets of independent patch embedding codes of different patch sizes on the input skin image feature map and photoacoustic imaging feature map of any scale; multiple sets of independent patch embedding codes of different patch sizes are used to reduce the spatial dimension to disperse the computational load and capture multi-scale features.

[0097] For the three state-space enhancement modules, the patch embedding encoding of multiple patch sizes is expressed as:

[0098]

[0099] in, , K means there are K kinds of patch sizes in total, 、 、 、 、 、 They are 、 、 、 、 and The overall patch embedding code obtained with the kth patch size, and The dimension is , which can be split along the channel dimension into Patch embedding coding of k-th patch size; and The dimension is ; and The dimension is .

[0100] Each state space enhancement module alternately connects the patch embedding codes of the corresponding patch sizes of the different modal feature maps generated by it at the pixel level to obtain a fusion embedding. Figure 9As shown in the figure, the process includes: traversing the patch embedding coding pixels of the corresponding specifications of the feature maps of different modalities by row and column; directly splicing the patch embedding coding pixels of all even columns in the same row of the two patch embedding codes of the corresponding specifications in alternating order; reversing the pixel index order of all odd columns in the same row of the two patch embedding codes of the corresponding specifications, and then splicing the patch embedding coding pixels alternately in order. The splicing results of the patch embedding coding pixels of each row are stacked in row order to form a fused embedding. By dividing and conquering the odd and even columns and retaining the order within the row for the even columns and reversing the order within the row for the odd columns, the spatial continuity within the patch embedding coding of the photoacoustic imaging feature map and the skin imaging feature map is maintained; the pixels of the patch embedding coding of the feature maps of different modalities are alternately spliced to promote cross-modal information interaction. Combined with the patch embedding coding form of multi-scale patch sizes, multi-granularity features can be extracted in parallel.

[0101] Example: The rows of any two patch embedding codes contain column indices 0 through 4. The even-numbered columns (0, 2, 4) are concatenated alternately in their original order, while the odd-numbered columns (1, 3) are reversed to (3, 1) and then concatenated alternately. After concatenation, the pixel column order for that row is [0, 0, 2, 2, 4, 4, 3, 3, 1, 1]. The adjacent rows are [0, 0, 2, 2, 4, 4, 3, 3, 1, 1] and [0, 0, 2, 2, 4, 4, 3, 3, 1, 1]. This preserves the row structure while introducing context after reversal to ensure spatial continuity and enhance the model's sensitivity to edges and textures. Through multi-scale embedding and structured concatenation, computational efficiency and feature expression are balanced.

[0102] For skin image feature maps and photoacoustic imaging feature maps of any scale, K patch embedding codes formed by K patch sizes will be generated, and K fusion embeddings will be generated, which can be expressed as ,in Indicated by 、 K kinds of fusion embeddings generated by K kinds of patch embedding encoding; Indicated by 、 K kinds of fusion embeddings generated by K kinds of patch embedding encoding; Indicated by 、 The K kinds of patch embedding encodings generate K kinds of fusion embeddings.

[0103] Each state space enhancement module uses an independent linear layer to project each fusion embedding, and uses an independent state space model to model the relationship between the two modalities based on each fusion embedding, so as to enhance its own modality with the data of the other modality; the output of each state space model is mapped by a linear layer and combined with the corresponding fusion embedding to obtain a sub-enhanced fusion embedding, and then all sub-enhanced fusion embeddings are aggregated to obtain an enhanced fusion embedding, and the enhanced fusion embedding is reshaped in reverse order in the form of any modality feature map and divided into enhanced feature maps of the two modalities, namely the above-mentioned skin image enhancement feature map and photoacoustic imaging enhancement feature map , Skin image enhancement feature map and photoacoustic imaging enhancement feature map , Skin image enhancement feature map and photoacoustic imaging enhancement feature map . Wherein, the state space model is formed by stacking mamba blocks.

[0104] like Figure 10 As shown in the figure, in the feature pyramid network, the enhanced feature maps of the two modes of the last three feature map encoding units of the backbone network are respectively adjusted to a fixed value through the horizontally connected 1×1 convolutional layer;

[0105] Skin image enhancement feature map or photoacoustic imaging enhancement feature map The output of the horizontally connected 1×1 convolution layer is used as the skin image enhancement pyramid feature map or photoacoustic imaging enhanced pyramid feature map ;

[0106] Enhance the pyramid feature map of skin image or photoacoustic imaging enhanced pyramid feature map Upsampling to skin image enhancement feature map or photoacoustic imaging enhancement feature map Resolution, and skin image enhancement feature map or photoacoustic imaging enhancement feature map The output features of the horizontally connected 1×1 convolutional layer are added to generate the skin image enhancement pyramid feature map or photoacoustic imaging enhanced pyramid feature map ;

[0107] Enhance the pyramid feature map of skin image or photoacoustic imaging enhanced pyramid feature map Upsampling to skin image enhancement feature map or photoacoustic imaging enhancement feature map Resolution, and skin image enhancement feature map and photoacoustic imaging enhancement feature map The output features of the horizontally connected 1×1 convolutional layer are added to generate the skin image enhancement pyramid feature map or photoacoustic imaging enhanced pyramid feature map .

[0108] The path aggregation network introduces a bottom-up path based on the feature pyramid network, including:

[0109] Enhance the pyramid feature map of skin image or photoacoustic imaging enhanced pyramid feature map Downsampling to skin image enhancement pyramid feature map or photoacoustic imaging enhanced pyramid feature map resolution and enhance the pyramid feature map with skin image or photoacoustic imaging enhanced pyramid feature map Add together to generate skin image enhanced pyramid feature aggregation feature map or photoacoustic imaging enhanced pyramid feature aggregation feature map ;

[0110] Skin image enhancement pyramid feature aggregation feature map or photoacoustic imaging enhanced pyramid feature aggregation feature map Downsampling to skin image enhancement pyramid feature map or photoacoustic imaging enhanced pyramid feature map resolution and enhance the pyramid feature map with skin image or photoacoustic imaging enhanced pyramid feature map Add together to generate skin image enhanced pyramid feature aggregation feature map or photoacoustic imaging enhanced pyramid feature aggregation feature map .

[0111] The enhanced feature maps of the two modalities after enhancement by each state space enhancement module are processed by a two-way feature pyramid network and a path aggregation network respectively, fused through a fully connected layer and transmitted to the skin disease analysis head and the mask prediction head respectively; the skin disease analysis head is based on a fully connected implementation and is used to predict the type of skin disease; the mask prediction head is based on a fully connected fusion implementation and is used to predict the multimodal skin data area mask related to the skin disease type.

[0112] The loss function used to train the skin analysis model is the sum of the skin disease classification loss and the multimodal skin data region mask IOU loss related to the skin disease type. The skin analysis model parameters are adjusted with the goal of minimizing the loss function.

[0113] The training and testing accuracy of the skin analysis model of this application is as follows Figure 11As shown, under the optimal parameters, the classification accuracy is stable at more than 90%. The confusion matrix of the skin analysis model of this application for the classification of different types of skin diseases is as follows Figure 12 As shown in the figure, the degree of confusion in predicting skin disease types is small. The confusion matrix is obtained by statistically analyzing the classification of squamous cell carcinoma, benign pigmented keratosis, vascular lesions, seborrheic keratosis, nevus, melanoma, basal cell carcinoma, dermatofibroma, and actinic keratosis in the dataset.

[0114] Table 1 below shows the performance data for four existing skin disease recognition methods. Existing technologies recognize fewer skin disease types than this application. Existing technology 1, using the EfficientNetB4 model, achieved an accuracy of 89.97%; using InceptionV3, achieved an accuracy of 86.95%; using ResNet-50, achieved an accuracy of 87.61%; and using DenseNet169, achieved an accuracy of 88.46%. Existing technology 2, using Custom CNN, achieved an accuracy of 84.71%. Existing technology 3, using ResNet-50, achieved the highest accuracy of 89.49%. Existing technology 4, using an improved VGG16, achieved an accuracy of 90.67%. Existing technologies, when recognizing fewer types of skin diseases, still perform worse than this application.

[0115] Table 1: Schematic table of the effects of existing technologies

[0116] Example 2

[0117] See Figure 13 As shown, an embodiment of the present invention provides a multimodal skin disease analysis device based on deep learning, comprising: at least one processing unit, the processing unit being connected to a storage unit via a bus unit, the storage unit being a computer-readable storage medium that can be used to store software programs, computer executable programs, and modules, such as the software programs, computer executable programs, and modules corresponding to a multimodal skin disease analysis method based on deep learning in an embodiment of the present invention. The processing unit implements the above-mentioned multimodal skin disease analysis method based on deep learning by running the software programs, computer executable programs, and modules stored in the storage unit, comprising:

[0118] Use photoacoustic imaging equipment to scan and dermatoscope to photograph the affected area of the skin disease to obtain photoacoustic imaging data and skin images of the affected area;

[0119] Align the photoacoustic imaging with the skin image, fill the space according to the skin image, and then combine it with the skin image to obtain multimodal skin data containing skin image and photoacoustic imaging information;

[0120] A skin analysis model for assisting skin disease analysis based on multimodal skin data is constructed and pre-trained; the skin analysis model comprises: two depthwise separable convolutions that respectively extract skin imaging and photoacoustic imaging modal features from the multimodal skin data, each of which is followed by a backbone network, each of which comprises a five-layer attention convolution block and a feature map encoding unit formed by downsampling; a state-space enhancement module connected to the two backbone networks, wherein the feature maps of any modality output by the last three feature map encoding units of the two backbone networks are enhanced by the state-space enhancement module; the enhanced feature maps of the two modalities enhanced by each state-space enhancement module are respectively processed by a two-way feature pyramid network and a path aggregation network, fused through a fully connected layer, and transmitted to the skin disease analysis head and the mask prediction head, respectively;

[0121] The pre-trained skin analysis model analyzes the collected multimodal skin data to obtain skin disease analysis results.

[0122] Of course, the computer program stored in the storage unit of the multimodal skin disease analysis device based on deep learning provided by an embodiment of the present invention is not limited to the method operations described above, and can also execute related operations in the multimodal skin disease analysis method based on deep learning provided by any embodiment of the present invention.

[0123] Example 3

[0124] An embodiment of the present invention provides a computer-readable storage medium storing a computer program. When the computer program is executed, the multimodal skin disease analysis method based on deep learning is implemented, including:

[0125] Use photoacoustic imaging equipment to scan and dermatoscope to photograph the affected area of the skin disease to obtain photoacoustic imaging data and skin images of the affected area;

[0126] Align the photoacoustic imaging with the skin image, fill the space according to the skin image, and then combine it with the skin image to obtain multimodal skin data containing skin image and photoacoustic imaging information;

[0127] A skin analysis model for assisting skin disease analysis based on multimodal skin data is constructed and pre-trained; the skin analysis model comprises: two depthwise separable convolutions that respectively extract skin imaging and photoacoustic imaging modal features from the multimodal skin data, each of which is followed by a backbone network, each of which comprises a five-layer attention convolution block and a feature map encoding unit formed by downsampling; a state-space enhancement module connected to the two backbone networks, wherein the feature maps of any modality output by the last three feature map encoding units of the two backbone networks are enhanced by the state-space enhancement module; the enhanced feature maps of the two modalities enhanced by each state-space enhancement module are respectively processed by a two-way feature pyramid network and a path aggregation network, fused through a fully connected layer, and transmitted to the skin disease analysis head and the mask prediction head, respectively;

[0128] The pre-trained skin analysis model analyzes the collected multimodal skin data to obtain skin disease analysis results.

[0129] A computer-readable storage medium provided by an embodiment of the present invention stores a computer program that is not limited to the method operations described above, but can also execute related operations in a multimodal skin disease analysis method based on deep learning provided by any embodiment of the present invention.

[0130] In the embodiments provided by the present invention, it should be understood that the disclosed structures and methods can be implemented in other ways. For example, the structural embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, structure or unit, which can be electrical, mechanical or other forms.

[0131] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0132] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0133] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is intended to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A multimodal skin disease analysis method based on deep learning, characterized in that: include: Use photoacoustic imaging equipment to scan and dermatoscope to photograph the affected area of the skin disease to obtain photoacoustic imaging data and skin images of the affected area; Align the photoacoustic imaging with the skin image, fill the space according to the skin image, and then combine it with the skin image to obtain multimodal skin data containing skin image and photoacoustic imaging information; Construct and pre-train a skin analysis model based on multimodal skin data to assist in skin disease analysis; the skin analysis model includes two depthwise separable convolutions that extract skin image and photoacoustic imaging modality features from the multimodal skin data, respectively. Each depthwise separable convolution is followed by a backbone network, each backbone network comprising a five-layer attention convolution block and a downsampling feature map encoding unit. The state-space enhancement module is connected to the two backbone networks. The feature maps of any modality output by the last three feature map encoding units of the two backbone networks are enhanced by the state-space enhancement module. The enhanced feature maps of the two modalities after each state-space enhancement module are processed by a two-way feature pyramid network and a path aggregation network, respectively, and then fused through a fully connected layer and transmitted to the skin disease analysis head and mask prediction head respectively. The pre-trained skin analysis model analyzes the collected multimodal skin data to obtain skin disease analysis results.

2. The multimodal skin disease analysis method based on deep learning according to claim 1, characterized in that: The step of aligning the photoacoustic imaging with the skin image, filling the space according to the skin image, and then combining the photoacoustic imaging with the skin image to obtain multimodal skin data including skin image and photoacoustic imaging information includes: Markers are set in advance on the affected area of the skin disease for aligning the data of different modalities, and the markers can provide at least three points for calculating the affine matrix; When collecting data on skin disease affected areas, the scanning area of the photoacoustic imaging device and the imaging area of the dermatoscope cover the mark; stacking the scanned photoacoustic imaging data and arranging them in order, and then transposing the stacked photoacoustic imaging data so that they are aligned with the space of the skin image; After spatial alignment, the pixels of the skin image and the corresponding photoacoustic imaging data are aligned based on the representation of the pre-set markers in the photoacoustic image and the skin image; After the photoacoustic imaging data is fully aligned with the skin image, the photoacoustic imaging data is filled according to the resolution of the skin image; Skin images and photoacoustic imaging data are normalized to form multimodal skin data.

3. The multimodal skin disease analysis method based on deep learning according to claim 2, characterized in that: After the spatial alignment, aligning the pixels of the skin image with the corresponding photoacoustic imaging data based on the representation of the pre-set markers in the photoacoustic imaging and the skin image includes: For skin images, edges are extracted by edge detection; and skin image marker edges are extracted from the extracted edges using marker edge features; For photoacoustic imaging data, according to the sensitivity of the laser to the fluorescent agent, the layer corresponding to the corresponding laser imaging data is obtained from the stacked photoacoustic imaging data, the edge is extracted by edge detection in the layer, and the photoacoustic imaging data marking edge is extracted from the extracted edge using the marking edge feature; Find the positions of at least three pairs of paired points from the edge of the photoacoustic imaging data mark and the edge of the skin image mark; An affine transformation matrix between the spatially aligned skin image and the photoacoustic imaging data is calculated using the pairing points, and the affine transformation matrix is applied to the skin image or the photoacoustic imaging data to achieve complete alignment.

4. The multimodal skin disease analysis method based on deep learning according to claim 1, characterized in that: The attention convolution block includes: a layer of depth-wise separable convolution, the output of which is copied into two branches, one of which is passed through channel attention and spatial attention, and the other is sent to the convolution layer in parallel with the channel attention and spatial attention; The outputs of the channel attention and spatial attention are added to the output of the convolutional layer, where the channel attention includes: adaptive average pooling for extracting channel attention weights, convolutional layer, ReLU activation function, convolutional layer and Sigmoid function, and two convolutional layers are used to compress and restore the channel dimension; the spatial attention includes: convolutional layer and Sigmoid function for extracting spatial attention weights.

5. The multimodal skin disease analysis method based on deep learning according to claim 1, characterized in that: Each state space enhancement module performs independent patch embedding coding of multiple groups of different patch sizes on the input skin image feature map and photoacoustic imaging feature map of any scale; Each state space enhancement module alternately connects the patch embedding codes of the corresponding patch sizes of the different modality feature maps generated by it at the pixel level to obtain a fused embedding; Each state-space augmentation module projects each fused embedding using a separate linear layer; Each state-space enhancement module projects each fused embedding using an independent linear layer and uses an independent state-space model to model the association between the two modalities based on each fused embedding to enhance its own modality with the data of the other modality; The output of each state space model is mapped by a linear layer and then added to the corresponding fusion embedding to obtain a sub-enhanced fusion embedding; All sub-enhanced fusion embeddings are aggregated to obtain the enhanced fusion embedding, and the enhanced fusion embedding is reshaped in reverse order in the form of arbitrary modality feature maps and split into enhanced feature maps of two modalities.

6. The multimodal skin disease analysis method based on deep learning according to claim 5, characterized in that: Each state space enhancement module alternately connects the patch embedding codes of the corresponding patch sizes of the different modal feature maps generated by it at the pixel level to obtain a fusion embedding including: Traverse the corresponding patch embedding encoding pixels of different modality feature maps by rows and columns; Directly and alternately splice two patches of corresponding specifications in sequence to embed the encoded patches of all even columns in the same row; Reverse the pixel index order of all odd-numbered columns in the same row of the two patches with corresponding specifications, and then alternately splice the patch embedding coded pixels in order; The concatenation results of the patch embedding coded pixels in each row are stacked in row order to form a fused embedding.

7. The multimodal skin disease analysis method based on deep learning according to claim 1, characterized in that: In the feature pyramid network, the enhanced feature maps of the two modalities of the last three feature map encoding units of the backbone network are respectively adjusted to a fixed value through the horizontally connected 1×1 convolutional layer; Skin image enhancement feature map or photoacoustic imaging enhancement feature map The output of the horizontally connected 1×1 convolution layer is used as the skin image enhancement pyramid feature map or photoacoustic imaging enhanced pyramid feature map ; Enhance the pyramid feature map of skin image or photoacoustic imaging enhanced pyramid feature map Upsampling to skin image enhancement feature map or photoacoustic imaging enhancement feature map Resolution, and skin image enhancement feature map or photoacoustic imaging enhancement feature map The output features of the horizontally connected 1×1 convolutional layer are added to generate the skin image enhancement pyramid feature map or photoacoustic imaging enhanced pyramid feature map ; Enhance the pyramid feature map of skin image or photoacoustic imaging enhanced pyramid feature map Upsampling to skin image enhancement feature map or photoacoustic imaging enhancement feature map Resolution, and skin image enhancement feature map and photoacoustic imaging enhancement feature map The output features of the horizontally connected 1×1 convolutional layer are added to generate the skin image enhancement pyramid feature map or photoacoustic imaging enhanced pyramid feature map .

8. The multimodal skin disease analysis method based on deep learning according to claim 7, characterized in that: The path aggregation network introduces a bottom-up path based on the feature pyramid network, including: Enhance the pyramid feature map of skin image or photoacoustic imaging enhanced pyramid feature map Downsampling to skin image enhancement pyramid feature map or photoacoustic imaging enhanced pyramid feature map resolution and enhance the pyramid feature map with skin image or photoacoustic imaging enhanced pyramid feature map Add together to generate skin image enhanced pyramid feature aggregation feature map or photoacoustic imaging enhanced pyramid feature aggregation feature map ; Skin image enhancement pyramid feature aggregation feature map or photoacoustic imaging enhanced pyramid feature aggregation feature map Downsampling to skin image enhancement pyramid feature map or photoacoustic imaging enhanced pyramid feature map resolution and enhance the pyramid feature map with skin image or photoacoustic imaging enhanced pyramid feature map Add together to generate skin image enhanced pyramid feature aggregation feature map or photoacoustic imaging enhanced pyramid feature aggregation feature map .

9. The multimodal skin disease analysis method based on deep learning according to claim 1, characterized in that: The skin disease analysis head is implemented based on full connection and is used to predict the type of skin disease; the mask prediction head is implemented based on full connection fusion and is used to predict the multimodal skin data area mask related to the skin disease type.

10. The multimodal skin disease analysis method based on deep learning according to claim 1, characterized in that: The loss function used to train the skin analysis model is the sum of the skin disease classification loss and the multimodal skin data region mask IOU loss related to the skin disease type. The skin analysis model parameters are adjusted with the goal of minimizing the loss function.

Citation Information

Patent Citations

  • Method for segmenting focus in dermatoscope image

    CN115731226A

  • Skin cancer auxiliary diagnosis method based on dermatoscope image

    CN118366642A