Multi-modal skin disease analysis method based on deep learning

Through photoacoustic imaging equipment combined with dermatoscopes, a deep learning multimodal skin disease analysis method is constructed, which solves the problem of insufficient imaging depth and resolution in the prior art, and achieves efficient and accurate diagnosis of skin diseases, enhancing the interpretability of the model.

CN120260892AActive Publication Date: 2025-07-04AFFILIATED HOSPITAL OF WEIFANG MEDICAL UNIV

Patent Information

Application Number
CN202510732510.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

In the diagnosis of dermatology, the depth of optical dermatoscope imaging is limited, the spatial resolution and microvascular display ability of ultrasound imaging are poor, biochemical information cannot be obtained, and multimodal skin disease auxiliary diagnosis technology for multi-layer skin tissue is lacking.

Method used

Photoacoustic imaging equipment is used to scan and dermatoscope to obtain photoacoustic imaging data and skin images, and extract features through depth separation convolution and attention convolution blocks, build skin analysis models, and use state space enhancement modules and feature pyramid networks to fuse multimodal skin data to predict skin disease types and regional masks.

Benefits of technology

It realizes high-resolution, multivariate data-supported skin disease analysis, improves the accuracy and reliability of diagnosis, enhances the interpretability of the model, and avoids the situation where the disease analysis results do not match the data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260892A_ABST
    Figure CN120260892A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-modal skin disease analysis method based on deep learning, and relates to the technical field of auxiliary diagnosis. A photoacoustic imaging device is used for scanning, a dermatoscope is used for shooting a skin disease affected part, and photoacoustic imaging data and a skin image of the affected part are obtained; aligning the photoacoustic image with a skin image, filling according to a skin image space, and combining with the skin image to obtain multi-modal skin data; constructing and pre-training a skin analysis model based on the multi-modal skin data; the pre-trained skin analysis model analyzes the collected multi-modal skin data to obtain a skin disease analysis result. The multi-modal skin data is constructed, multivariate data support is provided, and it is ensured that disease analysis is accurate and reliable. The skin analysis model extracts different scales of skin image feature maps and photoacoustic imaging feature maps so as to cover correlation between different sizes of features and different specifications of skin diseases; the other modal data is used to reinforce the own modal, and the two modal features are better combined; and the interpretability is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of auxiliary diagnosis of skin multimodal data, and in particular to a multimodal skin disease analysis method based on deep learning. Background Art

[0002] Skin diseases are one of the most common diseases in humans, and the characteristics of skin diseases are the structural and functional changes of skin tissue components. Imaging technology plays an important role in the diagnosis and analysis of skin diseases. As a non-invasive observation and diagnosis means, imaging technology provides valuable information for clinicians. Traditional optical dermoscopy, such as confocal microscopy and optical coherence tomography, can only observe the morphology of the epidermis and superficial dermis due to the imaging depth limitation. Ultrasound imaging can visualize the entire skin structure through deep penetration, but its spatial resolution and the ability to display microvessels are poor, and it cannot obtain biochemical information related to metabolism.

[0003] Photoacoustic imaging directly measures the optical absorption characteristics of tissues, thus promoting the diagnosis of skin diseases. Photoacoustic imaging combines the advantages of optical and ultrasonic imaging. By irradiating the skin with short-pulse lasers and then receiving the optical contrast generated by the different absorption of light by endogenous pigments (hemoglobin, melanin, lipids, collagen, glucose, etc.), and the ultrasonic signals generated by the thermal expansion effect of the endogenous pigments that absorb light, photoacoustic imaging can provide clinicians with high-contrast, high-resolution skin morphology, function, and pathological information. It is used for the analysis of melanoma, café-au-lait spots, psoriasis, and skin vascular diseases, etc. Although progress has been made in photoacoustic imaging research, multimodal skin disease auxiliary diagnosis technologies regarding the optical and acoustic characteristics of multi-layer skin tissues are still lacking. Summary of the Invention

[0004] To solve the above technical problems or at least partially solve the above technical problems, the present invention provides a multimodal skin disease analysis method based on deep learning.

[0005] In a first aspect, the present invention provides a multimodal skin disease analysis method based on deep learning, including: Scanning the affected area of the skin disease with a photoacoustic imaging device and photographing the affected area of the skin disease with a dermoscope to obtain photoacoustic imaging data and skin images of the affected area; Aligning the photoacoustic imaging with the skin image, filling it according to the skin image space, and then combining it with the skin image to obtain multimodal skin data including skin image and photoacoustic imaging information; Construct and pre-train a skin analysis model based on multi-modal skin data; the skin analysis model includes: two depthwise separable convolutions that respectively extract the skin image and photoacoustic imaging modal features in the multi-modal skin data, and each depthwise separable convolution is followed by a backbone network. Each backbone network contains an attention convolution block with 5 layers and a feature map encoding unit formed by downsampling; a state space enhancement module connected to the two backbone networks, and the feature maps of any modality output by the last three feature map encoding units of the two backbone networks are enhanced by the state space enhancement module; after the enhanced feature maps of the two modalities enhanced by each state space enhancement module are respectively processed by two feature pyramid networks and a path aggregation network, they are fused through a fully connected layer and respectively transmitted to a skin disease analysis head and a mask prediction head; The pre-trained skin analysis model analyzes the collected multi-modal skin data to obtain skin disease analysis results.

[0006] Furthermore, the method of aligning the photoacoustic imaging with the skin image, filling it according to the skin image space, and then combining it with the skin image to obtain multi-modal skin data including skin image and photoacoustic imaging information includes: Pre-set marks for aligning different modality data at the skin disease affected area, and the marks can provide at least three points for calculating the affine matrix; When collecting data at the skin disease affected area, the scanning area of the photoacoustic imaging device and the imaging area of the dermoscope cover the marks; Stack and arrange the photoacoustic imaging data formed by scanning in order, and then transpose it so that the stacked and arranged photoacoustic imaging data is aligned with the space of the skin image; After spatial alignment, based on the representations of the pre-set marks in the photoacoustic imaging and the skin image, align the pixels of the skin image with the corresponding photoacoustic imaging data; After the photoacoustic imaging data is completely aligned with the skin image, fill the photoacoustic imaging data according to the skin image resolution; Normalize the skin image and the photoacoustic imaging data to form multi-modal skin data.

[0007] Furthermore, the step of aligning the pixels of the skin image with the corresponding photoacoustic imaging data based on the representations of the pre-set marks in the photoacoustic imaging and the skin image after spatial alignment includes: For the skin image, extract the edges through edge detection; use the mark edge features to extract the skin image mark edges from the extracted edges; For the photoacoustic imaging data, according to the sensitivity of the used laser to the fluorescent agent, obtain the layer corresponding to the corresponding laser imaging data from the stacked photoacoustic imaging data, extract the edges through edge detection at this layer, and use the mark edge features to extract the photoacoustic imaging data mark edges from the extracted edges; Find the positions of at least three pairs of paired points from the labeled edges of the photoacoustic imaging data and the labeled edges of the skin image; Calculate the affine transformation matrix between the spatially aligned skin image and the photoacoustic imaging data through the paired points, and apply the affine transformation matrix to the skin image or the photoacoustic imaging data to achieve complete alignment.

[0008] Furthermore, the attention convolution block includes: a layer of depthwise separable convolution, and the output of the depthwise separable convolution is copied into two branches. One branch passes through channel attention and spatial attention, and the other branch is fed into a convolution layer in parallel with the channel attention and the spatial attention; the outputs of the channel attention and the spatial attention are added and combined with the output of the convolution layer. Among them, the channel attention includes: adaptive average pooling for extracting channel attention weights, a convolution layer, a ReLU activation function, a convolution layer, and a Sigmoid function. The two convolution layers are used to compress and restore the channel dimension; the spatial attention includes: a convolution layer for extracting spatial attention weights and a Sigmoid function.

[0009] Furthermore, each of the state space enhancement modules performs multiple groups of independent Patch embedding encodings with different Patch sizes on the input skin image feature maps and photoacoustic imaging feature maps of any scale; Each state space enhancement module alternately connects the Patch embedding encodings of the corresponding Patch sizes of the different modality feature maps generated by it from the pixel level to obtain a fused embedding; Each state space enhancement module projects each fused embedding using an independent linear layer; Each state space enhancement module projects each fused embedding using an independent linear layer, and uses an independent state space model to model the association between the two modalities based on each fused embedding, so as to strengthen its own modality with data of the other modality; The output of each of the state space models is mapped through a linear layer and then added and combined with the corresponding fused embedding to obtain a sub-strengthened fused embedding; Aggregate all the sub-strengthened fused embeddings to obtain a strengthened fused embedding, and reshape the strengthened fused embedding in reverse order in the form of any modality feature map and split it into enhanced feature maps of the two modalities.

[0010] Furthermore, the step that each state space enhancement module alternately connects the Patch embedding encodings of the corresponding Patch sizes of the different modality feature maps generated by it from the pixel level to obtain a fused embedding includes: Traverse the Patch embedding encoding pixels of the corresponding specifications of the different modality feature maps row by row and column by column; Directly and alternately splice the Patch embedding encoding pixels of all even-numbered columns in the same row of the two Patch embedding encodings of the corresponding specifications in order; Reverse the pixel index order of all odd columns in the same row of the two Patch embedding encodings corresponding to the specifications, and then alternately splice the Patch embedding encoding pixels in order; Stack the splicing results of the Patch embedding encoding pixels in each row in row order to form a fused embedding.

[0011] Furthermore, in the feature pyramid network, the enhanced feature maps of the two modalities of the last three feature map encoding units of the backbone network are respectively adjusted to a fixed number of channels through 1×1 convolutional layers connected horizontally; Skin image enhanced feature map Or photoacoustic imaging enhanced feature map The output of the 1×1 convolutional layer connected horizontally of Or photoacoustic imaging enhanced pyramid feature map ; Upsample the skin image enhanced pyramid feature map Or photoacoustic imaging enhanced pyramid feature map To the resolution of the skin image enhanced feature map Or photoacoustic imaging enhanced feature map , and add it to the output features of the 1×1 convolutional layer connected horizontally of the skin image enhanced feature map Or photoacoustic imaging enhanced feature map To generate the skin image enhanced pyramid feature map Or photoacoustic imaging enhanced pyramid feature map ; Upsample the skin image enhanced pyramid feature map Or photoacoustic imaging enhanced pyramid feature map To the resolution of the skin image enhanced feature map Or photoacoustic imaging enhanced feature map , and add it to the output features of the 1×1 convolutional layer connected horizontally of the skin image enhanced feature map And photoacoustic imaging enhanced feature map To generate the skin image enhanced pyramid feature map Or photoacoustic imaging enhanced pyramid feature map .

[0012] Furthermore, the path aggregation network introduces a bottom-up path on the basis of the feature pyramid network, including: Downsample the skin image enhanced pyramid feature map Or photoacoustic imaging enhanced pyramid feature map To the resolution of the skin image enhanced pyramid feature map Or photoacoustic imaging enhanced pyramid feature map , and combine it with the skin image enhanced pyramid feature map Or the photoacoustic imaging enhanced pyramid feature map are added together to generate a skin image enhanced pyramid feature aggregation feature map Or the photoacoustic imaging enhanced pyramid feature aggregation feature map ; The skin image enhanced pyramid feature aggregation feature map Or the photoacoustic imaging enhanced pyramid feature aggregation feature map is downsampled to the resolution of the skin image enhanced pyramid feature map Or the photoacoustic imaging enhanced pyramid feature map and added to the skin image enhanced pyramid feature map Or the photoacoustic imaging enhanced pyramid feature map to generate a skin image enhanced pyramid feature aggregation feature map Or the photoacoustic imaging enhanced pyramid feature aggregation feature map .

[0013] Furthermore, the skin disease analysis head is implemented based on a fully connected layer and is used to predict the type of skin disease; the mask prediction head is implemented based on a fully connected fusion and is used to predict the multi-modal skin data region mask related to the type of skin disease.

[0014] Furthermore, the loss function used to train the skin analysis model is the sum of the skin disease classification loss and the IOU loss of the multi-modal skin data region mask related to the type of skin disease. With the goal of minimizing the loss function, the parameters of the skin analysis model are adjusted.

[0015] In a second aspect, the present invention provides a multi-modal skin disease analysis device based on deep learning, including: at least one processing unit, the processing unit is connected to a storage unit through a bus unit, the storage unit stores a computer program, and when the computer program is executed by the processing unit, the multi-modal skin disease analysis method based on deep learning as described above is implemented.

[0016] In a third aspect, the present invention provides a computer-readable storage medium, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the multi-modal skin disease analysis method based on deep learning as described above is implemented.

[0017] The above technical solutions provided by the embodiments of the present invention have the following advantages compared with the prior art: This application uses a photoacoustic imaging device to scan and a dermoscope to photograph the affected area of skin diseases, so as to obtain photoacoustic imaging data and skin images of the affected area; align the photoacoustic imaging with the skin image, fill it according to the skin image space, and then combine it with the skin image to obtain multimodal skin data containing skin image and photoacoustic imaging information. The multimodal skin data combines the skin image reflecting the skin surface layer and the photoacoustic imaging reflecting the information of different tissue layers of the skin. It provides multi-source data support for subsequent auxiliary analysis of skin diseases, ensuring the accuracy and reliability of the analysis.

[0018] This application constructs and pre-trains a skin analysis model for auxiliary skin disease analysis based on multimodal skin data; the skin analysis model includes: two depthwise separable convolutions that respectively extract the modal features of the skin image and photoacoustic imaging in the multimodal skin data, and each depthwise separable convolution is respectively followed by a backbone network, and each of the backbone networks includes an attention convolutional block with 5 layers and a feature map encoding unit formed by downsampling; a state space enhancement module connected to the two backbone networks, and the feature maps of any modality output by the last three feature map encoding units of the two backbone networks are enhanced by the state space enhancement module; after the enhanced feature maps of the two modalities enhanced by each state space enhancement module are respectively processed by two feature pyramid networks and a path aggregation network, they are fused through a fully connected layer and respectively transmitted to the skin disease analysis head and the mask prediction head; the pre-trained skin analysis model analyzes the collected multimodal skin data to obtain the skin disease analysis result. The skin analysis model respectively extracts the skin image feature maps of different scales and the photoacoustic imaging feature maps of different scales through the two backbone networks. The feature maps of different scales focus on the relationships between different granularity multimodal features and skin diseases, so as to cover the associations between different sized features and different specifications of skin diseases. The state space enhancement module uses the skin image to enhance the photoacoustic imaging and uses the photoacoustic imaging to enhance the skin image, so as to strengthen its own modality with the data of another modality and better combine the two modality features for auxiliary analysis.

[0019] In addition, this application combines mask prediction and auxiliary skin disease analysis, and uses mask prediction related to skin diseases to enhance the interpretability of the model, avoiding the situation where the skin disease analysis is correct, but the information of the multimodal skin data on which the disease analysis is based does not match the disease analysis result. Description of the Drawings

[0020] The drawings here are incorporated into the specification and form a part of this specification, showing the embodiments in line with the present invention, and are used together with the specification to explain the principles of the present invention.

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0022] Figure 1 Flowchart of a multi-modal skin disease analysis method based on deep learning provided by an embodiment of the present invention; Figure 2 Flowchart of obtaining multi-modal skin data including skin images and photoacoustic imaging information provided by an embodiment of the present invention; Figure 3 Schematic diagram of the marking and two acquisition methods provided by an embodiment of the present invention; Figure 4 Schematic diagram of the spatial relationship between the scanning position of the acoustic imaging device and the skin image provided by an embodiment of the present invention; Figure 5 Schematic diagram of the signal data collected at a single scanning point provided by an embodiment of the present invention; Figure 6 Schematic diagram of the skin analysis model provided by an embodiment of the present invention; Figure 7 Schematic diagram of the attention convolution block provided by an embodiment of the present invention; Figure 8 Schematic diagram of the state space enhancement module provided by an embodiment of the present invention; Figure 9 Flowchart of each state space enhancement module obtaining the fusion embedding by alternately connecting the corresponding Patch-sized Patch embeddings of the different modal feature maps generated at the pixel level provided by an embodiment of the present invention; Figure 10 Schematic diagram of the feature pyramid network and the path aggregation network provided by an embodiment of the present invention; Figure 11 Schematic diagram of the model training and test accuracy provided by an embodiment of the present invention; Figure 12 Schematic diagram of the classification confusion matrix provided by an embodiment of the present invention; Figure 13 Schematic diagram of a multi-modal skin disease analysis device based on deep learning provided by an embodiment of the present invention. Detailed implementation manners

[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0024] It should be noted that in this text, the terms "include", "comprise", or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the phrase "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the said element.

[0025] Embodiment 1 As Figure 1 shown, the present invention provides a multi-modal skin disease analysis method based on deep learning, including: Scanning the affected area of the skin disease with a photoacoustic imaging device and photographing it with a dermoscope to obtain photoacoustic imaging data and skin images of the affected area.

[0026] Align the photoacoustic imaging with the skin image, fill it according to the skin image space, and then combine it with the skin image to obtain multi-modal skin data containing skin image and photoacoustic imaging information. As Figure 2 shown, the process includes: Pre-set marks for aligning different modal data at the affected area of the skin disease; the drawn marks can provide at least three points for calculating the affine matrix, and the marks can both affect the photoacoustic imaging and leave traces in the photoacoustic imaging data, and can also affect the dermoscopic imaging and leave traces in the skin image pixels; as Figure 3 shown, Figure 3 On the left is the scanning by the photoacoustic imaging device, and on the right is the photographing by the dermoscope. Draw a frame around the lesion of the skin disease with a colored fluorescent reagent along the black line as a mark, and the mark does not affect the detection of the lesion area of the skin disease. The given frame-shaped mark is only an example, and it can be a mark of other shapes.

[0027] When collecting data of the affected area of the skin disease, the scanning area of the photoacoustic imaging device and the imaging area of the dermoscope cover the said mark.

[0028] The space of the skin image is the area formed by its height and width; as Figure 4As shown, when the photoacoustic imaging device scans and acquires data, the scanning process of the photoacoustic imaging device is actually a process of discrete sampling in the space of the skin image. The obtained photoacoustic imaging data can establish a connection with the pixel positions in the skin image space. Figure 4 The discrete points in Figure 4 are the scanning points of photoacoustic imaging. Figure 4 Just for example, the discrete point spacing is only for clearly showing the principle and does not represent the actual scanning spacing. And in fact, the scanning direction may not be parallel to the width or height of the skin image.

[0029] The signal data acquired by the photoacoustic imaging device at a single scanning point is as Figure 5 shown, which contains multiple photoacoustic channels. Figure 5 Four photoacoustic channels are given; after scanning, the signal data of different channels in the photoacoustic imaging data formed by scanning are stacked in the time dimension and arranged in the order of the scanning point positions, and then transposed so that the arranged and stacked photoacoustic imaging data is aligned with the space of the skin image. At this time, the photoacoustic imaging data is not completely aligned with the pixels in the skin image space. It is only that the time (depth) dimension of photoacoustic imaging is parallel to the depth dimension of the skin image.

[0030] After spatial alignment, based on the representations of the preset markers in photoacoustic imaging and the skin image, the pixels of the skin image and the corresponding photoacoustic imaging data are aligned. Thus, the photoacoustic imaging data and the skin image are completely aligned, and the connection between the photoacoustic imaging data and the skin image pixels is established. The specific process includes: for the skin image, the edges are extracted through edge detection. An example of an edge detection method is the canny edge detection algorithm; the skin image marker edges are extracted from the extracted edges by using the marker edge features. An example of the marker edge features is that the number of corner points is four; for the photoacoustic imaging data, according to the sensitivity of the used laser to the fluorescent agent, the layer where the marker representation is located is obtained from the stacked photoacoustic imaging data. The edges are extracted through edge detection in this layer, and the photoacoustic imaging data marker edges are extracted from the extracted edges by using the marker edge features; at least three pairs of paired point positions are found from the photoacoustic imaging data marker edges and the skin image marker edges; the affine transformation matrix between the skin image and the photoacoustic imaging data after spatial alignment is calculated through the paired point pairs, and the affine transformation matrix is applied to the skin image or the photoacoustic imaging data to achieve complete alignment.

[0031] Since the scanning of the photoacoustic imaging device is actually a discrete sampling of other modality data in the space of the skin image, in the spatial domain of the aligned data, the photoacoustic imaging data is sparse relative to the pixels of the skin image and needs to be filled to unify the spatial dimensions of different modality data. After completely aligning the photoacoustic imaging data with the skin image, the photoacoustic imaging data is filled by interpolation and filtering according to the skin image resolution.

[0032] Normalize the skin image and photoacoustic imaging data to form multimodal skin data.

[0033] Construct and pre-train a skin analysis model. The skin analysis model uses cross-modal multimodal skin data for skin disease analysis and detection. To train the skin analysis model, doctors and experts need to annotate the multimodal skin data. The annotation content includes: giving a diagnosis label for the skin disease type based on the multimodal skin data, and using a mask to divide the region that determines the skin disease type diagnosis label from the multimodal skin data. Doctors and experts analyze the skin disease type indicated by the corresponding multimodal skin data based on the data in the region that determines the skin disease type diagnosis label.

[0034] In the specific implementation process, as Figure 6 shown, the skin analysis model includes: Two depthwise separable convolutions that respectively extract the skin image and photoacoustic imaging modal features from the multimodal skin data. The spatial and channel dimensions of the skin image feature map and the photoacoustic imaging feature map extracted by the two depthwise separable convolutions are the same, expressed as: ; Among them, is the depth convolution extracting the skin image feature map from the skin image , is the depth convolution extracting the photoacoustic imaging feature map from the photoacoustic imaging data , and the dimensions of and are exactly the same.

[0035] Backbone networks respectively following each depthwise separable convolution. The backbone network contains 5-layer attention convolution blocks and feature map encoding units formed by downsampling. Each feature map encoding unit halves the spatial dimension and doubles the channel dimension of the input feature map of any modality. For the skin image feature map , the five feature map encoding units of its backbone network are cascaded for encoding:

[0036] For the photoacoustic imaging feature map , the five feature map encoding units of its backbone network are cascaded for encoding:

[0037] Among them, , , , and They are different-scale skin image feature maps generated by five feature map encoding units of the backbone network of the skin image feature map; , , , and They are different-scale photoacoustic imaging feature maps generated by five feature map encoding units of the backbone network of the photoacoustic imaging feature map; is the skin image feature map Five feature map encoding units of the backbone network; is the photoacoustic imaging feature map Five feature map encoding units of the backbone network.

[0038] In the specific implementation process, as Figure 7 shown, the attention convolution block includes: one layer of depthwise separable convolution, and the output of the depthwise separable convolution is copied into two branches. One branch passes through channel attention and spatial attention, and the other branch is fed into a convolution layer in parallel with the channel attention and spatial attention; each feature map encoding unit uses its channel attention and spatial attention to model the channel correlation and spatial correlation of the corresponding-scale skin image feature map or photoacoustic imaging feature map it processes; the outputs of the channel attention and spatial attention are added and combined with the output of the convolution layer; the space of the output of the downsampling and dimensionality reduction attention convolution block. The channel attention includes: adaptive average pooling for extracting channel attention weights, a convolution layer, a ReLU activation function, a convolution layer, and a Sigmoid function. The two convolution layers are used to compress and restore the channel dimension; the spatial attention includes a convolution layer for extracting spatial attention weights and a Sigmoid function.

[0039] The outputs of the last three feature map encoding units of the two backbone networks are respectively connected to the state space enhancement module, and the output feature maps are enhanced by the state space enhancement module. That is, three state space enhancement modules are set to enhance any modality feature maps of three scales.

[0040] One state space enhancement module is based on the skin image feature map with dimensions H / 8, W / 8, C / 8 and the photoacoustic imaging feature map to generate the corresponding-scale skin image enhanced feature map and the photoacoustic imaging enhanced feature map , expressed as: ; One state space enhancement module is based on the skin image feature map with dimensions H / 16, W / 16, C / 16 and the photoacoustic imaging feature map to generate the corresponding-scale skin image enhanced feature map and the photoacoustic imaging enhanced feature map , expressed as: ; A state space enhancement module generates skin image enhancement feature maps and photoacoustic imaging enhancement feature maps at corresponding scales based on a skin image feature map and a photoacoustic imaging feature map with dimensions of H / 32, W / 32, and C / 32, expressed as: and the photoacoustic imaging feature map to generate skin image enhancement feature maps and photoacoustic imaging enhancement feature maps at corresponding scales and the photoacoustic imaging enhancement feature map , expressed as: ; wherein, represent state space enhancement modules at three different scales.

[0041] In the specific implementation process, as shown in Figure 8 , each of the state space enhancement modules performs independent Patch embedding encoding with multiple different Patch sizes on the input skin image feature map and photoacoustic imaging feature map at any scale; and uses the independent Patch embedding encoding with multiple different Patch sizes to reduce the spatial dimension to disperse the computational load and capture multi-scale features.

[0042] For the three state space enhancement modules, the Patch embedding encoding with multiple Patch sizes is expressed as:

[0043] wherein, , K represents that there are a total of K types of Patch sizes, , , , , , are respectively , , , , and the overall Patch embedding encoding obtained with the k-th type of Patch size, and have dimensions of , and can be split along the channel dimension into Patch embedding encodings of the k-th type of Patch size; and have dimensions of ; and have dimensions of .

[0044] Each state space enhancement module obtains a fused embedding by alternately connecting the Patch embeddings of corresponding Patch sizes of the different modality feature maps generated at the pixel level. As Figure 9 shown, the process includes: traversing the Patch embeddings of pixels of corresponding specifications of different modality feature maps row by row and column by column; directly and alternately splicing the Patch embeddings of corresponding specifications of the two Patch embeddings of all even-numbered columns in the same row; reversing the pixel index order of all odd-numbered columns in the same row of the two Patch embeddings of corresponding specifications, and then alternately splicing the Patch embeddings of pixels. Stack the splicing results of the Patch embeddings of pixels in each row in row order to form a fused embedding. By dividing and conquering odd and even columns and keeping the in-row order for even columns and reversing the in-row order for odd columns, the spatial continuity within the Patch embeddings of the photoacoustic imaging feature map and the skin imaging feature map is maintained; alternately splicing the pixels of the Patch embeddings of different modality feature maps promotes cross-modal information interaction. Combining the Patch embedding encoding forms of multiple scale Patch sizes enables parallel extraction of multi-granularity features.

[0045] Example illustration: The rows of any two Patch embeddings contain column indices 0 to 4. The even-numbered columns (0, 2, 4) of the two are alternately spliced in the original order, and the odd-numbered columns (1, 3) are reversed to (3, 1) and then alternately spliced. After splicing, the column order of the pixels in this row is [0, 0, 2, 2, 4, 4, 3, 3, 1, 1], and for two adjacent rows [0, 0, 2, 2, 4, 4, 3, 3, 1, 1][0, 0, 2, 2, 4, 4, 3, 3, 1, 1], it can be seen that both the in-row structure is retained and the reversed context association is introduced to ensure spatial continuity, enhancing the model's sensitivity to edges or textures. Through multi-scale embedding and structured splicing, the computational efficiency and feature expression ability are balanced.

[0046] For the skin image feature maps and photoacoustic imaging feature maps of any scale, K kinds of Patch embeddings formed by K kinds of Patch sizes will be generated, and K kinds of fused embeddings will be produced, denoted as , where represents the K kinds of fused embeddings generated by the K kinds of Patch embeddings of and ; represents the K kinds of fused embeddings generated by the K kinds of Patch embeddings of and ; represents the K kinds of fused embeddings generated by the K kinds of Patch embeddings of and .

[0047] Each state space enhancement module projects each fused embedding using an independent linear layer, and models the association between the two modalities based on each fused embedding using an independent state space model to reinforce its own modality with the data of the other modality; the output of each said state space model is mapped through a linear layer and then added to the corresponding fused embedding to obtain a sub-reinforced fused embedding. Then, all sub-reinforced fused embeddings are aggregated to obtain a reinforced fused embedding, and the reinforced fused embedding is reshaped in reverse order in the form of any modality feature map and segmented into enhanced feature maps of the two modalities, namely the above-mentioned skin image enhanced feature map and photoacoustic imaging enhanced feature map , skin image enhanced feature map and photoacoustic imaging enhanced feature map , skin image enhanced feature map and photoacoustic imaging enhanced feature map . Among them, the state space model is formed by stacking mamba blocks

[0048] As Figure 10 shown, in the feature pyramid network, the enhanced feature maps of the two modalities of the last three feature map encoding units of the backbone network are respectively adjusted to a fixed number of channels through a 1×1 convolutional layer connected horizontally Skin image enhanced feature map or photoacoustic imaging enhanced feature map The output of the 1×1 convolutional layer connected horizontally is used as the skin image enhanced pyramid feature map or photoacoustic imaging enhanced pyramid feature map ; The skin image enhanced pyramid feature map or photoacoustic imaging enhanced pyramid feature map is upsampled to the resolution of the skin image enhanced feature map or photoacoustic imaging enhanced feature map , and added to the output features of the 1×1 convolutional layer connected horizontally of the skin image enhanced feature map or photoacoustic imaging enhanced feature map to generate the skin image enhanced pyramid feature map or photoacoustic imaging enhanced pyramid feature map ; The skin image enhanced pyramid feature map or photoacoustic imaging enhanced pyramid feature map is upsampled to the resolution of the skin image enhanced feature map or photoacoustic imaging enhanced feature map , and added to the skin image enhanced feature map and photoacoustic imaging enhanced feature map The features output by the horizontally connected 1×1 convolutional layers are added together to generate a skin image enhanced pyramid feature map or a photoacoustic imaging enhanced pyramid feature map .

[0049] The Path Aggregation Network introduces a bottom-up path on the basis of the Feature Pyramid Network, including: Downsample the skin image enhanced pyramid feature map or the photoacoustic imaging enhanced pyramid feature map to the resolution of the skin image enhanced pyramid feature map or the photoacoustic imaging enhanced pyramid feature map , and add it to the skin image enhanced pyramid feature map or the photoacoustic imaging enhanced pyramid feature map to generate a skin image enhanced pyramid feature aggregation feature map or a photoacoustic imaging enhanced pyramid feature aggregation feature map ; Downsample the skin image enhanced pyramid feature aggregation feature map or the photoacoustic imaging enhanced pyramid feature aggregation feature map to the resolution of the skin image enhanced pyramid feature map or the photoacoustic imaging enhanced pyramid feature map , and add it to the skin image enhanced pyramid feature map or the photoacoustic imaging enhanced pyramid feature map to generate a skin image enhanced pyramid feature aggregation feature map or a photoacoustic imaging enhanced pyramid feature aggregation feature map .

[0050] After the enhanced feature maps of the two modalities enhanced by each state space enhancement module are processed by the two-path Feature Pyramid Network and the Path Aggregation Network respectively, they are fused through a fully connected layer and transmitted to the skin disease analysis head and the mask prediction head respectively; the skin disease analysis head is based on a fully connected layer and is used to predict the skin disease type; the mask prediction head is based on a fully connected fusion and is used to predict the multi-modal skin data region mask related to the skin disease type.

[0051] The loss function used to train the skin analysis model is the sum of the skin disease classification loss and the IOU loss of the multi-modal skin data region mask related to the skin disease type. With the goal of minimizing the loss function, the parameters of the skin analysis model are adjusted.

[0052] The training and test accuracy rates of the skin analysis model of the present application are as Figure 11As shown, when the parameters are optimal, the classification accuracy rate is stable above 90%. The confusion matrix of the skin analysis model of the present application for classifying different types of skin diseases is as follows Figure 12 As shown, the degree of confusion in predicting the skin disease type is small. The confusion matrix is obtained by statistically analyzing the classification of squamous cell carcinoma, pigmented benign keratosis, vascular lesions, seborrheic keratosis, nevus, melanoma, basal cell carcinoma, dermatofibroma, and actinic keratosis in the dataset.

[0053] The following Table 1 gives the effect data of four existing skin disease recognition methods, and the types of skin diseases recognized by the prior art are fewer than those of the present application; for the prior art 1, the accuracy rate of using the EfficientNetB4 model is 89.97%; the accuracy rate of using InceptionV3 is 86.95%; the accuracy rate of using ResNet-50 is 87.61%; the accuracy rate of using DenseNet169 is 88.46%. The prior art two uses Custom CNN with an accuracy rate of 84.71. The prior art 3 applies ResNet-50, and the highest accuracy rate is 89.49%. The prior art 4 uses the improved VGG16, and the accuracy rate reaches 90.67%. Even when the prior art has fewer types of recognition and classification, the effect is still worse than that of the present application.

[0054] Table 1: Schematic table of the effects of the prior art

[0055] Example 2 Refer to Figure 13 As shown, the embodiment of the present invention provides a multi-modal skin disease analysis device based on deep learning, including: at least one processing unit, the processing unit is connected to a storage unit through a bus unit, and the storage unit is used as a computer-readable storage medium, which can be used to store software programs, computer-executable programs, and modules, such as the software programs, computer-executable programs, and modules corresponding to a multi-modal skin disease analysis method based on deep learning in the embodiment of the present invention. The processing unit realizes the above-mentioned multi-modal skin disease analysis method based on deep learning by running the software programs, computer-executable programs, and modules stored in the storage unit, including: Scanning with a photoacoustic imaging device and photographing the affected area of the skin disease with a dermoscope to obtain photoacoustic imaging data and skin images of the affected area; Align the photoacoustic imaging with the skin image, fill it according to the skin image space, and then combine it with the skin image to obtain multi-modal skin data containing skin image and photoacoustic imaging information; Construct and pre-train a skin analysis model for assisting skin disease analysis based on multi-modal skin data; the skin analysis model includes: two depthwise separable convolutions that respectively extract the skin image and photoacoustic imaging modal features in the multi-modal skin data, and each depthwise separable convolution is followed by a backbone network, and each of the backbone networks includes a 5-layer attention convolution block and a feature map encoding unit formed by downsampling; a state space enhancement module connected to the two backbone networks, and the feature maps of any modality output by the last three feature map encoding units of the two backbone networks are enhanced by the state space enhancement module; after the enhanced feature maps of the two modalities enhanced by each state space enhancement module are respectively processed by two feature pyramid networks and a path aggregation network, they are fused through a fully connected layer and respectively transmitted to a skin disease analysis head and a mask prediction head; The pre-trained skin analysis model analyzes the collected multi-modal skin data to obtain skin disease analysis results.

[0056] Of course, the computer program stored in the storage unit of a multi-modal skin disease analysis device based on deep learning provided by the embodiments of the present invention is not limited to the method operations described above, and can also execute the related operations in a multi-modal skin disease analysis method based on deep learning provided by any embodiment of the present invention.

[0057] Embodiment 3 The embodiments of the present invention provide a computer-readable storage medium, and the computer-readable storage medium stores a computer program, and when the computer program is executed, the multi-modal skin disease analysis method based on deep learning is implemented, including: Using a photoacoustic imaging device to scan and a dermoscope to photograph the affected area of the skin disease to obtain photoacoustic imaging data and skin images of the affected area; Align the photoacoustic imaging with the skin image, fill it according to the skin image space, and then combine it with the skin image to obtain multi-modal skin data containing skin image and photoacoustic imaging information; Construct and pre-train a skin analysis model for assisting skin disease analysis based on multi-modal skin data; the skin analysis model includes: two depthwise separable convolutions that respectively extract the skin image and photoacoustic imaging modal features in the multi-modal skin data, and each depthwise separable convolution is followed by a backbone network, and each of the backbone networks includes a 5-layer attention convolution block and a feature map encoding unit formed by downsampling; a state space enhancement module connected to the two backbone networks, and the feature maps of any modality output by the last three feature map encoding units of the two backbone networks are enhanced by the state space enhancement module; after the enhanced feature maps of the two modalities enhanced by each state space enhancement module are respectively processed by two feature pyramid networks and a path aggregation network, they are fused through a fully connected layer and respectively transmitted to a skin disease analysis head and a mask prediction head; The pre-trained skin analysis model analyzes the collected multimodal skin data to obtain the skin disease analysis results.

[0058] The computer-readable storage medium provided by the embodiments of the present invention stores computer programs that are not limited to the method operations described above, and can also execute the related operations in a multimodal skin disease analysis method based on deep learning provided by any embodiment of the present invention.

[0059] In the embodiments provided by the present invention, it should be understood that the disclosed structures and methods can be implemented in other ways. For example, the structural embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of structures or units can be in electrical, mechanical or other forms.

[0060] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0061] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0062] The above are only the specific implementation manners of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A multi-modal skin disease analysis method based on deep learning, characterized in that, Including: Scanning with a photoacoustic imaging device and photographing the affected area of skin diseases with a dermoscope to obtain photoacoustic imaging data and skin images of the affected area; Aligning the photoacoustic imaging with the skin image, filling it according to the skin image space, and then combining it with the skin image to obtain multimodal skin data containing skin image and photoacoustic imaging information; Constructing and pre-training a skin analysis model for assisting skin disease analysis based on multimodal skin data; the skin analysis model includes: two depthwise separable convolutions that respectively extract the modal features of the skin image and photoacoustic imaging in the multimodal skin data, and each depthwise separable convolution is followed by a backbone network, and each of the backbone networks contains an attention convolution block with 5 layers and a feature map encoding unit formed by downsampling; A state space enhancement module connected to the two backbone networks, and the feature maps of any modality output by the last three feature map encoding units of the two backbone networks are enhanced by the state space enhancement module; after the enhanced feature maps of the two modalities enhanced by each state space enhancement module are respectively processed by two feature pyramid networks and a path aggregation network, they are fused through a fully connected layer and respectively transmitted to a skin disease analysis head and a mask prediction head; Analyzing the collected multimodal skin data with the pre-trained skin analysis model to obtain skin disease analysis results.

2. The multi-modal skin disease analysis method based on deep learning according to claim 1, wherein The step of aligning the photoacoustic imaging with the skin image, filling it according to the skin image space, and then combining it with the skin image to obtain multimodal skin data containing skin image and photoacoustic imaging information includes: Pre-setting marks for aligning different modality data at the affected area of skin diseases, and the marks can provide at least three points for calculating the affine matrix; When collecting data of the affected area of skin diseases, the scanning area of the photoacoustic imaging device and the imaging area of the dermoscope cover the marks; Stacking and arranging the photoacoustic imaging data formed by scanning in order, and then transposing it so that the stacked and arranged photoacoustic imaging data is aligned with the space of the skin image; After spatial alignment, based on the representations of the pre-set marks in the photoacoustic imaging and the skin image, aligning the pixels of the skin image and the corresponding photoacoustic imaging data; After completely aligning the photoacoustic imaging data with the skin image, filling the photoacoustic imaging data according to the skin image resolution; Normalizing the skin image and the photoacoustic imaging data to form multimodal skin data.

3. The multi-modal skin disease analysis method based on deep learning according to claim 2, wherein The step of, after spatial alignment, based on the representations of the pre-set marks in the photoacoustic imaging and the skin image, aligning the pixels of the skin image and the corresponding photoacoustic imaging data includes: For the skin image, extracting edges through edge detection; extracting the skin image mark edges from the extracted edges by using the mark edge features; For the photoacoustic imaging data, according to the sensitivity of the used laser to the fluorescent agent, obtaining the layer corresponding to the corresponding laser imaging data from the stacked photoacoustic imaging data, extracting edges through edge detection in this layer, and extracting the photoacoustic imaging data mark edges from the extracted edges by using the mark edge features; Finding the positions of at least three pairs of paired points from the photoacoustic imaging data mark edges and the skin image mark edges; Calculating the affine transformation matrix between the skin image and the photoacoustic imaging data after spatial alignment through the paired points, and applying the affine transformation matrix to the skin image or the photoacoustic imaging data to achieve complete alignment.

4. The multi-modal skin disease analysis method based on deep learning according to claim 1, characterized in that, The attention convolution block includes: a layer of depthwise separable convolution, the output of the depthwise separable convolution is copied into two branches, one branch passes through channel attention and spatial attention, and the other branch is fed into a convolution layer in parallel with the channel attention and spatial attention; The outputs of the channel attention and spatial attention are added and combined with the output of the convolution layer. Among them, the channel attention includes: adaptive average pooling for extracting channel attention weights, a convolution layer, a ReLU activation function, a convolution layer, and a Sigmoid function. The two convolution layers are used to compress and restore the channel dimension; the spatial attention includes: a convolution layer for extracting spatial attention weights and a Sigmoid function.

5. The multi-modal skin disease analysis method based on deep learning according to claim 1, wherein Each of the state space enhancement modules performs independent Patch embedding encoding with multiple different Patch sizes on the input skin image feature map and photoacoustic imaging feature map of any scale; Each state space enhancement module alternately connects the Patch embedding encodings of the corresponding Patch sizes of the different modality feature maps generated by it from the pixel level to obtain a fused embedding; Each state space enhancement module projects each fused embedding using an independent linear layer; Each state space enhancement module projects each fused embedding using an independent linear layer, and uses an independent state space model to model the association between the two modalities based on each fused embedding, so as to reinforce its own modality with data of the other modality; The output of each state space model is mapped through a linear layer and then added and combined with the corresponding fused embedding to obtain a sub-reinforced fused embedding; All sub-reinforced fused embeddings are aggregated to obtain a reinforced fused embedding, and the reinforced fused embedding is reshaped in reverse order in the form of any modality feature map and segmented into enhanced feature maps of the two modalities.

6. The multi-modal skin disease analysis method based on deep learning according to claim 5, characterized in that, Each of the state space enhancement modules alternately connects the Patch embedding encodings of the corresponding Patch sizes of the different modality feature maps generated by it from the pixel level to obtain a fused embedding, including: Traverse the Patch embedding encoded pixels of the corresponding specifications of the different modality feature maps row by row and column by column; Directly alternately splice the Patch embedding encoded pixels of all even columns in the same row of the two Patch embeddings of the corresponding specifications in order; Reverse the pixel index order of all odd columns in the same row of the two Patch embeddings of the corresponding specifications, and then alternately splice the Patch embedding encoded pixels in order; Stack the splicing results of the Patch embedding encoded pixels in each row in row order to form a fused embedding.

7. The multi-modal skin disease analysis method based on deep learning according to claim 1, wherein In the feature pyramid network, the enhanced feature maps of the two modalities of the last three feature map encoding units of the backbone network are respectively adjusted to a fixed value of the number of channels through a 1×1 convolution layer connected horizontally; Skin image enhancement feature map Or the output of the 1×1 convolutional layer connected horizontally of the photoacoustic imaging enhancement feature map Is used as the skin image enhancement pyramid feature map Or the photoacoustic imaging enhancement pyramid feature map ; Upsample the skin image enhanced pyramid feature map or the photoacoustic imaging enhanced pyramid feature map to the resolution of the skin image enhanced feature map or the photoacoustic imaging enhanced feature map , and add the output features of the 1×1 convolutional layer connected horizontally with the skin image enhanced feature map or the photoacoustic imaging enhanced feature map to generate the skin image enhanced pyramid feature map or the photoacoustic imaging enhanced pyramid feature map ; Upsample the skin image enhanced pyramid feature map or the photoacoustic imaging enhanced pyramid feature map to the resolution of the skin image enhanced feature map or the photoacoustic imaging enhanced feature map , and add the output features of the 1×1 convolutional layer connected horizontally to the skin image enhanced feature map and the photoacoustic imaging enhanced feature map to generate the skin image enhanced pyramid feature map or the photoacoustic imaging enhanced pyramid feature map .

8. The multi-modal skin disease analysis method based on deep learning according to claim 7, wherein The path aggregation network introduces a bottom-up path on the basis of the feature pyramid network, including: Downsample the skin image enhanced pyramid feature map or the photoacoustic imaging enhanced pyramid feature map to the resolution of the skin image enhanced pyramid feature map or the photoacoustic imaging enhanced pyramid feature map , and add it to the skin image enhanced pyramid feature map or the photoacoustic imaging enhanced pyramid feature map to generate the skin image enhanced pyramid feature aggregation feature map or the photoacoustic imaging enhanced pyramid feature aggregation feature map ; Aggregate the feature maps of the enhanced pyramid of skin images Or aggregate the feature maps of the enhanced pyramid of photoacoustic imaging Downsample to the feature maps of the enhanced pyramid of skin images Or the feature maps of the enhanced pyramid of photoacoustic imaging At the resolution of, and add it to the feature maps of the enhanced pyramid of skin images Or the feature maps of the enhanced pyramid of photoacoustic imaging To generate the aggregated feature maps of the enhanced pyramid of skin images Or the aggregated feature maps of the enhanced pyramid of photoacoustic imaging .

9. The multi-modal skin disease analysis method based on deep learning according to claim 1, wherein The skin disease analysis head is based on a fully connected layer and is used to predict the skin disease type; the mask prediction head is based on a fully connected fusion and is used to predict the multi-modal skin data region mask related to the skin disease type.

10. The multi-modal skin disease analysis method based on deep learning according to claim 1, wherein The loss function used to train the skin analysis model is the sum of the skin disease classification loss and the IOU loss of the multi-modal skin data region mask related to the skin disease type. With the goal of minimizing the loss function, the parameters of the skin analysis model are adjusted.

Citation Information

Patent Citations

  • Method for segmenting focus in dermatoscope image

    CN115731226A

  • Skin lesion image segmentation method, system and device and storage medium

    CN117893545A

  • Skin cancer auxiliary diagnosis method based on dermatoscope image

    CN118366642A

  • Heart ultrasound image segmentation method based on state space model and feature interactive perception

    CN119579614A

  • Multi-modal method for classifying thyroid nodule based on ultrasound and infrared thermal images

    US20240282090A1

Cited By

  • Disease data analysis intelligent chip and system based on multi-modal molecular marker

    CN121096614A

  • Multimodal molecular marker-based disease data analysis intelligent chip and system

    CN121096614B