Skin disease auxiliary diagnosis method and system based on multi-modal optical image fusion

By performing semantic segmentation and RCM image feature association on dermoscopic images, the problem of underutilization of multimodal data in skin disease diagnosis is solved, achieving automated and precise skin disease classification and improving diagnostic efficiency and accuracy.

CN121687461BActive Publication Date: 2026-05-01WEST CHINA HOSPITAL SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WEST CHINA HOSPITAL SICHUAN UNIV
Filing Date
2026-02-10
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing technologies, the image analysis of dermatoscopes and reflective confocal microscopes lacks an effective automated correlation mechanism, resulting in the incomplete exploitation of the comprehensive value of multimodal data. The identification process relies on manual integration, which is inefficient and inaccurate.

Method used

By semantically segmenting dermoscopy images to identify target feature regions, relevant feature regions are selected from the RCM image set based on semantic category and location information. Feature vectors from the two modalities are extracted and fused, and a trained classification model is used for disease classification.

Benefits of technology

It enables automated and precise semantic association between dermoscopy and RCM images, improves the efficiency and accuracy of multimodal data integration, generates more comprehensive disease classification results, and reduces reliance on professional experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121687461B_ABST
    Figure CN121687461B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of medical image processing, and particularly discloses a skin disease auxiliary diagnosis method and system based on multi-modal optical image fusion, which comprises the following steps: performing semantic segmentation on a dermoscope image of an acquired target skin area to identify at least one target feature area; based on semantic categories and position information of the target feature area, screening at least one associated feature area from an RCM image set of the acquired target skin area; extracting a first feature vector in the target feature area, and extracting a second feature vector in the associated feature area; fusing the first feature vector and the second feature vector, performing disease classification on the fused feature vector based on a trained classification model, and obtaining a skin disease classification result of the target skin area. According to the application, corresponding associated areas can be screened from the RCM image set according to the result of the dermoscope semantic segmentation, and the efficiency and reliability of multi-modal data integration are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

A method and system for auxiliary diagnosis of skin diseases based on multimodal optical image fusion Technical Field

[0001] This invention relates to the field of medical image processing technology, and in particular to a method and system for auxiliary diagnosis of skin diseases based on multimodal optical image fusion. Background Technology

[0002] Dermoscopy and Reflective Confocal Microscopy (RCM) are two important non-invasive imaging tools in the clinical diagnosis of skin diseases and have been widely used. Dermoscopy can provide information on the macroscopic surface structure and vascular patterns of skin lesions, while RCM can achieve "optical biopsy" of living skin at the cellular resolution level, revealing microscopic pathological changes in the epidermis and superficial dermis. Currently, in practice, evaluators usually observe images of these two modalities independently and make comprehensive judgments based on their personal experience. This mode is highly dependent on the evaluator's professional level, and the identification process is time-consuming, subjective, and difficult to standardize. Existing technologies also include some methods that use computer vision technology to assist in the analysis of single-modality images (such as analyzing dermoscopy images or RCM images separately), but these methods have significant limitations.

[0003] Specifically, existing methods cannot intelligently locate, filter, or associate corresponding images containing key microscopic evidence from the RCM image set based on specific pathological feature regions (such as annular vessels and thick scales) automatically identified in dermoscopy images. This results in the comprehensive value of multimodal data not being fully explored, and the final identification and classification process still heavily relies on manual integration, which is inefficient and makes it difficult to guarantee the accuracy and consistency of association, thus restricting further improvement of the overall performance of the auxiliary diagnostic system.

[0004] Therefore, there is an urgent need for an auxiliary recognition scheme that can accurately associate semantic information between dermatoscopy and RCM images to improve the objectivity and accuracy of dermatological image analysis. Summary of the Invention

[0005] In order to overcome the problems of low efficiency and low accuracy in the existing methods of identifying skin disease types through skin disease image fusion, this invention provides a method and system for auxiliary diagnosis of skin diseases based on multimodal optical image fusion.

[0006] In a first aspect, the present invention provides a method for auxiliary diagnosis of skin diseases based on multimodal optical image fusion, comprising:

[0007] Perform semantic segmentation on the acquired dermoscopic image of the target skin region to identify at least one target feature region;

[0008] Based on the semantic category and location information of the target feature region, at least one associated feature region is selected from the acquired RCM image set of the target skin region;

[0009] Extract a first feature vector from the target feature region, and extract a second feature vector from the associated feature region;

[0010] The first feature vector and the second feature vector are fused together, and the fused feature vector is used to classify diseases based on the trained classification model to obtain the skin disease classification result of the target skin region.

[0011] According to a specific implementation method, in the above-mentioned auxiliary diagnostic method for skin diseases, at least one associated feature region is selected from the acquired RCM image set of the target skin region, specifically including:

[0012] The spatial range corresponding to the RCM image set is determined based on the location information;

[0013] Within the spatial range, based on the preset correspondence between semantic categories and RCM image features, the correlation between each image or image region in the RCM image set and the target feature region is evaluated;

[0014] Based on the correlation degree, relevant feature regions are selected from the spatial range.

[0015] According to one specific implementation, in the above-mentioned auxiliary diagnostic method for skin diseases, the semantic categories include:

[0016] Color and pigment characteristics include erythematous areas, hyperpigmented areas, and hypopigmented areas;

[0017] Surface structure and scaling characteristics include thick scaling areas, fine scaling areas, collar-shaped scaling areas, and hair follicle plug areas;

[0018] Vascular morphological characteristics include punctate vascular regions, annular vascular regions, linear vascular regions, and dendritic vascular regions;

[0019] Special pattern feature classes include Wickham pattern areas, fissure areas, and widened skin groove areas.

[0020] According to a specific implementation, in the above-mentioned auxiliary diagnostic method for skin diseases, the correspondence between the semantic category and RCM image features includes at least one of the following:

[0021] The annular vascular region corresponds to the dilation and tortuosity of papillary dermal vessels in RCM images;

[0022] The thick scaly areas correspond to epidermal fusion parakeratosis and Munro microabscess features in RCM images;

[0023] Fine, fragmented scales or collar-shaped scales correspond to focal incomplete keratinization of the epidermis in RCM images;

[0024] The Wickham pattern area corresponds to the wedge-shaped thickening and hyperkeratosis of the epidermal granular layer in the RCM image;

[0025] The follicular plug area corresponds to the characteristics of thickened stratum corneum and enlarged follicular infundibulum in the RCM image.

[0026] According to a specific implementation, in the above-mentioned auxiliary diagnostic method for skin diseases, extracting the first feature vector in the target feature region includes:

[0027] Extract at least one of the color statistical features, texture features, geometric features, and morphological features from the target feature region and convert them into vectors to obtain the first feature vector;

[0028] Extracting the second feature vector from the associated feature region includes:

[0029] Extract at least one of the following features from the associated feature region: epidermal cell structure features, dermal papillary layer features, inflammatory cell infiltration pattern features, and collagen fiber features, and convert them into vectors to obtain the second feature vector.

[0030] According to a specific implementation, in the above-mentioned auxiliary diagnostic method for skin diseases, the epidermal cell structure characteristics include at least one of hyperkeratosis, spongiosis, and disordered arrangement of keratinocytes; the dermal papillary layer characteristics include at least one of dermal papillary edema and dermal papillary vascular dilation; and the inflammatory cell infiltration pattern characteristics include at least one of lichenified band-like infiltration and perivascular infiltration.

[0031] According to a specific implementation method, the above-mentioned auxiliary diagnostic method for skin diseases involves fusing the first feature vector and the second feature vector, and classifying the fused feature vector for disease based on a trained classification model, specifically including:

[0032] The first feature vector and the second feature vector are concatenated to obtain the fused feature vector;

[0033] The fused feature vector is input into the trained classification model to obtain the corresponding skin disease classification result; the skin disease classification result includes the probability that the target skin region belongs to at least one skin disease.

[0034] According to one specific implementation method, the training process of the classification model in the above-mentioned auxiliary diagnostic method for skin diseases includes:

[0035] The training sample set is based on the acquired dermoscopy image set, the corresponding RCM image set, and the pathologically confirmed diagnostic labels.

[0036] The dermoscopy images in the training sample set are identified and filtered, and the corresponding feature vectors are extracted and concatenated as input to the neural network model.

[0037] Using the diagnostic labels as supervision, the neural network model is trained to obtain a trained classification model.

[0038] According to one specific implementation, in the above-mentioned auxiliary diagnostic method for skin diseases, the RCM image set includes at least one of the following:

[0039] Image sequences obtained by performing a predefined grid scan within the target skin region;

[0040] Multiple frames of images obtained by video stream scanning within the target skin area;

[0041] Multi-layered images obtained by performing layered scanning at one or more specific depths within the target skin region.

[0042] According to one specific implementation, the above-mentioned auxiliary diagnostic method for skin diseases further includes:

[0043] An interpretability report is generated based on the skin disease classification results. The interpretability report is used to display the semantic category and location information of the target feature region, the associated feature region, and at least one or more combinations of the first feature vector and the second feature vector.

[0044] Secondly, the present invention provides a multimodal optical image fusion-based auxiliary diagnostic system for skin diseases, comprising:

[0045] The semantic segmentation module is used to perform semantic segmentation on the acquired dermoscopic image of the target skin region and identify at least one target feature region.

[0046] The association filtering module is used to filter out at least one associated feature region from the acquired RCM image set of the target skin region based on the semantic category and location information of the target feature region.

[0047] The feature extraction module is used to extract a first feature vector from the target feature region and to extract a second feature vector from the associated feature region.

[0048] The fusion decision module is used to fuse the first feature vector and the second feature vector, and to classify the fused feature vector for diseases based on the trained classification model, so as to obtain the skin disease classification result of the target skin region.

[0049] According to a specific implementation, in the above-mentioned auxiliary diagnostic system for skin diseases, the association screening module is specifically used for:

[0050] The spatial range corresponding to the RCM image set is determined based on the location information;

[0051] Within the spatial range, based on the preset correspondence between semantic categories and RCM image features, the correlation between each image or image region in the RCM image set and the target feature region is evaluated;

[0052] Based on the correlation degree, relevant feature regions are selected from the spatial range.

[0053] According to one specific implementation, the above-mentioned dermatology auxiliary diagnosis system further includes a report generation module, which is used to generate an interpretable report based on the dermatology classification results. The interpretable report is used to display the semantic category and location information of the target feature region, the associated feature region, and at least one or more combinations of the first feature vector and the second feature vector.

[0054] Thirdly, the present invention provides an auxiliary diagnostic system for skin diseases, the system comprising a memory and a processor;

[0055] The memory is used to store computer programs; the processor is used to call and execute the computer programs so that the system performs the skin disease auxiliary diagnosis method based on multimodal optical image fusion as described above.

[0056] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0057] This invention achieves automated and precise semantic association between dermoscopic image features and RCM image features by constructing an automated multimodal image processing workflow. Based on the results of dermoscopic semantic segmentation, this invention can select corresponding associated regions from the RCM image set, replacing inefficient and subjective manual matching, significantly improving the efficiency and reliability of multimodal data integration. Simultaneously, by fusing the two feature vectors selected through association, this invention can generate a more comprehensive and accurate integrated feature representation than any single-modal analysis, thus providing complementary and mutually corroborating information for the final disease classification, fundamentally improving the objectivity and accuracy of skin disease classification. The entire process achieves end-to-end automation from data association to feature fusion, reducing reliance on professional experience and providing core technical support for constructing standardized and reproducible multimodal classification and recognition of skin diseases. Attached Figure Description

[0058] Figure 1 is a flowchart illustrating an auxiliary diagnostic method for skin diseases based on multimodal optical image fusion provided in an embodiment of the present invention;

[0059] Figure 2 is a schematic diagram of a dermoscopic image of a region with thick scales marked according to an embodiment of the present invention;

[0060] Figure 3 is a schematic diagram of an RCM image of capillary loops labeled according to an embodiment of the present invention;

[0061] Figure 4 is a schematic diagram of an RCM image of a labeled Munro microabscess provided in an embodiment of the present invention. Detailed Implementation

[0062] The present invention will now be described in further detail with reference to specific embodiments. However, this should not be construed as limiting the scope of the present invention to the following embodiments; all technologies implemented based on the content of the present invention fall within the scope of the present invention.

[0063] Unless otherwise specified, the use of terms such as "first," "second," and "third" in the description of specific embodiments of the present invention is merely for distinguishing descriptions of identical or similar components and should not be construed as emphasizing or implying the relative importance of a particular component.

[0064] Furthermore, in the description of the embodiments of the present invention, "several", "more than", and "a number of" represent at least two. The number can be any number, such as two, three, four, five, six, seven, eight, or nine, and can even exceed nine.

[0065] Existing analyses of dermoscopy and reflex confocal microscopy (RCM) images are mostly conducted independently, lacking effective automated correlation mechanisms. This results in the macroscopic morphological and microscopic structural information contained in the two modalities being fragmented and difficult to integrate. Although some studies have attempted to simply stitch together or jointly analyze the features of the two images, a fundamental problem remains unsolved: how to automatically and accurately locate and filter corresponding key regions in the RCM image set based on the semantic features with clear pathological orientation inherent in dermoscopy images, thereby achieving effective semantic correlation and mutual verification between the two heterogeneous image data. The lack of such a correlation mechanism makes subsequent multimodal feature fusion lack focus, affecting the accuracy and reliability of the final image classification.

[0066] While RCM (Remote Clinical Massage) is non-invasive, the scanning process is time-consuming and requires a high level of operator skill. Traditionally, blind RCM scans of large areas are inefficient and wasteful of medical resources. Traditional result interpretation relies heavily on individual experience and knowledge. Insufficient experience or the presence of rare diseases can lead to poor diagnostic consistency and reliability.

[0067] Based on this, the present invention provides a multimodal skin image processing method and system, the core of which lies in constructing a cross-modal image association and fusion framework based on semantics. The method first performs semantic segmentation on dermoscopic images to identify target feature regions. Then, based on the semantic category and location of these regions, it filters out associated feature regions from the acquired RCM image set. Finally, it fuses feature information from two modalities to output a classification result.

[0068] The technical solution provided by the present invention will be further introduced and described below with reference to specific embodiments.

[0069] Please refer to Figure 1, which shows a flowchart of an auxiliary diagnostic method for skin diseases based on multimodal optical image fusion provided by an embodiment of the present invention, including:

[0070] Step 1: Perform semantic segmentation on the dermoscopic image of the acquired target skin region to identify at least one target feature region.

[0071] Specifically, dermoscopic images are digital images acquired using standardized dermoscopic equipment on clinically suspicious skin areas of the person being evaluated. Before semantic segmentation, preprocessing operations are typically performed to improve the accuracy and precision of subsequent semantic segmentation. These operations include color calibration, illumination normalization, hair occlusion removal, and preliminary image segmentation (used to separate lesion areas from normal skin).

[0072] The target feature regions are identified using existing semantic segmentation algorithms. After identification, a mask map with semantic category and location information is generated, such as "Region A: Thick scales" or "Region B: Circular blood vessels". These semantic categories are mainly divided into four categories: color / pigmentation, surface structure, blood vessel morphology, and special patterns. Each category can be identified by the algorithm and segmented into independent regions.

[0073] For example, color and pigment feature classes include:

[0074] Erythematous region: Visual description: red to purplish-red patches; quantifiable parameters: average hue (H value), saturation, brightness, uniformity of distribution (color standard deviation), and boundary sharpness (gradient variation).

[0075] Pigmented areas: Visual description: Brown to grayish-black dots, lines, networks, or homogeneous areas. Quantifiable parameters: Pigment density, particle size and distribution, color intensity.

[0076] Areas of hypopigmentation / loss: Visual description: Areas that are lighter than the surrounding skin. Quantifiable parameters: Brightness value, contrast with surrounding skin.

[0077] Surface structure and scaling characteristics include:

[0078] Thick scaly areas: Visual description: White or silvery-white, layered mica-like or plate-like covering. Quantifiable parameters: scaly thickness (estimated via shading / stereoscopic vision algorithms), percentage of coverage area, density, and degree of edge lifting.

[0079] Fine / fragmented scaly areas: Visual description: Fine, powdery, or bran-like white deposits. Quantifiable parameters: Scale particle size, uniformity of distribution.

[0080] Collar-like scaly region: Visual description: A collar-like structure with the free edge of the scales pointing outwards and the center adhering to it. Quantifiable parameters: Integrity and width of the ring structure.

[0081] Keratin plugs / follicular keratosis plug area: Visual description: Keratinous plugs at the opening of the hair follicle, appearing as punctate or cup-shaped depressions. Quantifiable parameters: Keratin plug density, size, and distribution pattern (whether it is limited to the hair follicle).

[0082] Vascular morphological characteristics include:

[0083] Dotted vascular region: Visual description: evenly distributed small red dots. Quantifiable parameters: density, spacing, and size uniformity of the punctate vessels.

[0084] Circular / globose vascular regions: Visual description: Clearly identifiable red ring-shaped or coil-like structures. Quantifiable parameters: Ring diameter, ring wall thickness, integrity score (whether closed), and distribution density.

[0085] Linear / curved vascular region: Visual description: thin, short, or curved red lines. Quantifiable parameters: vessel length, curvature, branching.

[0086] Dendritic / irregularly dilated vascular region: Visual description: large, irregularly shaped red blood vessels. Quantifiable parameters: changes in vessel diameter, complexity of vessel course.

[0087] Special pattern feature classes include:

[0088] Wickham pattern area: Visual description: fine, net-like or radial silvery-white linear stripes. Quantifiable parameters: stripe width, regularity of the net-like structure, contrast.

[0089] Crack / erosion area: Visual description: linear or star-shaped red fracture lines, with a moist or crusted surface. Quantifiable parameters: crack length, depth (indirectly), and number.

[0090] Skin groove / skin field pattern alteration area: Visual description: widening, deepening, or disappearance of normal skin texture. Quantifiable parameters: texture roughness, consistency of skin ridge direction.

[0091] Furthermore, for the identified target feature regions, corresponding location information can be generated based on the pixel coordinates of the center point of the region, as well as the size and boundaries of various identified features. For example, for a punctate blood vessel region, the generated location information label can be "coordinates (x, y), radius r", and for a thick scaly region, the generated location information label can be "coordinates (x1, y1) to (x2, y2)". The former can indicate the location information of a circular region, and the latter can indicate the location information of a rectangular region. Most feature regions can be represented in the above manner. In specific implementation, corresponding settings can be made according to different features, which are not limited in this embodiment.

[0092] Step 2: Based on the semantic category and location information of the target feature region, at least one associated feature region is selected from the acquired RCM image set of the target skin region.

[0093] In this step, the RCM image set for the target skin region is two-dimensional or three-dimensional volumetric data generated from a large-scale scan or video stream of the clinically suspicious skin region of the person being evaluated. This step aims to automatically and accurately select the most relevant and diagnostically valuable subset of images (or image regions) from a pre-obtained, large-scale, multi-frame RCM image dataset that is most closely related to specific semantic features discovered by dermoscopy. In one possible implementation, the RCM image set includes at least one of the following: an image sequence obtained by performing a predefined gridded scan within the target skin region; multiple frames of images obtained by performing a video stream scan within the target skin region; or multi-layered images obtained by performing a layered scan at one or more specific depths within the target skin region.

[0094] Specifically, the spatial range corresponding to the RCM image set is determined based on the location information. For example, the spatial coordinates of the dermoscopy images and the RCM image set can be roughly mapped first. This can be done by using physical location markers or matching common feature points through image features (such as pores, specific pigment spots, etc.), or by recording the initial position and scanning path of the RCM scan and establishing a coordinate system transformation relationship with the dermoscopy images. This constructs a mapping function F:(x derm ,y derm )≈(x rcm ,y rcm ,z rcm This is to ensure that each point on the dermoscopy image corresponds to an approximate spatial location in the RCM dataset. Where x... derm y derm Let x and y be the geometric center of the target feature region. rcm ,y rcm ,z rcmThe x, y, and z coordinates represent the spatial range within the RCM image set.

[0095] Specifically, the semantic segmentation results of the dermoscopy are obtained (e.g., "a 'circular blood vessel' exists within the range of coordinates (x,y) and radius r"). The results include a mask and a semantic category label. The mask marks all pixel positions of the target feature region in the dermoscopy image, as well as the geometric center coordinates.

[0096] According to the mapping function F, the geometric center coordinates of the target feature region on the dermoscopy are mapped to the RCM image set. Then, based on the range of the target feature region, considering registration error and the possible depth distribution of features, a three-dimensional candidate space range centered on this geometric center coordinate is defined, with the range size equivalent to a circle of radius r. According to the preset correspondence between semantic categories and RCM image features, the pathological change corresponding to "ring vessels" (dermal papillary vascular dilation) is mainly located in the dermal papillary layer; therefore, the depth range can be selected as [80μm, 150μm]. It is understood that the value of this depth range is only an example and is not the only implementation of this invention. In specific implementation, the feature categories contained in different target feature regions can be determined according to the preset correspondence between semantic categories and RCM image features, and their approximate spatial range in the RCM image can be determined. This is not limited in this embodiment of the invention.

[0097] Furthermore, within the spatial range, based on the preset correspondence between semantic categories and RCM image features, the correlation between each image or image region in the RCM image set and the target feature region is evaluated.

[0098] Specifically, within this spatial range, the correlation between each frame of the RCM image and the target feature region at each depth is calculated. The correlation can be obtained by weighted fusion of different sub-scores to arrive at a more accurate correlation region.

[0099] For example, a lightweight, pre-trained RCM feature classifier can be invoked. This RCM feature classifier can output the probability that an input RCM image patch contains RCM features. Based on the predefined correspondence between semantic categories and RCM image features, a probability vector containing the corresponding features can be calculated, which serves as the feature matching score of the RCM image.

[0100] Furthermore, based on the preferred depth that may appear in the RCM image features, the positional distance from different depth ranges to the preferred depth can be calculated as a positional score.

[0101] Furthermore, the feature consistency of different RCM images in their local regions can be evaluated to avoid selecting isolated noise points. For example, taking (i,j,k) as the center of the RCM image set within this depth range, a small 3x3x3 neighborhood is selected, and the mean and standard deviation of the texture details of all images within the neighborhood are calculated as a consistency score.

[0102] Finally, the feature matching score, location score, and consistency score of the RCM images at each depth range are weighted and fused to calculate the correlation. It is understood that the weighted fusion can assign different weights to each score to emphasize the primacy of its matching degree. This weight is related to the semantic category and can be specifically set in actual implementation; this embodiment of the invention does not impose any limitations.

[0103] Furthermore, based on the aforementioned correlation, the RCM image with the highest correlation is selected as the associated feature region within that spatial range. It can be understood that this associated feature region corresponds to the target feature region and may contain several frames of images with the same xy position but different depths, all exhibiting RCM images with high correlation that represent the semantic category obtained from identifying the target feature region.

[0104] Understandably, the aforementioned correlation calculation decomposes the abstract concept of "correlation" into three dimensions: computable microscopic feature matching degree, spatial proximity, and regional consistency. The entire process is driven by the semantic category of dermoscopy and relies on predefined "semantic-microscopic feature" correspondences that embody medical knowledge. Through comprehensive scoring, a carefully selected, high-quality subset of correlated images is ultimately output, providing accurate microscopic evidence for subsequent feature fusion.

[0105] Specifically, in this embodiment of the invention, the correspondence between semantic categories and RCM image features includes at least one of the following: circumferential vascular regions correspond to the dermal papillary layer vascular dilation and tortuosity features in RCM images; thick scaly regions correspond to epidermal confluent parakeratosis and Munro microabscess features in RCM images; fine, fragmented, or collar-shaped scaly regions correspond to epidermal focal parakeratosis features in RCM images; Wickham's lines correspond to the epidermal granular layer wedge-shaped thickening and hyperkeratosis features in RCM images; and follicular plug regions correspond to the epidermal stratum corneum thickening and follicular infundibulum enlargement features in RCM images. The above correspondences are the gold standard for various skin disease features in dermoscopy images and RCM images commonly used in actual implementation, and are not limited in this embodiment. For example, in the RCM image corresponding to the Wickham's lines region, there is also a corresponding epidermal granular layer acanthosis feature. Simultaneously, the hyperkeratosis is located above the epidermal granular layer, appearing dense, with its center surrounding sweat pores and hair follicles. For example, in the RCM image corresponding to the hair follicle plug area, the hair follicle usually also has the feature of being filled with keratin-like material.

[0106] It's important to note that hyperkeratosis refers to an abnormal thickening of the stratum corneum. Hyperkeratosis can be absolute, meaning the stratum corneum is significantly thicker than normal in the same area, or it can be relative, meaning the stratum corneum is relatively thickened due to the thinning of the stratum spinosum. Hyperkeratosis can consist of completely keratinized cells, i.e., orthohyperkeratosis or hyperkeratosis, and may also be accompanied by parakeratosis (nuclear remnants within the stratum corneum cells). Orthohyperkeratosis has three forms:

[0107] 1. Basket pattern: The stratum corneum is in a normal basket shape, but thicker than normal, such as in tinea versicolor;

[0108] 2. Dense type: such as in neurodermatitis;

[0109] 3. Lamellar type: such as in common ichthyosis.

[0110] Step 3: Extract the first feature vector from the target feature region, and extract the second feature vector from the associated feature region.

[0111] The extraction of the first feature vector in the target feature region includes: extracting at least one of the color statistical features, texture features, geometric features and morphological features in the target feature region, and converting them into vectors to obtain the first feature vector.

[0112] In one possible implementation, the above features can be quantized and combined into a one-dimensional vector to obtain the first feature vector. For example, the mean and standard deviation of pixels within the target feature region in the CIELAB color space are calculated as color statistical features. Contrast, correlation, energy, and homogeneity are calculated using the Gray-Level Co-occurrence Matrix (GLCM) as texture features. The region area, perimeter, roundness, and compactness are calculated as geometric morphological features. The model confidence score identifying the region as a morphological category of different diseases is used as morphological features, such as "thick scales" or "punctate blood vessels." Then, at least one of the above scalar values ​​is concatenated sequentially to form a one-dimensional vector array, which is then Z-score normalized to obtain the first feature vector.

[0113] In one possible implementation, a pre-trained convolutional neural network can be used as a feature extractor to extract a high-dimensional vector from the target feature region as the first feature vector.

[0114] In one possible implementation, the one-dimensional vector and the high-dimensional vector can be merged by concatenation or by projection and concatenation to obtain the first feature vector.

[0115] Further, extracting the second feature vector in the associated feature region includes: extracting at least one of the epidermal cell structure features, dermal papillary layer features, inflammatory cell infiltration pattern features, and collagen fiber features in the associated feature region, and converting them into vectors to obtain the second feature vector.

[0116] The epidermal cell structure features include at least one of hyperkeratosis, spongiosis, and disordered arrangement of keratinocytes; the dermal papillary layer features include at least one of dermal papillary edema and dermal papillary vascular dilation; and the inflammatory cell infiltration pattern features include at least one of lichenified band infiltration and perivascular infiltration.

[0117] Specifically, the following operations are performed in parallel on each frame of RCM image in the associated feature region:

[0118] Cellular-level quantization features: Keratinocytes and dermal papillae are automatically segmented from the image, and the following are calculated: epidermal cell density (cells / mm²), average brightness and coefficient of variation of cell nuclei; vascular diameter (pixel width) and vascular tortuosity in dermal papillae; area and number of Munro microabscesses; and these scalar values ​​are combined into a quantization array.

[0119] Deep learning feature extraction: Each frame of the RCM image is input into another pre-trained CNN feature extractor, which outputs a deep feature vector.

[0120] Since the associated feature region contains multiple frames of RCM images (which may be captured from different depths), it is necessary to aggregate the features of multiple frames into a unified feature representation.

[0121] First, quantization feature aggregation is performed. Specifically, the average and maximum values ​​(pooling) of the quantization arrays of all frames in the multi-frame RCM image are calculated. For example, for the feature "vessel diameter", the average value (reflecting the overall degree of dilation) and the maximum value (reflecting the most severe local dilation) are taken. The results of average pooling and max pooling are concatenated to form the aggregated quantization array.

[0122] Next, deep feature aggregation is performed. Specifically, attention pooling is applied to the deep feature vectors of all frames in the multi-frame RCM image. An attention weight is learned for each frame, and the aggregated deep features are obtained by weighted summation.

[0123] Attention weights can be calculated by a small neural network based on the deep feature vectors themselves, allowing the model to focus on frames that are more valuable for subsequent classification.

[0124] Furthermore, the aggregated quantized features are fused with the deep features to generate the final second feature vector. This can be achieved using a projection and concatenation method: the aggregated quantized array is projected to the same dimension as the aggregated deep features, and then concatenated with the aggregated deep features to obtain the second feature vector.

[0125] Step 4: Fuse the first feature vector and the second feature vector, and classify the fused feature vector for diseases based on the trained classification model to obtain the skin disease classification result of the target skin region.

[0126] According to a specific implementation method, the first feature vector and the second feature vector are fused, and disease classification is performed on the fused feature vector based on a trained classification model, specifically including:

[0127] The first feature vector and the second feature vector are concatenated to obtain the fused feature vector;

[0128] The fused feature vector is input into the trained classification model to obtain the corresponding skin disease classification result; the skin disease classification result includes the probability that the target skin region belongs to at least one skin disease.

[0129] Specifically, the first feature vector and the second feature vector can be concatenated along the feature dimension to obtain the fused feature vector.

[0130] The training process of the classification model includes:

[0131] The training sample set is based on the acquired dermoscopy image set, the corresponding RCM image set, and the pathologically confirmed diagnostic labels.

[0132] The dermoscopy images in the training sample set are identified and filtered, and the corresponding feature vectors are extracted and concatenated as input to the neural network model.

[0133] Using the diagnostic labels as supervision, the neural network model is trained to obtain a trained classification model.

[0134] Specifically, the classification model can be structured as a multilayer perceptron (MLP), for example, with the structure: 1024 -> 256 -> 64 -> N, where N is the number of disease categories (e.g., psoriasis, eczema, lichen planus, etc.). Each layer is followed by a ReLU activation function and a Dropout layer to prevent overfitting.

[0135] Forward propagation: The fused feature vectors pass through each layer of the MLP sequentially.

[0136] Output: The last layer of the MLP uses the Softmax activation function and outputs a probability distribution vector P=[p1,p2,...,pN].

[0137] For example, P=[0.87,0.10,0.03] means that the model judges the probability of having psoriasis in this area to be 87%, the probability of having eczema to be 10%, and the probability of having other diseases to be 3%.

[0138] The probability distribution vector P represents the skin disease classification result, which is the final decision-making embodiment after intelligent and non-linear fusion of dermoscopic image features and RCM image features. This step reveals the key transformation process from heterogeneous, multi-source feature information to unified, computable decision-making. It combines interpretable, manually quantified features (reflecting known medical knowledge) with powerful deep learning features (automatically mining latent patterns), balancing interpretability and performance. Furthermore, for multiple RCM images from the same region, it innovatively employs a combination of attention pooling and statistical pooling to achieve efficient, information-preserving compression from the image set to a single feature vector. The final fusion and classification are not simple weighting, but are achieved through a trainable neural network, enabling the model to automatically learn the complex, non-linear interaction between dermoscopic image features and RCM image features from the data, thereby making better classification and recognition.

[0139] It should be noted that the above-described identification and filtering of dermoscopic images in the training sample set is the same as the identification of target feature regions and the filtering of associated feature regions in the acquired dermoscopic images and RCM image sets of the target skin region in the embodiments of the present invention, as described in steps 1 to 3 above, thereby extracting the corresponding feature vectors. The training objective of the classification model is to obtain a model that can accurately output the probability of its class based on the fused feature vectors.

[0140] In addition, the aforementioned pathologically confirmed diagnostic labels are diagnostic results confirmed by skin tissue pathological biopsy, serving as the indisputable gold standard labels for supervised training of the classification model in the embodiments of the present invention.

[0141] In order to enable the final classification results provided by the embodiments of the present invention to intuitively display the correspondence between dermoscopic image features and RCM image features, and to improve the practical value of the embodiments of the present invention, the auxiliary diagnostic method for skin diseases provided by the embodiments of the present invention further includes generating an interpretability report based on the skin disease classification results. The interpretability report is used to display the semantic category and location information of the target feature region, the associated feature region, and at least one or more combinations of the first feature vector and the second feature vector.

[0142] Specifically, the interpretability report can be pre-set with templates and rules according to the habits of different individuals being assessed and users, as well as the requirements of the corresponding conditions. It can directly output the results of the above steps (images of the target feature regions and associated feature regions) or convert them into text descriptions. It can be output according to a standard clinical report format. For example, when the skin disease category of the individual being assessed is identified as psoriasis, the extracted target feature regions include annular vascular regions and thick scaly regions. Associated feature regions include significantly dilated and tortuous capillary loops visible in the papillary dermis corresponding to the annular vessels, and Munro microabscesses corresponding to the thick scales. At the same time, no other major features of skin diseases were found, such as spongiosis in RCM images of eczema and lichen planus in RCM images of lichen planus or Wickham's striae in dermoscopy images. The final interpretability report can be:

[0143] I. Imaging findings

[0144] Dermatoscope:

[0145] Circular vascular area: Highly regular red ring-shaped structures were detected, suggesting possible psoriasis.

[0146] Thick scaly areas: Homogeneous thickened silvery-white scaly coverage was detected.

[0147] Reflection confocal microscope - RCM:

[0148] Evidence A (corresponding to the annular vessels): Significantly dilated and tortuous capillary loops (diameter 28 μm) are visible in the papillary dermis.

[0149] Evidence B (corresponding to thick scales): epidermal confluent parakeratosis and Munro microabscess (neutrophil aggregation).

[0150] II. Multimodal Fusion Analysis and Diagnosis

[0151] Primary diagnosis: Psoriasis (Vulgaris Psoriasis)

[0152] Diagnostic confidence level: 91%

[0153] Main diagnostic criteria:

[0154] Deterministic evidence: Munro microabscesses were directly observed under RCM (a pathologically specific marker of psoriasis).

[0155] Strong supporting evidence: Dermal papillary vascular dilation was observed under RCM, which is consistent with the dermoscopic annular vascular pattern, together constituting the typical vascular changes in psoriasis.

[0156] Differential diagnosis and reasons for exclusion:

[0157] Eczema (7% probability): No key RCM feature "sponge edema" was observed.

[0158] Lichen planus (probability 2%): No RCM key feature "lichenoid lymphocytic band infiltration" or dermoscopic typical "Wickham's striae" was observed.

[0159] III. Conclusion

[0160] Based on the combined dermoscopic and RCM imaging features, this lesion is highly consistent with psoriasis.

[0161] Please refer to Figures 2 to 4, which respectively show schematic diagrams of dermoscopy images of areas with thick scales, RCM images of capillary loops, and RCM images of Munro microabscesses provided during the implementation of embodiments of the present invention, corresponding to the above-mentioned interpretability report.

[0162] Based on the above, the assisted diagnostic method provided in this invention does not merely output a "black box" conclusion. Instead, it automatically links the dermoscopic image feature regions, RCM image feature regions, and their key findings in a spatially and semantically corresponding manner, forming a complete logical system for disease classification / identification, and presenting it intuitively through a combination of text and graphics. The report reveals the quantitative contribution of different features to the final diagnosis (e.g., Munro microabscess contributes 40%), making the decision-making process using artificial intelligence in this invention transparent, understandable, and questionable, meeting the interpretability (XAI) requirements of assisted AI. Furthermore, the report transforms unstructured algorithmic data into a structured report that conforms to clinical reading habits, improving the report's professionalism and practicality, and allowing users to quickly grasp the key points.

[0163] Based on the above technical solution, this invention constructs an automated multimodal image processing workflow to achieve automated and accurate semantic association between dermoscopic image features and RCM image features. This invention can select corresponding associated regions from the RCM image set based on the results of dermoscopic semantic segmentation, replacing inefficient and subjective manual matching, significantly improving the efficiency and reliability of multimodal data integration. Simultaneously, by fusing the two feature vectors selected through association, this invention can generate a more comprehensive and accurate integrated feature representation than any single-modal analysis, thus providing complementary and mutually corroborating information for the final disease classification, fundamentally improving the objectivity and accuracy of skin disease classification. The entire process achieves end-to-end automation from data association to feature fusion, reducing reliance on professional experience and providing core technical support for constructing standardized and reproducible multimodal classification and recognition of skin diseases.

[0164] On the other hand, embodiments of the present invention also provide a multimodal optical image fusion-based auxiliary diagnostic system for skin diseases, comprising:

[0165] The semantic segmentation module is used to perform semantic segmentation on the acquired dermoscopic image of the target skin region and identify at least one target feature region.

[0166] The association filtering module is used to filter out at least one associated feature region from the acquired RCM image set of the target skin region based on the semantic category and location information of the target feature region.

[0167] The feature extraction module is used to extract a first feature vector from the target feature region and to extract a second feature vector from the associated feature region.

[0168] The fusion decision module is used to fuse the first feature vector and the second feature vector, and to classify the fused feature vector for diseases based on the trained classification model, so as to obtain the skin disease classification result of the target skin region.

[0169] Specifically, the association filtering module is used for:

[0170] The spatial range corresponding to the RCM image set is determined based on the location information;

[0171] Within the spatial range, based on the preset correspondence between semantic categories and RCM image features, the correlation between each image or image region in the RCM image set and the target feature region is evaluated;

[0172] Based on the correlation degree, relevant feature regions are selected from the spatial range.

[0173] Furthermore, the aforementioned auxiliary diagnostic system also includes a report generation module, which is used to generate an interpretable report based on the skin disease classification results. The interpretable report is used to display the semantic category and location information of the target feature region, the associated feature region, and at least one or more combinations of the first feature vector and the second feature vector.

[0174] The following section uses lichen planus as an example, whose clinical manifestation is purplish-red papules, to provide a detailed introduction and explanation of the multimodal optical image fusion-based auxiliary diagnostic system for skin diseases provided in this embodiment of the invention.

[0175] First, before the system runs, a dermoscopic color image of the skin area is obtained using a dermatoscope. Then, a grid scan or video stream scan is performed using a reflective confocal microscope to obtain a sequence of RCM grayscale images (i.e., an RCM image set) containing multiple frames and depths. This image set covers three-dimensional spatial information from the stratum corneum of the epidermis to the papillary layer of the dermis.

[0176] The images described above are input into the dermatology auxiliary diagnostic system, where the semantic segmentation module processes the acquired dermoscopic images. A pre-trained deep semantic segmentation network (such as U-Net or its variants) is used, trained on a large dataset of dermoscopic images labeled with skin pathological features. The dermoscopic images are input into this network, which outputs a pixel-level semantic segmentation map. In this embodiment, the network successfully identified and segmented two key target feature regions:

[0177] Region A: Classified as "Wickham stripes". This region appears as fine, net-like, or radial silvery-white linear stripes. The segmentation network outputs a binary mask for this region, along with a classification confidence score, such as 0.90.

[0178] Region B: Classified as "purplish-red spot". This region presents a homogeneous purplish-red background. The network outputs its mask and confidence score, for example, 0.85.

[0179] The semantic segmentation module also calculates attributes such as the geometric center coordinates of each region to provide spatial location information for subsequent association.

[0180] Furthermore, the association filtering module receives the semantic segmentation results. Its core task is to automatically filter out the most relevant subset of images, i.e., the associated feature regions, from the huge RCM image set for each identified target feature region based on its semantic category and location.

[0181] For region A, the association filtering module maps the center position of region A to the RCM three-dimensional space based on the pre-stored dermoscopy-RCM coordinate mapping relationship, and delineates a candidate cube range mainly composed of the epidermis and the dermal-epidermal junction (DEJ), taking into account registration errors. The system calls the predefined "semantic-feature correspondence database". This database indicates that the "Wickham striae" of the dermoscopy mainly correspond to the two microscopic features of "superficial dermal lichenoid lymphocytic infiltration" and "serrated epidermal ridges" under RCM.

[0182] The association screening module uses a lightweight RCM feature detector to quickly scan each frame of RCM images within the candidate range. The detector evaluates the likelihood probability of the presence of "dense, highly refractive circular cells arranged in a band below the DEJ" (i.e., lichenoid infiltration) in the image, as well as the prominence of the "zigzag" appearance of the epidermal-dermal junction. Combining the likelihood probability, prominence, distance weights between the image and the mapping center, and local consistency, a comprehensive association score is calculated between each frame of RCM image and the "Wickham's striae" region. The top K frames (e.g., K=2) of RCM images with the highest association scores are selected to form the subset of images {RCM_A} associated with region A. These images are expected to exhibit a typical pattern of lichenoid infiltration at the cellular level.

[0183] Furthermore, for region B, similarly, a candidate range is defined (possibly with a greater focus on the papillary dermis). Based on the relational database, "purplish-red spots" correspond to features such as "vasodilation and congestion in the papillary dermis" and "basal cell liquefaction and degeneration." The degree of vasodilation and the integrity of the basal layer structure in the candidate RCM images are evaluated, and a correlation score is calculated. A subset of images with high correlation, {RCM_B}, is selected.

[0184] The feature extraction module extracts features from regions A (Wickham lines) and B (purplish-red spots) in the dermoscopy image. For region A, it extracts texture features (such as the grid regularity and line width of the Wickham lines), color features (brightness and saturation statistics of silvery-white), and geometric features of the region to form the first feature vector A. A similar operation is performed on region B to obtain the first feature vector B.

[0185] The selected RCM-related image subsets {RCM_A} and {RCM_B} are processed separately. Taking {RCM_A} as an example, for each frame in the subset, a dedicated RCM CNN network is used to extract high-dimensional depth features. Simultaneously, quantization analysis is performed: for example, calculating the density and uniformity of highly refractive inflammatory cells in suspected infiltration areas; measuring the depth and spacing of the "teeth" of epidermal ridges. The depth features of all frames in the subset are aggregated using attention pooling, and the quantization features are aggregated using statistical pooling (e.g., taking the mean and maximum values), finally fused into a second feature vector A representing the region. Similarly, {RCM_B} is processed to obtain the second feature vector B.

[0186] The fusion decision module pairs and concatenates the first and second feature vectors from the same semantic target. For example, it concatenates the first feature A with the second feature vector A to form a fused feature vector A; and it concatenates the first feature vector B with the second feature vector B to form a fused feature vector AB. All the paired and fused feature vectors are then input into a pre-trained classification neural network model (e.g., a multilayer perceptron, MLP). This classification model, trained on multimodal data containing various skin diseases (such as lichen planus, discoid lupus erythematosus, and eczema), learns the complex mapping relationship between different feature combinations and disease categories. Finally, the classification model outputs a probability distribution vector, for example: lichen planus: 0.89, discoid lupus erythematosus: 0.08, others: 0.03. This vector is the final classification result obtained by this system after processing the input multimodal skin images.

[0187] The report generation module can automatically generate an interpretable report based on all the intermediate data and the final classification results. The report is presented in a graphic format: the original dermoscopy image with regions A and B highlighted; the selected key RCM images {RCM_A} and {RCM_B} are displayed, and the second feature vectors identified are marked on the images (e.g., arrows are used to indicate band-like infiltration); finally, the classification results and the main feature pairs used as the basis are listed (e.g., "Wickham striae" corresponds to "moss-like infiltration"), making the entire processing clear and traceable.

[0188] Based on the above technical solution, this embodiment achieves automated and precise semantic association between macroscopic findings from dermoscopy and microscopic evidence from RCM through the core steps of "semantic segmentation -> association filtering," overcoming the problem of data fragmentation between the two modalities in existing technologies. Subsequent targeted feature fusion enables the final classification model to comprehensively utilize validated and associated multi-scale information, significantly improving the accuracy, robustness, and interpretability of image classification.

[0189] On the other hand, embodiments of the present invention also provide an auxiliary diagnostic system for skin diseases, the system including a memory and a processor;

[0190] The memory is used to store computer programs; the processor is used to call and execute the computer programs so that the system performs the skin disease auxiliary diagnosis method based on multimodal optical image fusion as described above.

[0191] In embodiments of the present invention, the processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0192] The various methods, steps, and logic diagrams disclosed in the embodiments of this invention can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The processor reads information from the storage medium and, in conjunction with its hardware, completes the steps of the above methods.

[0193] The storage medium can be memory, such as volatile memory or non-volatile memory, or may include both volatile and non-volatile memory.

[0194] Among them, non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory.

[0195] Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), sync link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM).

[0196] The storage media described in the embodiments of the present invention are intended to include, but are not limited to, these and any other suitable types of memory.

[0197] It should be understood that the system disclosed in the embodiments of the present invention can be implemented in other ways. For example, the division of modules is merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the communication connection between modules can be through some interfaces, indirect coupling or communication connections between servers or units, and can be electrical or other forms.

[0198] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each module can exist physically separately, or two or more modules can be integrated into one processing unit. The integrated unit described above can be implemented in hardware or as a software functional unit.

[0199] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0200] Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner for ease of understanding.

[0201] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for auxiliary diagnosis of skin diseases based on multimodal optical image fusion, characterized in that, include: Perform semantic segmentation on the acquired dermoscopic image of the target skin region to identify at least one target feature region; Based on the semantic category and location information of the target feature region, at least one associated feature region is selected from the acquired RCM image set of the target skin region; a first feature vector is extracted from the target feature region, and a second feature vector is extracted from the associated feature region; the first feature vector and the second feature vector are fused, and disease classification is performed on the fused feature vector based on a trained classification model to obtain the skin disease classification result of the target skin region; wherein, selecting at least one associated feature region from the acquired RCM image set of the target skin region specifically includes: determining the spatial range corresponding to the RCM image set according to the location information; within the spatial range, evaluating the correlation between each image or image region in the RCM image set and the target feature region according to the preset semantic category and RCM image feature correspondence; and selecting associated feature regions from the spatial range according to the correlation; the semantic category includes: color and pigment features. The semantic categories and RCM image features include: erythematous areas, hyperpigmented areas, and hypopigmented areas; surface structure and scaling features, including thick scaling areas, fine scaling areas, collar-shaped scaling areas, and follicular plug areas; vascular morphology features, including punctate vascular areas, annular vascular areas, linear vascular areas, and dendritic vascular areas; and special pattern features, including Wickham's lines, fissure areas, and widened skin grooves. The correspondence between the semantic categories and RCM image features includes at least one of the following: annular vascular areas correspond to the dilation and tortuosity of papillary dermal vessels in RCM images; thick scaling areas correspond to epidermal confluent parakeratosis and Munro microabscesses in RCM images; fine scaling or collar-shaped scaling areas correspond to focal epidermal parakeratosis in RCM images; Wickham's lines correspond to wedge-shaped thickening and hyperkeratosis of the granular layer of the epidermis in RCM images; and follicular plug areas correspond to thickening of the stratum corneum and enlargement of the follicular infundibulum in RCM images.

2. The method for auxiliary diagnosis of skin diseases based on multimodal optical image fusion according to claim 1, characterized in that, Extracting a first feature vector from the target feature region includes: extracting at least one of color statistical features, texture features, geometric features, and morphological features from the target feature region and converting them into vectors to obtain a first feature vector; extracting a second feature vector from the associated feature region includes: extracting at least one of epidermal cell structure features, dermal papillary layer features, inflammatory cell infiltration pattern features, and collagen fiber features from the associated feature region and converting them into vectors to obtain a second feature vector.

3. The method for auxiliary diagnosis of skin diseases based on multimodal optical image fusion according to claim 2, characterized in that, The epidermal cell structure features include at least one of hyperkeratosis, spongiosis, and disordered arrangement of keratinocytes; the dermal papillary layer features include at least one of dermal papillary edema and dermal papillary vascular dilation; the inflammatory cell infiltration pattern features include at least one of lichenified band infiltration and perivascular infiltration.

4. The method for auxiliary diagnosis of skin diseases based on multimodal optical image fusion according to claim 2, characterized in that, The first feature vector and the second feature vector are fused, and disease classification is performed on the fused feature vector based on a trained classification model. Specifically, this includes: concatenating the first feature vector and the second feature vector to obtain a fused feature vector; inputting the fused feature vector into a trained classification model to obtain the corresponding skin disease classification result; the skin disease classification result includes the probability that the target skin region belongs to at least one skin disease.

5. The method for auxiliary diagnosis of skin diseases based on multimodal optical image fusion according to claim 4, characterized in that, The training process of the classification model includes: using the acquired dermoscopic image set, the corresponding RCM image set, and the pathologically confirmed diagnostic labels as the training sample set; identifying and filtering the dermoscopic images in the training sample set, and extracting the corresponding feature vectors and concatenating them as the input of the neural network model; and training the neural network model with the diagnostic labels as supervision to obtain the trained classification model.

6. The method for auxiliary diagnosis of skin diseases based on multimodal optical image fusion according to claim 1, characterized in that, The RCM image set includes at least one of the following: an image sequence obtained by performing a predefined gridded scan within the target skin region; multiple frames of images obtained by performing a video stream scan within the target skin region; and multi-layer images obtained by performing a layered scan at one or more depths within the target skin region.

7. The method for auxiliary diagnosis of skin diseases based on multimodal optical image fusion according to claim 1, characterized in that, The method further includes: generating an interpretability report based on the skin disease classification results, wherein the interpretability report is used to display the semantic category and location information of the target feature region, the associated feature region, and at least one or more combinations of the first feature vector and the second feature vector.

8. A multimodal optical image fusion-based auxiliary diagnostic system for skin diseases, characterized in that, include: The semantic segmentation module is used to perform semantic segmentation on the acquired dermoscopic image of the target skin region and identify at least one target feature region. The association filtering module is used to filter out at least one associated feature region from the acquired RCM image set of the target skin region based on the semantic category and location information of the target feature region. The feature extraction module is used to extract a first feature vector from the target feature region and a second feature vector from the associated feature region; the fusion decision module is used to fuse the first feature vector and the second feature vector, and perform disease classification on the fused feature vector based on a trained classification model to obtain the skin disease classification result of the target skin region; wherein, the association filtering module is specifically used to: determine the spatial range corresponding to the RCM image set according to the location information; within the spatial range, evaluate the correlation degree between each image or image region in the RCM image set and the target feature region according to the preset semantic category and the correspondence relationship between RCM image features; and filter out associated feature regions from the spatial range according to the correlation degree; the semantic categories include: color and pigmentation feature categories, including erythema regions, hyperpigmentation regions, and hypopigmentation regions; surface structure and scales The semantic categories include: thick scaly areas, fine scaly areas, collar-shaped scaly areas, and follicular plug areas; vascular morphology features include: punctate vascular areas, annular vascular areas, linear vascular areas, and dendritic vascular areas; special pattern features include: Wickham's lines, fissure areas, and widened skin grooves. The correspondence between the semantic categories and RCM image features includes at least one of the following: annular vascular areas correspond to the dilation and tortuosity of papillary dermal vessels in RCM images; thick scaly areas correspond to epidermal confluent parakeratosis and Munro microabscesses in RCM images; fine scaly or collar-shaped scaly areas correspond to focal epidermal parakeratosis in RCM images; Wickham's lines correspond to wedge-shaped thickening and hyperkeratosis of the granular layer of the epidermis in RCM images; and follicular plug areas correspond to thickening of the stratum corneum and enlargement of the follicular infundibulum in RCM images.

9. The multimodal optical image fusion-based auxiliary diagnostic system for skin diseases according to claim 8, characterized in that, The system also includes a report generation module, which is used to generate an interpretable report based on the skin disease classification results. The interpretable report is used to display the semantic category and location information of the target feature region, the associated feature region, and at least one or more combinations of the first feature vector and the second feature vector.

10. A skin disease auxiliary diagnostic system, characterized in that, The system includes a memory and a processor; wherein the memory is used to store a computer program; and the processor is used to call and execute the computer program so that the system performs the auxiliary diagnostic method for skin diseases based on multimodal optical image fusion as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Skin disease picture classification method and device, product and storage medium

    CN115005768A

  • Ultrasonic vein puncture system integrating image recognition and data analysis

    CN120605101A