Skin condition classification system based on multimodal feature fusion
The skin condition classification system based on multimodal feature fusion solves the problem of low accuracy in skin disease classification by utilizing text information and image feature fusion model, thus achieving more efficient skin disease diagnosis.
Patent Information
- Application Number
- CN202411645616.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-18
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2044-11-18
AI Technical Summary
In the existing technology, when identifying and classifying skin diseases, due to the small dimensions and limited features of structured oral text, it is difficult to accurately and comprehensively characterize the characteristics of various skin diseases, resulting in low accuracy in skin condition classification.
A skin condition classification system based on multimodal feature fusion is adopted. Structured and unstructured text information and skin lesion images are obtained through the patient condition text information entry device and the skin image acquisition device. The main control chip performs text vectorization and image feature extraction, and combines the pre-trained skin condition classification model to perform multimodal feature fusion to improve classification accuracy.
Through multimodal feature fusion, the classification accuracy of skin conditions is improved, assisting doctors in improving the efficiency of skin disease diagnosis.
Smart Images

Figure CN119600342B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the field of computer technology, and in particular to a skin condition classification system based on multimodal feature fusion. Background Art
[0002] Due to the wide variety of skin diseases and the confusing nature of their features, doctors' visual examinations are inefficient and prone to misdiagnosis. Therefore, there is an urgent need for AI-based skin condition classification systems to improve diagnostic efficiency and quality. Currently, the common approach for skin disease identification and classification is for the skin condition classification system to obtain the patient's skin lesion images and simple structured oral text information (such as age and skin type). The skin lesion classification model then uses the fusion features of the lesion images and structured text to classify the patient's skin condition.
[0003] However, when using the above method to identify and classify skin diseases, the following technical problems often arise:
[0004] Structured oral text usually has fewer dimensions and limited features, making it difficult to characterize the characteristics of various skin diseases. This makes it difficult to accurately and comprehensively characterize the patient's skin lesion status by combining the fusion features of skin lesion images and structured text, resulting in low classification accuracy of skin conditions.
[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure concept and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art. Summary of the Invention
[0006] The content of this disclosure is used to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this disclosure is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0007] Some embodiments of the present disclosure propose a skin condition classification system based on multimodal feature fusion to solve one or more of the technical problems mentioned in the above background technology section.
[0008] Some embodiments of the present disclosure provide a skin condition classification system based on multimodal feature fusion, and the skin condition classification system based on multimodal feature fusion includes: a patient condition text information input device, a skin image acquisition device, a main control chip and a classification result display, wherein the patient condition text information input device is configured to input a patient condition structured text information sequence and a patient condition unstructured text information sequence, and send the patient condition structured text information sequence and the patient condition unstructured text information sequence to the main control chip; the skin image acquisition device is configured to acquire a patient skin lesion image, and send the patient skin lesion image to the main control chip; the main control chip is configured to The state structured text information and the patient state unstructured text information sequence are subjected to text vectorization processing to obtain an initial first state text feature vector and an initial second state text feature vector sequence; the above-mentioned main control chip is configured to perform image feature extraction processing on the received patient skin lesion image to obtain an initial skin lesion area feature vector; the above-mentioned main control chip is configured to input the above-mentioned initial first state text feature vector, the above-mentioned initial second state text feature vector sequence and the above-mentioned initial skin lesion area feature vector into a pre-trained skin state classification model to obtain a skin state classification result, and send the above-mentioned skin state classification result to the above-mentioned classification result display; the above-mentioned classification result display is configured to display the received skin state classification result.
[0009] The above-mentioned embodiments of the present disclosure have the following beneficial effects: through the skin condition classification system based on multimodal feature fusion of some embodiments of the present disclosure, the classification accuracy of skin conditions can be improved. Specifically, the reason for the low accuracy of skin condition classification is that structured oral text usually has fewer dimensions and limited features, making it difficult to characterize the characteristics of various skin diseases. This makes it difficult for the fusion features of skin lesion images and structured text to accurately and comprehensively characterize the patient's skin lesion status, thereby resulting in low skin condition classification accuracy. Based on this, some embodiments of the present disclosure include a skin state classification system based on multimodal feature fusion, including: a patient state text information entry device, a skin image acquisition device, a main control chip, and a classification result display, wherein the patient state text information entry device is configured to enter a patient state structured text information sequence and a patient state unstructured text information sequence, and send the patient state structured text information sequence and the patient state unstructured text information sequence to the main control chip; the skin image acquisition device is configured to acquire a patient skin lesion image, and send the patient skin lesion image to the main control chip; the main control chip is configured to process the received patient state structured text information and The patient's unstructured text information sequence is subjected to text vectorization processing to obtain an initial first-state text feature vector and an initial second-state text feature vector sequence; the above-mentioned main control chip is configured to perform image feature extraction processing on the received patient's skin lesion image to obtain an initial skin lesion area feature vector; the above-mentioned main control chip is configured to input the above-mentioned initial first-state text feature vector, the above-mentioned initial second-state text feature vector sequence and the above-mentioned initial skin lesion area feature vector into a pre-trained skin state classification model to obtain a skin state classification result, and send the above-mentioned skin state classification result to the above-mentioned classification result display; the above-mentioned classification result display is configured to display the received skin state classification result. Therefore, in some embodiments of the present disclosure, the skin state classification system based on multimodal feature fusion can improve the feature dimension of the skin lesion feature by performing multimodal feature fusion on the patient's skin lesion image, the structured text corresponding to the skin state and the unstructured text through the main control chip, so as to more accurately and comprehensively characterize various types of skin diseases, thereby improving the classification accuracy of the skin state. Furthermore, the skin condition classification results can be displayed through the classification result display, thereby assisting doctors in diagnosing skin diseases and improving the efficiency of skin disease diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.
[0011] Figure 1 is an exemplary structural diagram of some embodiments of a skin condition classification system based on multimodal feature fusion according to the present disclosure;
[0012] Figure 2 is an exemplary structural diagram of other embodiments of a skin condition classification system based on multimodal feature fusion according to the present disclosure;
[0013] Figure 3 is a data flow diagram of a skin lesion image feature extraction step in some embodiments of the skin condition classification system based on multimodal feature fusion according to the present disclosure;
[0014] Figure 4 Schematic diagram of experimental data on skin lesion classification performance according to some embodiments of the skin condition classification system based on multimodal feature fusion according to the present disclosure. DETAILED DESCRIPTION
[0015] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0016] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.
[0017] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0018] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0019] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0020] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0021] First, see Figure 1 , Figure 1 This is an exemplary structural diagram of some embodiments of the skin condition classification system based on multimodal feature fusion disclosed herein. The skin condition classification system based on multimodal feature fusion includes: a patient condition text information input device 101, a skin image acquisition device 102, a main control chip 103, and a classification result display 104.
[0022] In some embodiments, the patient status text information entry device 101 can be configured to enter a sequence of structured text information and an unstructured text information sequence of patient status, and to send the structured text information and the unstructured text information sequence to the main control chip 103. The patient status text information entry device can be a terminal with a touch screen display, allowing a user to enter data through keystrokes and interactive selections. The structured text information of the patient status in the structured text information sequence can correspond one-to-one with a patient's structured attributes. The structured attributes can be attributes related to the patient, with corresponding attribute values in a fixed encoding format. The structured text information of the patient status in the structured text information sequence can be structured text describing the patient's status. The structured text can be text corresponding to the attribute values of the structured attributes. Each piece of unstructured text information in the unstructured text information sequence can be natural language text describing the patient's condition, distinct from structured text. The main control chip can be a chip used to diagnose and classify a patient's skin lesions based on the text describing the patient's condition and skin images. The above main control chip can be deployed separately on a server terminal.
[0023] As an example, the above-mentioned patient structured attributes may be, but are not limited to, one of the following: age, gender, skin type, and whether or not the patient has a scarring physique. The patient's age is usually expressed as a numerical value. Gender is usually represented by a single character as "male" or "female". The above-mentioned skin type may be the type of skin. The skin type may be one of the following: normal skin, dry skin, oily skin, mixed skin, and sensitive skin. Whether or not the patient has a scarring physique may be represented by "yes" and "no" respectively to indicate "scarring physique" and "not scarring physique". When the patient structured attribute corresponding to the patient status structured text information is "skin type", the patient status structured text information may be "sensitive skin". The patient status unstructured text information may be "more obvious brown patches on the left cheekbone".
[0024] In practice, the above-mentioned patient status text information entry device can display each patient structured attribute and corresponding attribute value options corresponding to the structured text, a text input box for filling in the unstructured text describing the user's skin condition, and a submit button on the touch screen. First, based on the patient's response to each patient structured attribute corresponding to the structured text, the user can select the attribute value for each attribute on the touch screen included in the patient status text information entry device using a finger or a stylus. Then, based on the description of the patient's main complaint, type the unstructured text in the displayed text input box using a stylus input or a soft keyboard. Finally, the user clicks the submit button to send the above-mentioned patient status structured text information sequence and the above-mentioned patient status unstructured text information sequence to the above-mentioned main control chip via a wired connection or a wireless connection. It should be noted that the above-mentioned wireless connection method may include but is not limited to 3G / 4G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other wireless connection methods currently known or to be developed in the future.
[0025] In some embodiments, the skin image acquisition device 102 can be configured to acquire skin lesion images of patients and send the skin lesion images of patients to the main control chip 103. The skin image acquisition device can be a device equipped with a camera. The skin lesion images of patients can be skin images taken including the skin lesion site of the patient and the area of the patient's body surrounding the lesion site. The skin image acquisition device can acquire skin lesion images of patients through a camera and send the skin lesion images of patients to the main control chip through a wired connection or a wireless connection. It should be noted that the area of the patient's body surrounding the lesion site can be a skin area not covered by clothes or a skin area covered by clothes, which is not specifically limited here.
[0026] In some embodiments, the main control chip 103 can be configured to perform text vectorization processing on the received patient status structured text information sequence and the patient status unstructured text information sequence to obtain an initial first state text feature vector and an initial second state text feature vector sequence. The initial first state text feature vector can be an embedded representation of the patient status structured text information sequence. The initial second state text feature vector in the initial second state text feature vector sequence can correspond one-to-one to the patient status unstructured text information in the patient status unstructured text information sequence. The initial second state text feature vector in the initial second state text feature vector sequence can be an embedded representation of the corresponding patient status unstructured text information.
[0027] In some optional implementations of some embodiments, each patient status structured text information in the above-mentioned patient status structured text information sequence may include a status attribute value. The above-mentioned status attribute value may be an attribute value of a corresponding patient structured attribute. The above-mentioned main control chip 103 may be further configured to perform the following steps to perform text vectorization processing on the received patient status structured text information and patient status unstructured text information sequence to obtain an initial first state text feature vector and an initial second state text feature vector sequence:
[0028] In the first step, for each patient status structured text information in the above patient status structured text information sequence, perform the following steps:
[0029] In a first sub-step, in response to determining that the patient status structured text information satisfies a preset attribute condition, normalizing the status attribute value included in the patient status structured text information to obtain a numerical attribute characteristic value. The preset attribute condition may be that the attribute value of the patient structured attribute corresponding to the patient status structured text information is a numerical value. The numerical attribute characteristic value may represent the status attribute value of the numerical attribute. For example, the numerical attribute may be, but is not limited to, age.
[0030] As an example, the main control chip can normalize the state attribute values included in the patient state structured text information using a preset numerical normalization method to obtain numerical attribute feature values. For example, the numerical normalization method can be a Min-Max normalization method.
[0031] In a second sub-step, in response to determining that the patient status structured text information does not meet the preset attribute conditions, word embedding processing is performed on the status attribute values included in the patient status structured text information to obtain a categorical attribute feature vector. The categorical attribute feature vector may represent the status attribute values of a categorical attribute. For example, the categorical attribute may be, but is not limited to, one of the following: skin type, gender, or whether the patient has a scar tissue.
[0032] As an example, the main control chip can perform word embedding processing on the state attribute value included in the structured text information of the patient's state using a preset first word embedding processing method to obtain a category attribute feature vector. For example, the first word embedding processing method can be a one-hot encoding method.
[0033] In the second step, the obtained feature vectors of each category attribute and the feature values of each numerical attribute are concatenated to obtain an initial first-state text feature vector.
[0034] As an example, the main control chip can perform horizontal splicing on the obtained category attribute feature vectors and numerical attribute feature values according to the order of the patient status structured text information sequence to obtain an initial first status text feature vector.
[0035] In the third step, for each piece of patient status unstructured text information in the above patient status unstructured text information sequence, word embedding processing is performed on the above patient status unstructured text information to obtain an initial second state text feature vector.
[0036] As an example, the main control chip can perform word embedding processing on the unstructured text information of the patient's status using a preset second word embedding processing method to obtain an initial second-state text feature vector. For example, the second word embedding processing method can be, but is not limited to, one of: a bidirectional encoder representation model and a general language model for conversational applications.
[0037] In some embodiments, the main control chip 103 may be configured to perform image feature extraction on the received patient skin lesion image to obtain an initial skin lesion region feature vector, wherein the initial skin lesion region feature vector may represent features of a local image of the skin lesion region.
[0038] In some optional implementations of some embodiments, the main control chip 103 may be configured to perform the following steps to perform image feature extraction processing on the received patient skin lesion image to obtain an initial skin lesion area feature vector:
[0039] In the first step, target detection is performed on the received patient skin lesion image to obtain skin lesion area bounding box information. The skin lesion area bounding box information may include the coordinates of the upper left corner of the bounding box, the bounding box width, and the bounding box height. The upper left corner coordinates may be the coordinates of the upper left corner of the skin lesion area bounding box in the image coordinate system. The skin lesion area bounding box may be a bounding box used to locate the skin lesion area. The bounding box width may be the width of the skin lesion area bounding box. The bounding box height may be the height of the skin lesion area bounding box.
[0040] As an example, the main control chip can perform target detection on a received patient skin lesion image using a preset target detection model to obtain bounding box information for the skin lesion region. The target detection model can be, but is not limited to, one of the following: a Faster R-CNN (Faster Region-based Convolutional Neural Network) model or a YOLO (You Only Look Once) model.
[0041] In practice, the main control chip can input the patient's skin lesion image into the target detection model. The convolutional layer of the target detection model first extracts features from the patient's skin lesion image. The target detection model then predicts the bounding box of the lesion area based on the extracted features. Non-maximum suppression is then used to filter out overlapping bounding boxes, retaining the bounding box with the highest confidence as the lesion area bounding box. This effectively removes interference from background information such as clothing, hair, and normal skin through target detection, ensuring more accurate feature extraction, further improving the accuracy of lesion classification, and reducing misdiagnosis by doctors.
[0042] In the second step, based on the skin lesion area bounding box information, the patient's skin lesion image is cropped to obtain a skin lesion area image. The skin lesion area image may be an image of the skin lesion area framed by the skin lesion area bounding box.
[0043] As an example, the main control chip can crop a local image of the lesion area from the patient's skin lesion image based on the upper left corner coordinates, bounding box width and bounding box height included in the lesion area bounding box information to obtain a skin lesion area image.
[0044] The third step is to perform normalization processing on the skin lesion area image to obtain a normalized skin lesion area image. The normalized skin lesion area image can be a skin lesion area image after size and color channel adjustment.
[0045] As an example, the main control chip can be configured to, first, scale the skin lesion area image according to preset image size information to obtain an updated skin lesion area image. The image size information can be the preset width and height information of the image. For example, the image size information can be: {"height": 224px, "width": 224px}. Then, in response to determining that the updated skin lesion area image is not in RGB format, perform color channel conversion on the updated skin lesion area image to obtain a normalized skin lesion area image. It should be noted that when the updated skin lesion area image is in RGB format, the updated skin lesion area image can be used as a normalized skin lesion area image.
[0046] The fourth step is to input the normalized skin lesion area image into a pre-trained image feature extraction model to obtain an initial skin lesion area feature vector. The image feature extraction model can be a ResNet18 network with the fully connected layer removed.
[0047] Optionally, the above-mentioned skin condition classification system based on multimodal feature fusion may further include a first model training server and a second model training server.
[0048] As an example, Figure 2Schematic diagrams of some other exemplary embodiments of the skin condition classification system based on multimodal feature fusion disclosed in the present invention are shown. Figure 2 As shown, the skin condition classification system based on multimodal feature fusion includes a patient condition text information input device 101, a skin image acquisition device 102, a main control chip 103, a classification result display 104, a first model training server 105, and a second model training server 106. The first model training server 105 and the second model training server 106 can both be communicatively connected to the main control chip 103. The first model training server can be used to train the skin condition classification model. The second model training server can be used to train the image feature extraction model.
[0049] Optionally, the second model training server 106 may be configured to perform the following steps:
[0050] The first step is to obtain a positive sample skin image information set, a negative sample skin image set and an initial image feature extraction model. Each positive sample skin image information in the positive sample skin image information set may include a positive sample skin image and a positive sample skin category label. The positive sample skin image may be a patient skin lesion image corresponding to the target skin lesion category. The target skin lesion category may be a category of skin damage to be identified. The target skin lesion category may be one of the following: pigmented nevus, nevus of Ota, freckles, chloasma, café au lait spots, brown-blue nevus, seborrheic keratosis, freckle-like nevus, pigmented hair epidermal nevus and melanosis, etc. Each negative sample skin image in the negative sample skin image set may be an image containing clothing, hair, normal skin and skin with other skin diseases, which is different from the positive sample skin image. The initial image feature extraction model may be a residual network including a fully connected layer.
[0051] As an example, the second model training server may obtain a positive sample skin image information set, a negative sample skin image set, and an initial image feature extraction model from a database.
[0052] The second step is to annotate each negative sample skin image in the negative sample skin image set to obtain a negative sample annotated skin image information set. Each negative sample annotated skin image in the negative sample annotated skin image information set may include a negative sample annotated image and a negative sample skin category label. The negative sample annotated image may be an image with a non-lesion area annotated. The non-lesion area may be a skin area without lesions corresponding to the target lesion category. The negative sample skin category label may be text used to characterize the skin category of the non-lesion area.
[0053] As an example, the non-lesional area can be an area of normal skin. The negative skin category label can be normal skin or other skin disease types. The main control chip manually labels each negative skin image in the negative skin image set. First, the non-lesional area in the negative skin image is labeled using a polygonal box or other image box, and the non-lesional area is labeled with a skin category. Then, the image with the labeled box is used as the negative sample labeled image, and the labeled skin category text is determined as the negative sample skin category label.
[0054] In a third step, for each positive sample skin image in the positive sample skin image information set, the positive sample skin image included in the positive sample skin image information is preprocessed to obtain a positive sample normalized skin lesion area image, and the positive sample skin category label and the positive sample normalized skin lesion area image included in the positive sample skin image information are determined as positive samples. The positive sample normalized skin lesion area image may be a normalized skin lesion area image corresponding to the positive sample skin image.
[0055] As an example, the main control chip can pre-process the positive sample skin image included in the positive sample skin image information through operations such as target detection, image cropping and normalization processing to obtain a positive sample normalized skin lesion area image.
[0056] In a fourth step, for each negative sample annotated skin image in the negative sample annotated skin image information set, the negative sample annotated image included in the negative sample annotated skin image information is preprocessed to obtain a negative sample normalized skin lesion area image, and the negative sample skin category label and the negative sample normalized skin lesion area image included in the negative sample annotated skin image information are determined as negative samples. The negative sample normalized skin lesion area image may be the normalized skin lesion area image corresponding to the negative sample annotated skin image.
[0057] As an example, the main control chip may pre-process the negative sample annotated image included in the negative sample annotated skin image information through operations such as image cropping and normalization processing to obtain a negative sample normalized skin lesion area image.
[0058] In the fifth step, the initial image feature extraction model is trained based on the determined positive samples and negative samples, and the fully connected layer in the trained initial image feature extraction model is deleted to generate an image feature extraction model.
[0059] As an example, the main control chip can train the initial image feature extraction model based on the determined positive samples and negative samples using a conventional training method. After the initial image feature extraction model is trained, the fully connected layer in the initial image feature extraction model is deleted, and the initial image feature extraction model that does not include the fully connected layer is determined as the image feature extraction model.
[0060] It should be noted that Figure 3 A data flow diagram of the skin image feature extraction step in some embodiments of the skin condition classification system based on multimodal feature fusion of the present disclosure is shown. Figure 3 Taking mole as an example, first, the original mole image is subjected to target detection, image cropping and normalization to obtain a normalized skin lesion area image. Then, the above-mentioned image feature extraction model is used to extract features from the normalized skin lesion area image to obtain a skin lesion area image feature map. In addition, Figure 3 Based on the steps shown above, the above image feature extraction model can also be used to perform global average pooling. Figure 3 The skin lesion area image feature map shown in is converted into an initial skin lesion area feature vector.
[0061] In some embodiments, the main control chip 103 can be configured to input the initial first-state text feature vector, the initial second-state text feature vector sequence, and the initial skin lesion area feature vector into a pre-trained skin state classification model to obtain a skin state classification result, and transmit the skin state classification result to the classification result display 104. The skin state classification model can be obtained by training an initial skin state classification model. The initial skin state classification model can include a feature dimension alignment module, a text fusion module, a text and image fusion module, a classification module, and a multimodal feature adaptive contrast enhancement module. The feature dimension alignment module can be used to map different feature vectors to a preset feature space through linear transformation to obtain vectors of the same dimension. The preset feature space can be a feature space pre-set for shared representation of images and text. The text fusion module can be used to fuse structured text features and unstructured text features. The text and image fusion module can be used to fuse text features and image features. The classification module can be composed of a neural network for predicting skin lesion categories. For example, the classification module can be a multi-layer perceptron or a CNN model. The multimodal feature adaptive contrast enhancement module can be used to introduce contrast loss for both intra-class consistency and inter-class variability to enhance the ability to discriminate patient characteristics. This module is only used during the model training phase and not during the model application and inference phase. The skin condition classification results can be the predicted probabilities corresponding to different skin lesion categories for a patient's skin. The classification result display can be a terminal with a display screen.
[0062] Optionally, the first model training server 105 may be configured to perform the following steps to train a skin condition classification model:
[0063] The first step is to obtain a patient sample set and an initial skin state classification model. Each patient sample in the patient sample set may include a sample patient identifier, a sample initial skin lesion area feature vector, a sample initial first state text feature vector, a sample initial second state text feature vector sequence, and a sample skin lesion category label. The sample patient identifier may be a unique identifier of the sample patient. The sample patient may be a patient whose corresponding patient data is used to train the model. The sample initial skin lesion area feature vector may be an initial skin lesion area feature vector corresponding to the sample patient's skin lesion area image. The sample initial first state text feature vector may be an initial first state text feature vector corresponding to the sample patient. The sample initial second state text feature vector sequence may be an initial second state text feature vector sequence corresponding to the sample patient. The sample skin lesion category label may represent the actual skin lesion category to which the sample patient's skin belongs.
[0064] As an example, the first model training server may obtain a patient sample set and an initial skin condition classification model from a database.
[0065] Continuing, in the process of solving the technical problems in the background technology, the present disclosure is also accompanied by the following technical problem 2: how to fuse image features and text features to improve the classification accuracy of skin conditions. In response to the above technical problem 2, the conventional solution is generally to set different fixed weights for features of different modalities. However, the above conventional solution has the following problems: since congenital skin pigmentation diseases such as brown-blue nevus are not closely related to the structured text information of the patient's skin, when fusing the skin lesion image and structured text, if a fixed weight distribution method is used to perform weighted fusion of the two features, the feature information will be redundant, resulting in overfitting, further reducing the difference between categories, and thus easily leading to lower classification accuracy of skin conditions. Therefore, in the face of the above technical problem 2, combined with the technical advantages of the solution development team itself in the field of deep learning, the present disclosure decided to adopt the following solution.
[0066] In the second step, at least one patient sample is selected from the patient sample set and the following training steps are performed:
[0067] In the first sub-step, the initial skin lesion region feature vector of each patient sample in the at least one patient sample is mapped to a preset feature space using a feature dimension alignment module included in the initial skin condition classification model to obtain at least one sample skin lesion region feature vector. The sample skin lesion region feature vector can be generated using the following formula:
[0068] I′=W 图像 I+b 图像 .
[0069] Where I represents the initial lesion area feature vector of the sample. 图像 Represents the weight matrix corresponding to the image features. b 图像 Represents the bias term corresponding to the image feature. I′ represents the feature vector of the sample lesion area. It should be noted that when the model starts training, W 图像 It is initialized using the Xavier initialization method, which can solve the problem of gradient disappearance or gradient explosion during model training.
[0070] In the second sub-step, the sample initial first-state text feature vector and the sample initial second-state text feature vector sequence included in each of the at least one patient sample are mapped to the feature space using the feature dimension alignment module included in the initial skin state classification model to obtain at least one sample first-state text feature vector and at least one sample second-state text feature vector sequence. The sample first-state text feature vector and the sample second-state text feature vector sequence can be generated using the following formula group:
[0071]
[0072] Where S represents the initial first state text feature vector of the sample. s Represents the weight matrix corresponding to the structured text features. b s Represents the bias term corresponding to the structured text feature. S′ represents the sample’s first state text feature vector. T represents the sample’s initial second state text feature vector. W t The weight matrix representing the unstructured text features. b t Represents the bias term. T′ represents the second state text feature vector of the sample. It should be noted that when the model starts training, W s and W t It is initialized by the Xavier initialization method.
[0073] In a third sub-step, the text fusion module included in the initial skin condition classification model determines, based on the sequence of the at least one sample first-state text feature vector and the at least one sample second-state text feature vector, a sample patient text feature vector corresponding to each patient sample in the at least one patient sample, thereby obtaining at least one sample patient text feature vector. The sample patient text feature vector in the at least one sample patient text feature vector may represent features of the text describing the skin lesion condition of the sample patient.
[0074] In some optional implementations of some embodiments, the first model training server 105 may be further configured to perform the following steps to generate at least one sample patient text feature vector through a text fusion module included in the initial skin condition classification model:
[0075] Step 1: Normalize the preset first learnable weight parameter and the second learnable weight parameter to obtain the first normalized weight parameter and the second normalized weight parameter. The first learnable weight parameter may be the initial weight parameter of the structured text feature. The second learnable weight parameter may be the initial weight parameter of the unstructured text feature. The first normalized weight parameter may be a weight parameter between 0 and 1 that is applicable to the structured text feature. The second normalized weight parameter may be a weight between 0 and 1 that is applicable to the unstructured text feature. It should be noted that the first learnable weight parameter and the second learnable weight parameter can be learned and updated through the back propagation of the model during the training phase, and the sum of the first normalized weight parameter and the second normalized weight parameter may be 1.
[0076] As an example, the first model training server may perform normalization processing on the first learnable weight parameter and the second learnable weight parameter through a Softmax function to obtain a first normalized weight parameter and a second normalized weight parameter.
[0077] Step 2: For each patient sample in the at least one patient sample, perform the following steps:
[0078] Sub-step 1: Perform feature aggregation on each sample second-state text feature vector in the sample second-state text feature vector sequence corresponding to the patient sample to obtain a sample second-state fused feature vector. The sample second-state fused feature vector may represent a fusion result of each unstructured text describing the patient's skin lesion status.
[0079] As an example, the first model training server may perform feature aggregation processing on each sample second-state text feature vector in the sample second-state text feature vector sequence corresponding to the patient sample using a preset feature aggregation processing method to obtain a sample second-state fused feature vector. The feature aggregation processing method may include, but is not limited to, at least one of the following: average pooling, maximum pooling, or an attention mechanism method.
[0080] Sub-step 2, based on the preset first learnable gating factor, the second learnable gating factor, the above-mentioned first normalization weight parameter and the above-mentioned second normalization weight parameter, the sample first state text feature vector and the sample second state fusion feature vector corresponding to the above-mentioned patient sample are subjected to feature weighted fusion processing to obtain the sample patient text feature vector. Among them, the above-mentioned first learnable gating factor can be a learnable gating factor for adjusting the first normalization weight parameter. The above-mentioned second learnable gating factor can be a learnable gating factor for adjusting the second normalization weight parameter. The sample patient text feature vector can be generated by the following formula:
[0081] F 文本 =g t α t ·T′+g s α s ·S′.
[0082] Among them, F 文本 g represents the sample patient text feature vector. s represents the first learnable gating factor. g t Represents the second learnable gating factor. α s Represents the first normalized weight parameter. α t Represents the second normalization weight parameter.
[0083] In a fourth sub-step, the text and image fusion module included in the initial skin condition classification model determines, based on the at least one sample patient text feature vector and the at least one sample skin lesion area feature vector, a sample multimodal fusion feature vector corresponding to each patient sample, thereby obtaining at least one sample multimodal fusion feature vector. The sample multimodal fusion feature vector in the at least one sample multimodal fusion feature vector may represent fusion features of the text and image describing the patient's skin lesion condition.
[0084] In some optional implementations of some embodiments, the first model training server 105 may be further configured to perform the following steps to generate at least one sample multimodal fusion feature vector through the text and image fusion module included in the initial skin condition classification model:
[0085] Step 1, normalize the preset third learnable weight parameter and the fourth learnable weight parameter to obtain a third normalized weight parameter and a fourth normalized weight parameter. The third learnable weight parameter may be an initial weight parameter for text features. The fourth learnable weight parameter may be an initial weight parameter for image features. The third normalized weight parameter may be a weight parameter between 0 and 1 that is applicable to text features. The fourth normalized weight parameter may be a weight between 0 and 1 that is applicable to image features. It should be noted that the third learnable weight parameter and the fourth learnable weight parameter can be learned and updated through back propagation of the model during the training phase, and the sum of the third normalized weight parameter and the fourth normalized weight parameter can be 1.
[0086] As an example, the first model training server may perform normalization processing on the third learnable weight parameter and the fourth learnable weight parameter through a Softmax function to obtain a third normalized weight parameter and a fourth normalized weight parameter.
[0087] Step 2: For each patient sample in the at least one patient sample, based on the preset third learnable gating factor, the fourth learnable gating factor, the third normalization weight parameter and the fourth normalization weight parameter, the sample patient text feature vector and the sample skin lesion area feature vector corresponding to the patient sample are subjected to feature weighted fusion processing to obtain a sample multimodal fusion feature vector. Among them, the third learnable gating factor can be a learnable gating factor for adjusting the third normalization weight parameter. The fourth learnable gating factor can be a learnable gating factor for adjusting the fourth normalization weight parameter. The sample patient text feature vector can be generated by the following formula:
[0088] F 多模态 =g 图像 α 图像 ·I′+g 文本 α 文本 ·F 文本 .
[0089] Among them, F 多模态 Represents the sample multimodal fusion feature vector. g 文本 Represents the third learnable gating factor. α 文本 Represents the third normalization weight parameter. g 图像 Represents the fourth learnable gating factor. α 图像 Represents the fourth normalization weight parameter.
[0090] The fifth sub-step is to determine the skin condition classification result corresponding to each patient sample in the at least one patient sample according to the at least one sample multimodal fusion feature vector through the classification module included in the initial skin condition classification model.
[0091] Continuing, in the process of solving the technical problems in the background technology and the above-mentioned technical problem 2, the present disclosure is also accompanied by the following technical problem 3: different types of skin diseases often have similar image features or description features. How to train the model based on limited samples to improve the classification performance of the model. In response to the above-mentioned technical problem 3, the conventional solution is generally: based on the existing positive sample images and positive sample texts, new sample images and sample texts are generated through image generation models and text generation models. However, the above-mentioned conventional solutions have the following problems: the images and texts generated by the neural network model depend on the performance of the neural network model, and are greatly different from the real environment, and the generation quality is also greatly different, which leads to poor classification performance of the trained skin condition classification model, and further, leads to low classification accuracy of the skin condition. Therefore, in the face of the above-mentioned technical problem 3, combined with the technical advantages of the solution development team itself in the field of deep learning, the present disclosure decided to adopt the following solution.
[0092] In a sixth sub-step, a multimodal feature adaptive contrast enhancement module included in the initial skin condition classification model is used to determine a sample feature distance information set based on the sample skin lesion region feature vector and the sample patient text feature vector corresponding to each of the at least one patient sample. The sample feature distance information in the sample feature distance information set may represent the degree of similarity or difference between the sample skin lesion region feature vector and the sample patient text feature vector.
[0093] In some optional implementations of some embodiments, the first model training server 105 may be further configured to perform the following steps to generate a sample feature distance information set using a multimodal feature adaptive contrast enhancement module included in the initial skin condition classification model:
[0094] Step 1: For each patient sample in the at least one patient sample, a preset positive sample label, a sample skin lesion area feature vector corresponding to the patient sample, and a sample patient text feature vector are determined as a positive training sample. The preset positive sample label may be a pre-set text representing the positive sample. For example, the preset positive sample label may be represented by "1."
[0095] Step 2: For each skin lesion category identifier in the preset skin lesion category identifier group, perform the following steps:
[0096] Sub-step 1: Based on the above-mentioned lesion category identifier, classify and process the skin lesion area feature vectors of each sample and the text feature vectors of each sample patient corresponding to the above-mentioned at least one patient sample to obtain a set of skin lesion area feature vectors of similar samples, a set of text feature vectors of similar sample patients, a set of skin lesion area feature vectors of heterogeneous samples, and a set of text feature vectors of heterogeneous sample patients. The skin lesion category identifier in the above-mentioned skin lesion category identifier group may be a unique identifier of a pre-set skin lesion category. The above-mentioned skin lesion area feature vector set of similar samples may be a set of skin lesion area feature vectors of each sample corresponding to the above-mentioned skin lesion category identifier. The above-mentioned text feature vector set of similar sample patients may be a set of text feature vectors of each sample patient corresponding to the above-mentioned skin lesion category identifier. The above-mentioned skin lesion area feature vector set of heterogeneous samples may be a set of skin lesion area feature vectors of each sample that does not correspond to the above-mentioned skin lesion category identifier. The text feature vector set of heterogeneous sample patients may be a set of text feature vectors of each sample patient that does not correspond to the above-mentioned skin lesion category identifier.
[0097] As an example, the first model training server is configured to, first, select the sample lesion area feature vectors that match the lesion category identifier from the sample lesion area feature vectors corresponding to the at least one patient sample as the same-class sample lesion area feature vectors, and obtain a set of similar-class sample lesion area feature vectors. Wherein, matching the lesion category identifier may mean that the sample lesion category label corresponding to the sample lesion area feature vector is the same as the lesion category identifier. Then, select the sample patient text feature vectors that match the lesion category identifier from the sample patient text feature vectors corresponding to the at least one patient sample as the same-class sample patient text feature vectors, and obtain a set of similar-class sample patient text feature vectors. Wherein, matching the lesion category identifier may mean that the sample lesion category label corresponding to the sample patient text feature vector is the same as the lesion category identifier. Afterwards, the sample lesion area feature vectors that do not match the lesion category identifier among the sample lesion area feature vectors corresponding to the at least one patient sample are determined as the different-class sample lesion area feature vector set. Finally, among the sample patient text feature vectors corresponding to the at least one patient sample, the sample patient text feature vectors that do not match the skin lesion category identifier are determined as a heterogeneous sample patient text feature vector set.
[0098] Sub-step 2: Generate a first negative training sample set based on the aforementioned similar sample skin lesion region feature vector set and the aforementioned heterogeneous sample patient text feature vector set. The first negative training sample in the aforementioned first negative training sample set may be a negative sample consisting of a similar sample skin lesion region feature vector and a heterogeneous sample patient text feature vector.
[0099] As an example, the first model training server may randomly select a vector element from the skin lesion area feature vector set of the same type of samples and the text feature vector set of the different type of samples of patients multiple times to form a first negative training sample.
[0100] Sub-step three: generating a second negative training sample set based on the aforementioned similar sample patient text feature vector set and the aforementioned heterogeneous sample skin lesion area feature vector set. The second negative training sample in the aforementioned second negative training sample set may be a negative sample consisting of a similar sample patient text feature vector and a heterogeneous sample skin lesion area feature vector.
[0101] As an example, the first model training server may randomly select a vector element from the patient text feature vector set of the same type sample and the skin lesion area feature vector set of the heterogeneous sample multiple times each time to form a first negative training sample.
[0102] Step 3: Generate a comparison sample set based on the determined positive training samples, the first negative training sample set, and the second negative training sample set. The comparison sample set may be a set of non-repeated comparison samples.
[0103] As an example, the first model training server first determines each positive training sample, each first negative training sample, and each second negative training sample as a comparison sample to be deduplicated, thereby obtaining a comparison sample set to be deduplicated. Then, the comparison sample set to be deduplicated is deduplicated to obtain a comparison sample set.
[0104] Step 4: Determine the sample feature distance information corresponding to each comparison sample in the comparison sample set to obtain a sample feature distance information set.
[0105] As an example, the first model training server may determine the sample feature distance information corresponding to each comparison sample using the following formula:
[0106]
[0107] Where, d(I′, F 文本 ) represents the distance between the text features and image features included in the comparison sample.
[0108] The seventh sub-step is to determine the total loss value based on the sample feature distance information set, the sample skin lesion category label corresponding to each patient sample in the at least one patient sample, and the skin condition classification result.
[0109] In some optional implementations of some embodiments, the first model training server 105 may be further configured to perform the following steps to determine the total loss value:
[0110] Step 1: Determine a classification loss value based on the skin lesion category label and skin condition classification result corresponding to each patient sample in the at least one patient sample using a preset classification loss function. The classification loss function may be:
[0111]
[0112] Among them, L cls represents the classification loss value. N represents the total number of patient samples in the at least one patient sample. C represents the total number of skin lesion categories. i represents the sample number. j represents the skin lesion category number. y j,i Represents the sample skin lesion category label corresponding to the j-th patient sample. It represents the predicted probability that the jth sample belongs to the i-th category.
[0113] Step 2: Determine the regularization loss value based on the first learnable gating factor, the second learnable gating factor, the third learnable gating factor, and the fourth learnable gating factor using a preset regularization loss function. The regularization loss function may be:
[0114] L reg =λ·(|g 图像 |+|g 文本 |+|g t |+|g s |).
[0115] Among them, L reg represents the regularization loss value. λ represents the hyperparameter.
[0116] The above-mentioned steps for generating the sample multimodal fusion feature vector and regularization loss value and their related contents, as an inventive point of an embodiment of the present disclosure, solve the above-mentioned technical problem No. 2, "How to fuse image features and text features to improve the classification accuracy of skin conditions". The reason why conventional solutions have low classification accuracy of skin conditions is that congenital skin pigmentation diseases such as brown-blue nevus are not closely related to the structured text information of the patient's skin. When performing feature fusion of skin lesion images and structured text, if a fixed weight distribution method is used to perform weighted fusion of the two features, the feature information will be redundant, resulting in overfitting, and further reducing the difference between categories. If the above-mentioned problems are solved, the effect of improving the accuracy of skin condition classification can be achieved. To achieve this effect, the present disclosure first maps the skin lesion image features, structured text features, and unstructured text features to the same feature space, thereby ensuring the dimensional consistency of image features and text features during the feature fusion process. Then, through a gating mechanism, the weights in the multimodal feature fusion process are dynamically adjusted. Finally, the gating factors are sparsified using a subsequent L1 regularization loss to ensure that features with little correlation to skin lesions have a weight of 0. This reduces redundant feature information, increases the diversity of skin lesion categories, and avoids overfitting. This improves the robustness and accuracy of the skin condition classification model, and ultimately, enhances skin condition classification accuracy.
[0117] Step 3: Determine the contrast loss value corresponding to the above-mentioned contrast sample set based on the above-mentioned sample feature distance information set through a preset contrast loss function. The above-mentioned contrast loss function can be:
[0118]
[0119] Among them, L contrastive Represents the contrast loss value. I′ i F represents the feature vector of the sample skin lesion area in the i-th comparison sample.文本,i y represents the sample patient text feature vector in the i-th comparison sample. M represents the total number of comparison samples in the comparison sample set. i Represents the indicator variable. When the i-th comparison sample is a positive training sample, y i is 1, when the i-th comparison sample is a negative training sample, y i is 0. margin is a set hyperparameter, which represents the minimum distance threshold between two feature vectors in negative training samples.
[0120] The above-mentioned contrast loss value generation step and its related contents, as an inventive point of an embodiment of the present disclosure, solve the technical problem three mentioned in the background technology, "how to train a model based on limited samples to improve the classification performance of the model". The reason why the conventional solution has poor classification performance and low skin state classification accuracy of the trained skin state classification model is that the images and texts generated by the neural network model depend on the performance of the neural network model, which are quite different from the real environment, and the generation quality difference is also quite large. Therefore, if the above-mentioned problem is solved, the effect of improving the accuracy of skin state classification can be achieved. In order to achieve this effect, the present disclosure first constructs a contrast sample set including positive training samples and negative training samples according to the skin lesion area feature vectors and patient text feature vectors corresponding to each patient sample through a multimodal feature adaptive contrast enhancement module during the training process of the skin state classification model. In this way, it can be ensured that each generated contrast sample is consistent with the image and text in the real environment. Then, for each contrast sample in the contrast sample set, the feature distance between the skin lesion image features and text features included in the contrast sample is determined. The contrastive loss function can then be used to increase the similarity between the lesion image features and text features included in the positive training samples, while decreasing the similarity between the lesion image features and text features included in the negative training samples. Finally, the contrastive loss value is integrated into the model's backpropagation process to continuously optimize the embedding distribution of the lesion image and text fusion features. This ensures that the image and text features corresponding to the same lesion category are similar, while those corresponding to different lesion categories are different. The fused feature representations are more separable between lesion categories. This provides a more discriminative input to the classification module, thereby improving the classification performance of the skin condition classification model. Furthermore, the classification accuracy of skin conditions can be improved.
[0121] Step 4: The sum of the classification loss value, the regularization loss value, and the contrast loss value is determined as the total loss value.
[0122] In an eighth sub-step, in response to determining that the total loss value is less than a preset loss threshold, the trained initial skin condition classification model is determined as the skin condition classification model.
[0123] Optionally, the first model training server 105 may be further configured to, in response to determining that the total loss value is greater than or equal to the preset loss threshold, adjust network parameters of the initial skin condition classification model, and form a patient sample set with unused patient samples, and perform the training step again using the adjusted initial skin condition classification model. The first model training server 105 may optimize the network parameters of the initial skin condition classification model based on the total loss value through backpropagation.
[0124] In some embodiments, the classification result display 104 may be configured to display the received skin condition classification result, thereby allowing the user to view the skin condition classification result and assisting the user in making a diagnosis.
[0125] As an example, Figure 4 The following table shows the experimental data of skin lesion classification performance of some embodiments of the skin condition classification system based on multimodal feature fusion disclosed in the present invention. Figure 4 The table lists different skin lesion types and the various performance metrics used to measure lesion classification. These lesion types include pigmented nevus, nevus of Ota, freckles, melasma, café-au-lait spots, chondroitinus, seborrheic keratosis, lentigo, pigmented piloepidermal nevus, melanosis, normal skin, and others. The performance metrics include accuracy, sensitivity, specificity, precision, and F1 value. The F1 value is the harmonic mean of accuracy and sensitivity. The values in the table represent the model's performance on the corresponding performance metrics when identifying the corresponding lesion type.
[0126] The above-mentioned embodiments of the present disclosure have the following beneficial effects: through the skin condition classification system based on multimodal feature fusion of some embodiments of the present disclosure, the classification accuracy of skin conditions can be improved. Specifically, the reason for the low accuracy of skin condition classification is that structured oral text usually has fewer dimensions and limited features, making it difficult to characterize the characteristics of various skin diseases. This makes it difficult for the fusion features of skin lesion images and structured text to accurately and comprehensively characterize the patient's skin lesion status, thereby resulting in low skin condition classification accuracy. Based on this, some embodiments of the present disclosure include a skin state classification system based on multimodal feature fusion, including: a patient state text information entry device, a skin image acquisition device, a main control chip, and a classification result display, wherein the patient state text information entry device is configured to enter a patient state structured text information sequence and a patient state unstructured text information sequence, and send the patient state structured text information sequence and the patient state unstructured text information sequence to the main control chip; the skin image acquisition device is configured to acquire a patient skin lesion image, and send the patient skin lesion image to the main control chip; the main control chip is configured to process the received patient state structured text information and The patient's unstructured text information sequence is subjected to text vectorization processing to obtain an initial first-state text feature vector and an initial second-state text feature vector sequence; the above-mentioned main control chip is configured to perform image feature extraction processing on the received patient's skin lesion image to obtain an initial skin lesion area feature vector; the above-mentioned main control chip is configured to input the above-mentioned initial first-state text feature vector, the above-mentioned initial second-state text feature vector sequence and the above-mentioned initial skin lesion area feature vector into a pre-trained skin state classification model to obtain a skin state classification result, and send the above-mentioned skin state classification result to the above-mentioned classification result display; the above-mentioned classification result display is configured to display the received skin state classification result. Therefore, in some embodiments of the present disclosure, the skin state classification system based on multimodal feature fusion can improve the feature dimension of the skin lesion feature by performing multimodal feature fusion on the patient's skin lesion image, the structured text corresponding to the skin state and the unstructured text through the main control chip, so as to more accurately and comprehensively characterize various types of skin diseases, thereby improving the classification accuracy of the skin state. Furthermore, the skin condition classification results can be displayed through the classification result display, thereby assisting doctors in diagnosing skin diseases and improving the efficiency of skin disease diagnosis.
[0127] The above description is only an illustration of some preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.
Claims
1. A skin condition classification system based on multimodal feature fusion, the skin condition classification system based on multimodal feature fusion comprising: Patient status text information input device, skin image acquisition device, main control chip, first model training server and classification result display, wherein, The patient status text information input device is configured to input a patient status structured text information sequence and a patient status unstructured text information sequence, and send the patient status structured text information sequence and the patient status unstructured text information sequence to the main control chip; The skin image acquisition device is configured to acquire a patient's skin lesion image and send the patient's skin lesion image to the main control chip; The main control chip is configured to perform text vectorization processing on the received patient status structured text information and patient status unstructured text information sequence to obtain an initial first state text feature vector and an initial second state text feature vector sequence; The main control chip is configured to perform image feature extraction processing on the received patient skin lesion image to obtain an initial skin lesion area feature vector, wherein the image feature extraction processing on the received patient skin lesion image to obtain the initial skin lesion area feature vector includes: Performing target detection on the received patient skin lesion image to obtain skin lesion area bounding box information; performing image cropping processing on the patient's skin lesion image based on the skin lesion area bounding box information to obtain a skin lesion area image; performing normalization processing on the skin lesion area image to obtain a normalized skin lesion area image; Inputting the normalized skin lesion area image into a pre-trained image feature extraction model to obtain an initial skin lesion area feature vector, wherein the initial skin lesion area feature vector represents the features of the local image of the skin lesion area; The main control chip is configured to input the initial first state text feature vector, the initial second state text feature vector sequence, and the initial skin lesion area feature vector into a pre-trained skin state classification model to obtain a skin state classification result, and send the skin state classification result to the classification result display; The classification result display is configured to display the received skin condition classification result; The first model training server is configured to: Obtaining a patient sample set and an initial skin state classification model, wherein each patient sample in the patient sample set includes a sample patient identifier, a sample initial skin lesion area feature vector, a sample initial first state text feature vector, a sample initial second state text feature vector sequence, and a sample skin lesion category label, and the initial skin state classification model includes a feature dimension alignment module, a text fusion module, a text and image fusion module, a classification module, and a multimodal feature adaptive contrast enhancement module; Select at least one patient sample from the patient sample set and perform the following training steps: Mapping the sample initial skin lesion area feature vector included in each patient sample of the at least one patient sample to a preset feature space through a feature dimension alignment module included in the initial skin state classification model to obtain at least one sample skin lesion area feature vector; Mapping the sample initial first-state text feature vector and the sample initial second-state text feature vector sequence included in each of the at least one patient sample to the feature space respectively through a feature dimension alignment module included in the initial skin state classification model, thereby obtaining at least one sample first-state text feature vector and at least one sample second-state text feature vector sequence; Determining, by a text fusion module included in the initial skin state classification model, a sample patient text feature vector corresponding to each patient sample in the at least one patient sample based on the at least one sample first state text feature vector and the at least one sample second state text feature vector sequence, thereby obtaining at least one sample patient text feature vector; Determining, by a text and image fusion module included in the initial skin condition classification model, a sample multimodal fusion feature vector corresponding to each patient sample in the at least one patient sample based on the at least one sample patient text feature vector and the at least one sample skin lesion area feature vector, thereby obtaining at least one sample multimodal fusion feature vector; Determining, by a classification module included in the initial skin condition classification model, a skin condition classification result corresponding to each patient sample in the at least one patient sample according to the at least one sample multimodal fusion feature vector; Determining a sample feature distance information set based on a sample skin lesion area feature vector and a sample patient text feature vector corresponding to each patient sample in the at least one patient sample by a multimodal feature adaptive contrast enhancement module included in the initial skin state classification model, wherein determining the sample feature distance information set based on the sample skin lesion area feature vector and the sample patient text feature vector corresponding to each patient sample in the at least one patient sample includes: For each patient sample in the at least one patient sample, determining a preset positive sample label, a sample skin lesion area feature vector corresponding to the patient sample, and a sample patient text feature vector as a positive training sample; For each lesion category identifier in the preset lesion category identifier group, perform the following steps: Classify and process each sample skin lesion region feature vector and each sample patient text feature vector corresponding to the at least one patient sample according to the skin lesion category identifier to obtain a set of similar sample skin lesion region feature vectors, a set of similar sample patient text feature vectors, a set of heterogeneous sample skin lesion region feature vectors, and a set of heterogeneous sample patient text feature vectors; Generate a first negative training sample set based on the skin lesion area feature vector set of the same type of samples and the text feature vector set of the different type of sample patients; generating a second negative training sample set based on the patient text feature vector set of the same type of samples and the skin lesion area feature vector set of the different type of samples; Generate a comparison sample set based on the determined positive training samples, the first negative training sample set, and the second negative training sample set; Determine the sample feature distance information corresponding to each comparison sample in the comparison sample set to obtain a sample feature distance information set; determining a total loss value according to the sample feature distance information set, the sample skin lesion category label corresponding to each patient sample in the at least one patient sample, and the skin condition classification result; In response to determining that the total loss value is less than a preset loss threshold, the trained initial skin condition classification model is determined as the skin condition classification model.
2. The skin condition classification system based on multimodal feature fusion according to claim 1, wherein: Each patient status structured text information in the patient status structured text information sequence includes a status attribute value; and the main control chip is further configured to: For each piece of patient status structured text information in the patient status structured text information sequence, perform the following steps: In response to determining that the patient status structured text information satisfies a preset attribute condition, normalizing the status attribute value included in the patient status structured text information to obtain a numerical attribute feature value; In response to determining that the patient status structured text information does not meet the preset attribute condition, performing word embedding processing on the status attribute value included in the patient status structured text information to obtain a category attribute feature vector; The obtained feature vectors of each category attribute and the feature values of each numerical attribute are concatenated to obtain an initial first-state text feature vector; For each piece of patient status unstructured text information in the patient status unstructured text information sequence, word embedding processing is performed on the patient status unstructured text information to obtain an initial second state text feature vector.
3. The skin condition classification system based on multimodal feature fusion according to claim 1, wherein: The first model training server is further configured to: In response to determining that the total loss value is greater than or equal to the preset loss threshold, the network parameters of the initial skin condition classification model are adjusted, and the unused patient samples are combined into a patient sample set, and the training step is performed again using the adjusted initial skin condition classification model.
4. The skin condition classification system based on multimodal feature fusion according to claim 1, wherein: The skin condition classification system based on multimodal feature fusion further includes a second model training server; and the second model training server is configured to: Obtaining a positive sample skin image information set, a negative sample skin image set, and an initial image feature extraction model, wherein each positive sample skin image information in the positive sample skin image information set includes a positive sample skin image and a positive sample skin category label, and the initial image feature extraction model is a residual network including a fully connected layer; Performing labeling processing on each negative sample skin image information in the negative sample skin image set to obtain a negative sample labeled skin image information set, wherein each negative sample labeled skin image information in the negative sample labeled skin image information set includes a negative sample labeled image and a negative sample skin category label; For each positive sample skin image information in the positive sample skin image information set, preprocessing the positive sample skin image included in the positive sample skin image information to obtain a positive sample normalized skin lesion area image, and determining the positive sample skin category label included in the positive sample skin image information and the positive sample normalized skin lesion area image as a positive sample; For each negative sample annotated skin image information in the negative sample annotated skin image information set, preprocessing the negative sample annotated image included in the negative sample annotated skin image information to obtain a negative sample normalized skin lesion area image, and determining the negative sample skin category label included in the negative sample annotated skin image information and the negative sample normalized skin lesion area image as a negative sample; Based on the determined positive samples and negative samples, the initial image feature extraction model is trained, and the fully connected layer in the trained initial image feature extraction model is deleted to generate an image feature extraction model.
Citation Information
Patent Citations
Image classification model processing method
CN115222982A