Polymorphic model for distinguishing vulva sclerotic moss and vulva chronic simple moss
By fusing medical history information with image features and using Transformer model for training, the problem of difficulty in accurately distinguishing between vulvar sclerosis and vulvar chronic lichen simplex in the prior art is solved, achieving higher diagnostic accuracy and lower risk of misdiagnosis.
Patent Information
- Application Number
- CN202510114309.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to accurately distinguish between lichen sclerosis vulva from lichen chronic lichen simple vulva, resulting in misdiagnosis or misdiagnosis, making it difficult to achieve early accurate diagnosis and early treatment.
The multimodal fusion model is used to fuse the patient's medical history information with the image features of the lesion area. The image features are extracted through the ResNet50 model, and the quantized medical history information is spliced after LayerNormalization operation. Finally, the Transformer model is used for training to achieve automatic classification and risk assessment of VLS and VLSC.
It significantly improves the accuracy of distinguishing between VLS and VLSC, reduces the risk of misdiagnosis and misdiagnosis, provides more objective diagnostic results, and improves the accuracy and credibility of clinical diagnosis.
Smart Images

Figure CN120048446A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer-aided diagnosis and medical image processing, and particularly relates to a multimodal model for distinguishing vulvar lichen sclerosus and vulvar lichen simplex chronicus. Background Art
[0002] Both vulvar lichen sclerosus (VLS) and vulvar lichen simplex chronicus (VLSC) are difficult gynecological diseases. Patients suffer from severe long-term vulvar itching, hypopigmentation, fissures, and even atrophy, adhesions, difficulties in sexual life and urination, which seriously reduce the quality of life and may even lead to malignancy. There are significant differences in pathological characteristics, progression risks, and treatment methods between VLS and VLSC. In particular, VLS has a certain risk of malignancy and requires long-term monitoring. Currently, the distinction between VLS and VLSC relies on the preliminary judgment of clinicians, and a final diagnosis is made after subsequent pathological examinations of suspected VLSC cases.
[0003] However, due to the high similarity of these lesions in image features, especially in terms of the shape, color, texture of the erythema and leukoplakia that appear, as well as the normal folds and pathological fissures of the local skin, these difficulties increase the technical difficulty of preliminary clinical judgment. On the other hand, the patient's medical history information (such as itching time, past medical history, drug use, etc.) also has important diagnostic value for distinguishing these two lesions. Therefore, the current diagnosis generally adopts the method of outpatient visual inspection combined with interrogation. Since the affected area images and medical history information are two different modalities of data and there is no unified quantification method, clinicians need to rely on experience to make judgments based on the above two aspects. This brings great uncertainty to the diagnosis work, easily leads to misdiagnosis or missed diagnosis, and it is difficult to achieve early accurate diagnosis and early treatment.
[0004] Although current artificial intelligence technologies, especially in the fields of image classification, object detection, and even image segmentation, have demonstrated powerful performance, due to the similarity of the visual features of the two diseases, the field of using artificial intelligence technology to distinguish between VLS and VLSC diseases is still blank.
[0005] To solve the above problems, the present invention proposes a multi-modal fusion model, which improves the discrimination accuracy of VLS and VLSC by fusing patient medical history information and lesion image features. The model combines medical history information (such as itching time, past medical history, drug use, etc.) and image features to enhance the expressive ability of multi-modal data and effectively make up for the deficiencies brought by pure image analysis. Specifically, the present invention first quantifies the medical history information to generate a standardized medical history information vector; then, extracts features from the lesion image through the ResNet50 model, splices the obtained image feature vector with the medical history information after performing LayerNormalization operation, and forms an input vector for multi-modal fusion; finally, uses the Transformer model to train the fusion information to achieve automatic classification and risk assessment of VLS and VLSC.
[0006] The model design of the present invention can process multi-modal data in the same framework, utilize the attention mechanism to capture the implicit correlation relationship between medical history information and image features, thereby significantly improving the specificity and accuracy of classification and reducing the risk of misdiagnosis and missed diagnosis. This automated diagnostic system based on multi-modal fusion can not only provide more objective diagnostic results, but also greatly improve the diagnostic accuracy and credibility of clinicians, bring safer and more effective early diagnostic services to patients, and break through the diagnostic bottleneck of vulvar leukoplakia, a difficult gynecological disease, which has existed for a long time. Summary of the Invention
[0007] Object of the Invention:
[0008] The present invention aims to propose a multi-modal classification model by fusing the medical history information of patients and the image features of the lesion area, enabling the model to evaluate the target object from a more comprehensive perspective, solving the problem of the deficiency in the ability of traditional single-modal image recognition to distinguish VLS and VLSC, and providing a reliable auxiliary diagnostic tool for clinical practice.
[0009] Technical Solution:
[0010] To achieve the above object, the present invention adopts the following technical solutions: A polymorphic model for distinguishing vulvar lichen sclerosus and vulvar chronic simple lichen mainly includes a medical history information record (1), an RGB image of the affected area (2), a quantization mapping module (3), a ResNet network (4), a medical history information vector (5), a feature map sequence (6), a feature map vector (7), a layer normalization parameter of the medical history information vector (8), a normalized medical history information vector (9), a normalization parameter of the feature map vector (10), a normalized feature map vector (11), a vector splicing module (12), a category information vector (13), a fusion information vector group (14), a Transformer module (15), a category information encoding (16), a multi-layer perceptron (MLP) module (17), and an inference result vector (18).
[0011] The implementation of a polymorphic model for distinguishing vulvar lichen sclerosus and vulvar chronic simple lichen is divided into two parts: model training and model inference:
[0012] A polymorphic model for distinguishing vulvar lichen sclerosus and vulvar chronic simple lichen, according to its structural design, adopts a two-stage implementation plan for its model training part, training ResNet and Transformer respectively. The specific plan is as follows:
[0013] Step 1, for individual medical record cases, collect their vulvar pictures as picture samples and collect their medical history information as medical history information samples. Collect picture samples and medical history information samples of a large number of cases to make a data set.
[0014] Step 2, divide the data set. Use healthy cases without VLS and VLSC as the negative sample data set, and use cases with VLS and VLSC as the positive sample data set.
[0015] Step 3, as a preferred solution of the present invention, perform random rotation, inversion, cropping, and mosaic operations on the picture samples in the data set to implement enhancement.
[0016] Step 4, use the above data set to train the ResNet network. After the model reaches a certain accuracy, fix the ResNet parameters.
[0017] As a preferred solution of the present invention, this embodiment adopts a cross-entropy loss function during the ResNet training process,
[0018]
[0019] where a i is the sample label, and a i =1 indicates that the i-th case is a VLS or VLSC case, and a i= 0 indicates that the i-th case is a non-VLS or VLSC case; indicates that ResNet's prediction for the i-th case is a VLS or VLSC case, indicates that ResNet's prediction for the i-th case is a non-VLS or VLSC case; N is the total number of samples. As a preferred solution of the present invention, the learning rate L in this embodiment R = 1×10 -4 , epoch = 300.
[0020] Step 5, use the trained ResNet to infer the dataset, and record the inference results of each medical record case and the feature maps output by ResNet. Record the cases with inference results of TP (True Positive) and TN (True Negative), and take the feature maps and medical history information of this medical record case as a sample in the training dataset of Transformer.
[0021] Step 6, in the training dataset of Transformer obtained in Step 5, take one case as a data sample. For the feature maps in this data sample, perform a flatten operation on the feature maps of each channel to obtain the feature map vector of this channel.
[0022] Step 7, perform a layer normalization operation on the feature maps of each channel to obtain the normalized feature map vectors of each channel.
[0023] Step 8, perform a layer normalization operation (Layer Normal-ization) on the medical history information vector to obtain the normalized medical history information vector.
[0024] Step 9, for the same case sample, append the normalized medical history information vector behind the normalized feature map vectors of each channel to form a fused information vector.
[0025] Step 10, combine the fused information vector with the category information vector to form the input vector of this case at the Transformer end.
[0026] Step 11, take the output vector corresponding to the category information at the Transformer decoder end and input it into the multi-layer perceptron (MLP) end for inference.
[0027] Step 12, as a preferred solution of the present invention, this embodiment adopts a cross-entropy loss function during the training of Transformer,
[0028]
[0029] Among them, y i is the sample label, and y i = 1 indicates that the i-th case is VLS, and y i = 0 indicates that the i-th case is VLSC; indicates that the prediction of the Transformer for the i-th case is VLS, indicates that the prediction of the Transformer for the i-th case is VLSC; N is the total number of samples. As a preferred solution of the present invention, the learning rate L r = 1×10 -4 , and epoch = 500.
[0030] In step 13, set the training epoch value, and repeat steps 6 to 11 to implement the training of the Transformer and the MLP.
[0031] A polymorphic model for distinguishing vulvar lichen sclerosus and vulvar chronic simple lichen. According to its structural design, the implementation scheme of the model inference part is as follows:
[0032] Step 1: Input data preprocessing. Obtain the vulvar area picture of the individual to be detected, and through preprocessing operations (such as resizing, removing the reflective area, etc.) to ensure that the picture quality meets the input requirements of the ResNet network. At the same time, collect the medical history information of this individual and convert it into a medical history information vector for combination with the image features and input into the Transformer.
[0033] Step 2: Input the preprocessed picture into the trained ResNet network to extract multi-channel feature maps.
[0034] Step 3: Perform a flatten operation on the feature maps of each channel output by ResNet to generate feature map vectors, and perform Layer Normalization processing on these feature map vectors to standardize the numerical range of the features and make it convenient for subsequent fusion.
[0035] Step 4: Medical history information processing. Perform Layer Normalization operation on the medical history information vector to make it consistent with the image feature vector numerically. The standardized medical history information vector can be concatenated with the image features to form a multi-modal fusion input.
[0036] Step 5: Multi-modal fusion. Concatenate the normalized medical history information vector with the normalized feature map vectors of each channel to form a multi-modal fusion information vector. This vector contains image features and medical history information, and can provide more comprehensive input information for the Transformer.
[0037] Step 6: Transformer Inference. Combine the multi-modal fusion information vector and the category information vector to form a complete Transformer input token. Through the self-attention mechanism of Transformer, perform deep learning on the interaction relationship between the medical history information and the image features to capture the complex features of the lesion.
[0038] Step 7: Classification Output. The category information vector output by the Transformer decoder is used as the final feature vector and input into a multi-layer perceptron (MLP). Calculate the probability values belonging to VLS or VLSC through its internal softmax layer. Determine the lesion category (VLS or VLSC) of this individual based on the probability values and give the final inference result.
[0039] This invention patent provides a method with the following beneficial effects:
[0040] 1. Innovative application of multi-modal fusion technology: Compared with traditional deep learning methods in the field of computer vision, such as models like YOLO, ResNet, and Fast R-CNN, this invention innovatively adopts a fusion method of medical history information and image feature information to construct a multi-modal detection model, which unifies and quantifies the empirical information that could only be obtained by clinicians based on physical examinations and past medical histories in the past. Furthermore, the above two key preliminary diagnostic bases of different modalities can be applied to the relatively mature field of deep learning.
[0041] 2. Higher-dimensional input data and higher classification accuracy: Compared with existing typical vision detection methods based on the Transformer structure, such as the ViT model and the DETR model, the polymorphic model for distinguishing vulvar lichen sclerosus and vulvar chronic simple lichen in this invention integrates medical history information and image information. The input data has higher-dimensional relevant information and higher classification and detection accuracy.
[0042] 3. Optimized feature extraction and multi-modal information fusion process: Compared with existing end-to-end multi-modal detection models, such as ViLT, this invention adopts a method of initially extracting image features using a convolutional neural network, which significantly reduces the number of self-attention in the ViLT structure, thereby reducing the computing power load required for training. As a preferred technical solution of this invention patent, the number of parameters of ResNet50 adopted in this invention is approximately 23.5M, and the number of parameters of Transformer is approximately 86M, which is in the same order of magnitude as the number of parameters of the lightweight YOLOv8x model, 68M. This enables the model proposed in this invention to be conveniently trained and deployed on medium and low computing power platforms, facilitating the promotion and use of this patent.
[0043] 4. Interpretability in Clinical Applications: The multi-modal fusion model of the present invention integrates medical history information and image features, making the classification results output by the model more interpretable clinically. Doctors can understand the decision-making basis of the model based on the weights of medical history information and the visual analysis of image features, which lays a foundation for the clinical promotion and application of the model and enhances the practical value of the model in the medical field. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 is the model structure diagram described in the present application, Figure 2 is the internal structure diagram of the ResNet network described in the present application, Figure 3 is the internal structure diagram of the Transformer described in the present application.
[0045] Figure 1 is the model structure diagram described in the present application, where: 1. Medical history information record; 2. RGB image of the affected area; 3. Quantization mapping module; 4. ResNet network; 5. Medical history information vector; 6. Feature map sequence; 7. Feature map vector; 8. Layer normalization parameter of the medical history information vector; 9. Normalized medical history information vector; 10. Feature map vector normalization parameter; 11. Normalized feature map vector; 12. Vector concatenation module; 13. Class information vector; 14. Fusion information vector group; 15. Transformer module; 16. Class information encoding; 17. Multi-layer perceptron (MLP) module; 18. And inference result vector (18).
[0046] Figure 2 is the internal structure diagram of the ResNet network in the model structure diagram described in the present application. Among them, BTNK1 is the bottleneck (BottleNeck) structure, where C is the number of input channels, W is the width and height of the input image, C1 is the number of output channels, and S is the stride parameter; BTNK2 is a bottleneck (BottleNeck) structure different from BTNK1, where C is the number of channels and W is the width and height of the input image. The ResNet network in the model structure diagram outputs a feature map sequence with a dimension of 256×32×32 in this embodiment.
[0047] Figure 3 is the internal structure diagram of the Transformer module in the model structure diagram described in the present application. Figure 3 On the left is the Encoder part, and its input is 256 tokens, namely p 0 , p 1 ,..., p 256 ( Figure 1In (14). The dimension of each token is 1024×1 (one output channel of ResNet is a feature map of 32×32, which becomes a vector of 1024×1 after flattening, and there are 256 output channels in total); L×6 in the upper left corner of the Encoder part indicates that the Encoder part is composed of 6 Encoder layers with the same structure in series. During inference, the Encoder part provides two parameter matrices K and V to the Decoder. Figure 3 On the right is the Decoder part, and in this embodiment its output is q 0 , q 1 ,..., q 256 , among which the meaningful one is the class information encoding q 0 ( Figure 1 in (16), q 0 is output to the MLP module ( Figure 1 in (17)) for inference of the prediction result. Specific implementation manners
[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments (including the selection of models, the selection of parameters, and the selection of medical history information) are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those skilled in the art without creative efforts belong to the protection scope of the present invention.
[0049] The working principle and data flow of the present invention in this embodiment are as follows:
[0050] Step 1: Collect relevant medical history information (1) of the patient through clinical interviews, including itching time, past medical history, drug use situation, etc. For the convenience of subsequent processing, different states of each medical history feature are quantified into values in the interval [0,1] to form a medical history information vector. As a preferred technical solution of the present invention, Table 1 provides the medical history information quantification scheme adopted in the implementation cases of the present invention.
[0051] Table 1 Quantification table of past medical history information
[0052]
[0053]
[0054] According to the interview situation, record the values of the patient's information in Table 1, and perform sample screening to discard samples with illegal data and missing information, and finally form a medical history information sample.
[0055] Step 2: According to the records in the medical history information statistical table (Table 1), vectorize the medical history information to generate a medical history information vector, x = [x 1 , x 2 ,..., x n (5), In this embodiment, n = 15, and x k ∈[0, 1] is the quantization value of the medical history information corresponding to the serial number k in Table 1, where k = 1, 2,..., n.
[0056] Step 3: Use a color camera to take an image (2) of the patient's lesion area, and screen the image to remove blurred images, images with a large amount of obvious reflection, and other images with poor shooting quality, and perform a resize operation of uniform size to ensure the quality and size specification consistency of the images. As a preferred solution, in this example, the size of the modified image is set to W pixels in width and H pixels in height.
[0057] Step 4: Input the 3×W×H - dimensional image obtained in the previous step into the ResNet network (4) to obtain a feature map (6) with c channels, w pixels in width, and h pixels in height.
[0058] Step 5: For the c×w×h - dimensional feature map output by the ResNet network, let m = w×h, then the vector formed after flattening the feature map of the i - th channel is denoted as f i , f i is the feature map vector of the i - th channel. The j - th element of the feature map vector f i is denoted as f i,j . The feature map vectors of c channels form the feature vector group [f i (7), where i = 1, 2,..., c and j = 1, 2,..., m.
[0059] Step 6: For the medical history information vector x, calculate its mean μ x j and variance σ x j (8). Perform a LayerNormalization operation on x k to obtain a vector x' = [x' k (9), where k = 1, 2,..., n:
[0060]
[0061] Step 7: For the feature map vector f i , calculate the mean μ i of the information f fi Sum of variances σ f i (10). For the element f i,j , perform the Layer Normalization operation, and the feature map vector of the i-th channel is obtained as f' i = [f' i,j (11), j = 1, 2,..., m:
[0062]
[0063] Step 8: Concatenate the feature map vector f' i with the medical history information vector x to obtain the fused information vector p after embedding the medical history information i = [{f' i}, {x'}](12), i = 1, 2,..., c.
[0064] Step 9: Let the category information token be p 0 (13), p 0 is a learnable parameter. As a preferred technical solution of this invention patent, p 0 is initialized as a zero vector.
[0065] Step 10: The category information vector and the fused information vector are concatenated and fused again to form the input vector p, that is, p = [p 0 , p 1 ,..., p c (14). p is used as the input token of the next-stage Transformer module (15).
[0066] Step 11: Take the vector corresponding to the output end of the Transformer network as the category information encoding q 0 (16), input it into the MLP module (17), and finally obtain the inference result of whether the input sample belongs to VLSC or VLS through the softmax layer (18).
[0067] The above are only the preferred embodiments of this invention patent and are not used to limit this invention patent. Although the technical solutions described in the foregoing embodiments have been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this invention patent shall be included within the protection scope of this invention patent.
Claims
1. A multimodal model for distinguishing vulvar lichen sclerosus from vulvar chronic simple lichen sclerosus, characterized in that: It includes and uses medical history information record (1), RGB image of affected area (2), quantization mapping module (3), ResNet network (4), medical history information vector (5), feature map sequence (6), feature map vector (7), medical history information vector layer normalization parameter (8), normalized medical history information vector (9), feature map vector normalization parameter (10), normalized feature map vector (11), vector concatenation module (12), category information vector (13), fusion information vector group (14), Transformer module (15), category information encoding (16), multi-layer perceptron (MLP) module (17) and inference result vector (18).
2. A multimodal model for distinguishing vulvar lichen sclerosus from vulvar chronic simplex lichen according to claim 1, characterized in that: A feature information fusion method was designed that combines the patient's medical history information (such as itching duration, previous medical history, medication use, etc.) with the image information of the lesion area. A deep algorithm model based on the attention mechanism was trained on the fused information to better distinguish vulvar lichen sclerosus (VLS) from vulvar simplex chronicus (VLSC).
3. A multimodal model for distinguishing vulvar lichen sclerosus from vulvar chronic simplex lichen according to claim 1, characterized in that: A medical history information quantification table is designed, and the information vector and value range are designed and made consistent in dimension for the medical history information vectors of different cases, as a medical history information vector modality in a multimodal model for distinguishing between vulvar lichen sclerosus and vulvar chronic simple lichen sclerosus as described in claim 1.
4. A multimodal model for distinguishing vulvar lichen sclerosus from vulvar chronic simplex lichen according to claim 1, characterized in that: ResNet (residual network) is used to extract image features for the affected area image, and the extracted multi-channel feature map is used as a feature map mode in a multimodal model for distinguishing vulvar lichen sclerosus from vulvar chronic simple lichen as described in claim 1.
5. A multimodal model for distinguishing vulvar lichen sclerosus from vulvar chronic simplex lichen according to claim 1, characterized in that: Each medical record sample corresponds to a medical history information vector; each medical record sample corresponds to a multi-channel feature map extracted by the ResNet.
6. A multimodal model for distinguishing vulvar lichen sclerosus from vulvar chronic simplex lichen according to claim 1, characterized in that: The multi-channel feature map is flattened according to the feature map on each channel to form a feature map vector corresponding to each channel; the feature map vectors all have the same width and height.
7. A multimodal model for distinguishing vulvar lichen sclerosus from vulvar chronic simplex lichen according to claim 1, characterized in that: For a medical record sample, a layer normalization operation is performed on the medical history information vector; a layer normalization operation is also performed on the feature map on each channel in the multi-channel feature map of the medical record sample; this design smoothes out the information disturbance caused by the differences between the medical history information vector modality and the feature map modality, and between different feature maps in different numerical ranges, while making the data of two different modalities tend to the same distribution, while maintaining their respective information expressions; this design is conducive to promoting the convergence of the loss function and accelerating the training process of the model.
8. A multimodal model for distinguishing vulvar lichen sclerosus from vulvar chronic simplex lichen according to claim 1, characterized in that: For a medical record sample, the medical history information vector after layer normalization is concatenated with the normalized feature map vector on each channel. This operation realizes the feature fusion of the medical history information vector modality and the feature map modality. Because the medical history information is attached to the feature map vector of each channel, the amount of information in the medical history information is enhanced, which is beneficial for the model to extract deep information from the medical history information modality other than the feature map modality.
9. A multimodal model for distinguishing vulvar lichen sclerosus from vulvar lichen simplex chronicus according to claim 1, characterized in that: A two-stage model training strategy is adopted; in the first stage of the two-stage model training strategy, the ResNet network is trained with "containing VLS or containing VLSC" and "not containing VLS and VLSC" as two categories, and a certain accuracy is achieved; in the second stage, the "containing VLS or containing VLSC" samples predicted correctly in the first stage training process and the medical history information of their corresponding medical records are used as multimodal data sets to train the Transformer model and a certain accuracy is achieved, thereby realizing multimodal data fusion and recognition; in the first stage of the two-stage model training strategy, the ResNet network is guaranteed to effectively extract features of the image modality; in the second stage of the two-stage model training strategy, the training of multimodal data samples is realized.