Acne Grading Method, System, Device and Medium Based on Multimodal Fusion

By constructing a multimodal fusion acne grading model, using cross-attention mechanism and perceptron technology, the problem of poor acne grading effect on facial acne is solved, and a more accurate and explainable acne severity grading is achieved.

CN119600398BActive Publication Date: 2025-07-08SICHUAN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411608634.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2025-07-08
Estimated Expiration
2044-11-12

AI Technical Summary

Technical Problem

In the prior art, the severity classification effect of facial acne is poor, and the existing methods fail to effectively integrate multimodal characteristics, resulting in insufficient perception of acne characteristics and affecting the accuracy of grading.

Method used

A acne grading model based on multimodal fusion is constructed, including VGG feature extraction module, multimodal feature fusion module, acne feature perception module and comprehensive grading module. The feature fusion and acne feature perception are enhanced through the cross attention mechanism, and acne feature perception is used to extract acne features and perform comprehensive grading.

Benefits of technology

It improves the grading accuracy of facial acne severity, enhances attention to acne characteristics, realizes comprehensive diagnostic grading, and has good generalization and interpretability, reducing data collection costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119600398B_ABST
    Figure CN119600398B_ABST
Patent Text Reader

Abstract

The present invention discloses an acne grading method, system, device and medium based on multimodal fusion, belonging to the severity grading of facial acne in the field of artificial intelligence technology. The purpose is to solve the technical problem of poor grading effect of the severity of facial acne in the prior art. The constructed facial acne grading model includes a VGG feature extraction module, a multimodal feature fusion module, an acne feature perception module and a comprehensive grading module. The VGG feature extraction module extracts facial features in different modalities. The multimodal feature fusion module fuses the facial features in different modalities by cross-attention modeling the relationship between them. The acne feature perception module enhances the perception of acne features through the cross-attention mechanism. The comprehensive grading module makes a comprehensive and accurate acne severity grading result. This method realizes the comprehensive grading of the severity of acne by combining multimodal patient face pictures, and greatly improves the facial acne severity grading effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence, relates to the grading of the severity of facial acne, and particularly relates to a method, system, device and medium for acne grading based on multimodal fusion. Background Art

[0002] Acne is one of the most common skin diseases. According to surveys, being affected by hormonal levels during puberty, as well as unhealthy eating and sleeping habits, etc., are all likely to cause skin acne. During the daily treatment of acne, doctors will grade the severity of acne, which can help doctors formulate more targeted treatment plans for patients and also facilitate patients to timely understand their own conditions.

[0003] The grading of the severity of acne usually has different standards in different regions, but most are comprehensively judged based on doctors' experience and the number of skin lesions on the patient's face. For example, some grading methods will comprehensively divide the overall situation of the patient's face into five grades: Very Mild (very mild or no drug treatment required), Mild (mild), Moderate (moderate), Severe (severe), and Very Severe (very severe), indicating an urgent need for professional treatment. Of course, there are also those that divide the severity of facial acne into 4 grades.

[0004] When grading the severity of acne, most are still manually graded by doctors, which greatly increases the workload and intensity of doctors, and the grading efficiency is low. With the development of artificial intelligence and deep neural network technologies, it is very necessary to automatically grade the severity of acne with the help of a computer for patients to understand their own conditions and assist doctors in formulating treatment plans.

[0005] The invention patent application with the application number 202310795157.7 discloses a method, device, equipment and medium for acne grading. The method includes: obtaining a facial image of a target user, wherein the facial image includes an image corresponding to acne; detecting the facial image through a pre-trained detection model to obtain at least one acne type and the number of acne corresponding to the acne type; grading the acne type and the number of acne corresponding to the acne type through a pre-trained grading model to obtain the acne grade of the target user. This solution realizes the analysis of the facial image of the target user to obtain the acne grade of the target user, making the acne grade independent of doctors' experience and making the division of the acne grade more objective and accurate.

[0006] The invention patent application with the application number 202410122717.7 discloses an acne grading method, device, equipment and storage medium. It inputs an unlabeled pre-trained sample image set into the first feature extraction network and the second feature extraction network of the initial model; determines a target similarity matrix according to the first pre-trained sample feature set output by the first feature extraction network, the second pre-trained sample feature set output by the second feature extraction network, and the target pre-trained sample feature set obtained from the database; adjusts the network parameters in the initial model according to the loss function value calculated from the target similarity matrix to obtain an optimized initial model; selects one of the feature extraction networks in the optimized initial model to construct a pre-trained model; trains the pre-trained model with a training sample image set with label data to obtain a target acne grading model, which strengthens the feature learning ability of the model and improves the accuracy of acne grading.

[0007] Similar to the acne grading method in the above-mentioned invention patent, when the prior art uses computer-aided acne grading, it only focuses on the unimodal pictures of the patient's face, which is different from the doctor's actual diagnosis process of paying attention to the overall situation of the patient's face; and the existing multimodal methods for natural images are not suitable for acne image tasks because they do not consider the impact on the subsequent perception of acne lesion areas when fusing multimodal features. As a result, in the prior art, the network's perception ability of facial acne is low, the feature extraction effect of facial acne is poor, and the grading effect of the severity of facial acne is not good. Summary of the Invention

[0008] The purpose of the present invention is to provide an acne grading method, system, equipment and medium based on multimodal fusion to solve the technical problem of poor grading effect of the severity of facial acne in the prior art.

[0009] In order to achieve the above purpose, the present invention specifically adopts the following technical solutions:

[0010] An acne grading method based on multimodal fusion includes the following steps:

[0011] Step S1, obtain acne image sample data;

[0012] Obtain facial image samples in different modalities and label the facial image samples to obtain label data;

[0013] Step S2, construct a facial acne grading model;

[0014] Construct a facial acne grading model. The facial acne grading model includes a VGG feature extraction module, a multi-modal feature fusion module, an acne feature perception module, and a comprehensive grading module. The VGG feature extraction module extracts facial features in different modalities. The multi-modal feature fusion module fuses the facial features in different modalities by cross-attention modeling the relationships between them. The acne feature perception module enhances the perception of acne features through a cross-attention mechanism. The comprehensive grading module produces a comprehensive and accurate acne severity grading result;

[0015] Step S3: Train the facial acne grading model;

[0016] Use the facial image samples and label data obtained in step S1 to train the facial acne grading model constructed in step S2 to obtain a mature facial acne grading model;

[0017] Step S4: Facial acne grading;

[0018] Obtain acne images of the face to be graded in different modalities and input them into the mature facial acne grading model obtained in step S3 to obtain an acne severity grading result.

[0019] Furthermore, in step S1, preprocess the facial image samples. The specific method is as follows:

[0020] Step S1.1: Uniformly represent the facial image samples in a 3D format and use the bilinear interpolation method to uniformly scale the facial image samples to a size of 3 * 520 * 520;

[0021] Step S1.2: Center-crop the facial image samples to 3 * 512 * 512 and perform data augmentation in sequence by channel normalization, random flipping, and rotation;

[0022] Step S1.3: Digitally number the severity levels of each case of acne in the facial image samples and process them into one-hot label form.

[0023] Furthermore, in step S2, when the multi-modal feature fusion module performs feature fusion, specifically:

[0024] Step S2.2.1: First, use a 1 * 1 convolution operation to adjust the multi-view features extracted by the VGG feature extraction module 、 、 to dimension L, and then expand in the spatial dimension to obtain new features representing the left, middle, and right three modalities respectively 、 、 ,

[0025]

[0026] Among them, , , represents a 1*1 convolution, represents unfolding in the spatial dimension;

[0027] Step S2.2.2, perform cross-attention calculation between the features of the left modality and the features of the middle modality, and between the features of the right modality and the features of the middle modality, to perform interaction between different modalities.

[0028] Furthermore, in step S2.2.2, when performing cross-attention calculation between the features of the left modality and the features of the middle modality, specifically:

[0029] Step S2.2.2.1, use the first fully connected layer to linearly transform the middle modality features and use it as the query matrix Q in the attention calculation. Use the second fully connected layer to linearly transform the left modality features and use it as the key matrix K and value matrix V in the attention calculation;

[0030]

[0031] Among them, represents linear transformation, , , all represent the parameters of the linear transformation;

[0032] Step S2.2.2.2, calculate the product between the query matrix Q and the transpose of the key matrix K through matrix algorithm to obtain the cross-attention weight matrix :

[0033]

[0034] Among them, represents the number of features, represents the feature dimension;

[0035] Step S2.2.2.3, perform normalization processing on the cross-attention weight matrix :

[0036]

[0037] Among them, represents the scaling factor, represents performing normalization operation on each row of the cross-attention weight matrix ;

[0038] Step S2.2.2.4, the attention weight matrix Multiply with the value matrix V, and update the features of the left - hand modality using the relationship between the features of the left - hand modality and the middle - hand modality to obtain the updated left - hand auxiliary modality features ;

[0039]

[0040]

[0041] Among them, 、 represent the learnable parameters of the first fully - connected layer, 、 represent the learnable parameters of the second fully - connected layer, 、 respectively represent the multiplication operations of the attention weight matrix and the value matrix V, represents the left - hand modality features before update, represents the left - hand modality features after update.

[0042] Furthermore, in step S2, when the acne feature perception module perceives the acne area, specifically:

[0043] Step S2.3.1, initialize a learnable perception matrix , ; Each row vector in the perception matrix is a perceptron, ;

[0044] Step S2.3.2, use the perception matrix as the query matrix in the attention calculation, and use the fused features as the key and value matrices; Enhance the perception of acne features through the cross - attention mechanism, calculate the multiplication of the query matrix and the key matrix to obtain the perception weight matrix , and multiply the perception weight matrix with the value matrix to extract the features centered on the lesion area ;

[0045] Among them, , , N represents the number of perceptron vectors, L represents the dimension of the perceptron vectors, and Y represents the perception results of different perceptrons.

[0046] Furthermore, in step S2, the comprehensive grading module includes x fully - connected layers, and each fully - connected layer predicts the corresponding feature vector in the feature . The output of the nth fully - connected layer is expressed as:

[0047]

[0048] Among them, represents the learnable parameters of the nth fully connected layer;

[0049] The final acne severity prediction result is obtained by weighted summation of the prediction results of each feature according to the contribution degree, specifically:

[0050]

[0051] Among them, represents the contribution degree of the ith feature vector extracted by the perceptron, represents the prediction result of the ith feature vector extracted.

[0052] Furthermore, in step S3, when training the facial acne grading model, the loss function is:

[0053]

[0054] Among them, C represents the total number of grades, G represents the total number of data used for training, j represents the data of the jth patient, represents the probability that the severity prediction grade of the jth patient is z, represents the probability that the label of the jth patient is on grade z.

[0055] An acne grading system based on multimodal fusion, comprising:

[0056] A facial image sample acquisition module, configured to acquire facial image samples in different modalities and label the facial image samples to obtain label data;

[0057] A facial acne grading model construction module, configured to construct a facial acne grading model, the facial acne grading model includes a VGG feature extraction module, a multimodal feature fusion module, an acne feature perception module, and a comprehensive grading module. The VGG feature extraction module extracts features in images in different modalities, the multimodal feature fusion module fuses the features in images in different modalities by cross-attention to model the relationship between the features, the acne feature perception module enhances the perception of acne features through a cross-attention mechanism, and the comprehensive grading module makes a comprehensive and accurate acne severity grading result;

[0058] A facial acne grading model training module, configured to train the facial acne grading model constructed by the facial acne grading model construction module by using the facial image samples and label data acquired by the facial image sample acquisition module to obtain a mature facial acne grading model;

[0059] The facial acne grading module is used to obtain acne images of the face to be graded in different modalities, and input them into the facial acne grading model training module to obtain a mature facial acne grading model, and obtain the acne severity grading result.

[0060] A computer device includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of the above method.

[0061] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the processor executes the steps of the above method.

[0062] The beneficial effects of the present invention are as follows:

[0063] 1. In the present invention, the constructed facial acne grading model includes a VGG feature extraction module, a multi-modal feature fusion module, an acne feature perception module, and a comprehensive grading module. The VGG feature extraction module extracts facial features in different modalities. The multi-modal feature fusion module fuses the facial features in different modalities by cross-attention modeling the relationship between them. The acne feature perception module enhances the perception of acne features through the cross-attention mechanism. The comprehensive grading module makes a comprehensive and accurate acne severity grading result; by combining multi-modal patient face pictures, a comprehensive grading of acne severity is realized, and the grading effect of facial acne severity is improved.

[0064] 2. In the present invention, the modal fusion method adopted by the multi-modal feature fusion module effectively reduces the redundant information between different modalities, enhances the consistency of modalities, and takes into account the impact of modal fusion on subsequent acne grading; the perceptron mechanism is used in the multi-modal feature fusion module, which greatly enhances the model's attention to acne features, improves the effect of feature fusion, and finally improves the grading effect of facial acne severity.

[0065] 3. In the present invention, considering the multi-modal views of the face, in the acne grading task, doctors usually make a comprehensive judgment on the condition of the patient's face. However, previous methods are all for diagnosing single-modal pictures of the patient's face, which does not conform to the actual clinical needs, and will also affect the accuracy of diagnosis due to the lack of diagnostic basis. The proposed method fuses pictures of different modalities of the patient's face to achieve comprehensive diagnostic grading.

[0066] 4. In the present invention, the proposed grading method uses a perceptron to extract acne features, and the attention of the perceptron can more finely display acne features, which can be visualized as the basis for explaining the diagnostic results.

[0067] 5. In the present invention, this grading method takes into account the influence of acne characteristics on acne diagnosis and grading. Similar to the process of diagnosis by dermatologists in clinical practice, it is necessary to pay attention to the acne characteristics on the patient's face and use this as the basis for judging the severity of acne. Previous methods either lack attention to acne characteristics or require additional labeled data, seriously affecting the effectiveness and generality of the model. The proposed perceptron mechanism enhances the attention to acne characteristics without additional labeled information.

[0068] 6. In the present invention, this grading method has strong generalization ability. First, in the case of missing modalities, through modality transformation, the original grading and diagnosis effect can also be enhanced. On the other hand, since no additional labeled information such as acne location and acne quantity is used, it can be easily migrated to different diagnostic criteria, reducing the human and material costs of data collection.

[0069] 7. In the present invention, this grading method has a certain degree of diagnostic interpretability. The visualization result of the perceptual attention matrix of the perceptron for acne characteristics can be regarded as the diagnostic basis of the model, which provides a more reliable explanation for the diagnostic result of the model. Previous models are completely black-box models, which is a big problem for assisting medical diagnosis tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 is a schematic flowchart of the present invention;

[0071] Figure 2 is a schematic structural diagram of the facial acne grading model in the present invention;

[0072] Figure 3 is a schematic diagram of feature fusion of the multi-modal feature fusion module in the present invention;

[0073] Figure 4 is a schematic structural diagram of the residual connection and multi-layer perception in the multi-modal feature fusion module of the present invention;

[0074] Figure 5 is a schematic structural diagram of the comprehensive grading module in the present invention;

[0075] Figure 6 is a schematic diagram of acne perception visualization in the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0076] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.

[0077] Therefore, based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0078] Embodiment 1

[0079] This embodiment provides an acne grading method based on multimodal fusion. As Figure 1 shown, it includes the following steps:

[0080] Step S1: Obtain acne image sample data;

[0081] Obtain facial image samples in different modalities, and label the facial image samples to obtain label data.

[0082] The facial image samples include two parts: The first part is from the publicly available facial acne severity grading dataset ACNE04, which contains single-modal pictures from each patient. There are a total of 1457 cases and 1457 patient pictures, and pictures in other modalities are generated through data augmentation; The second part is the diagnostic dataset HX-MMACNE of patients in different modalities taken by the right VISIA collected in cooperation with dermatologists. This dataset contains a total of 652 cases and 1956 pictures.

[0083] Both of the above two parts of facial image samples contain label data, which are all labeled by professional doctors.

[0084] During training, it is necessary to perform unified preprocessing and enhancement on the facial images so that the picture size and format of the facial images are suitable for computer processing, and the generalization of the network is improved through data augmentation.

[0085] Perform preprocessing on the facial image samples. The specific method is as follows:

[0086] Step S1.1: Uniformly represent the facial image samples in a 3D format, and uniformly scale the facial image samples to a size of 3 * 520 * 520 using the bilinear interpolation method;

[0087] Step S1.2: Center-crop the facial image samples to 3 * 512 * 512, and perform data augmentation in sequence through channel normalization, random flipping, and rotation;

[0088] Step S1.3: Digitally number the severity levels of each case of acne in the facial image samples and process them into one-hot label form.

[0089] The above-mentioned image samples are divided into a training set and a test set according to a ratio of 4:1, which are used to train and optimize the model and evaluate the final model respectively. To enhance the stability of the model results, both datasets are divided by five-fold cross-validation, and the average of the five results is taken as the final evaluation result.

[0090] Step S2, construct a facial acne grading model;

[0091] Construct a facial acne grading model. The facial acne grading model includes a VGG feature extraction module, a multi-modal feature fusion module, an acne feature perception module, and a comprehensive grading module. The VGG feature extraction module extracts facial features in different modalities. The multi-modal feature fusion module fuses the facial features in different modalities by cross-attention to model the relationship between them. The acne feature perception module enhances the perception of acne features through the cross-attention mechanism. The comprehensive grading module makes a comprehensive and accurate acne severity grading result, as Figure 2 shown.

[0092] For the VGG feature extraction module, an existing and conventional VGG network can be used without creative labor. The face image is input into the VGG feature extraction module for feature extraction, and facial features in different modalities are extracted , where represent the height, width, and channel dimensions of the features respectively.

[0093] Next, these multi-modal features are sent to the multi-modal feature fusion module, where the redundancy between different modal features is reduced through the interaction between features, the consistency between features is enhanced, and multi-modal feature fusion is performed.

[0094] In the prior art, the acne severity of a patient's face is usually graded based on a single-modal image. However, a single-modal image usually only contains a partial area of the face, ignoring acne lesions in other parts of the face, resulting in inaccurate diagnosis results and unable to fully reflect the patient's condition. Therefore, in this embodiment, three different-modal images are used to obtain the complete information of the patient's face, and the multi-modal feature fusion module is designed. By interacting between features of different modalities, the redundancy between features is reduced, the consistency information is enhanced, and the operation of splicing and fusing the channel dimensions is used to avoid interference with subsequent acne grading. When the multi-modal feature fusion module performs feature fusion, as Figure 3 shown, specifically:

[0095] Step S2.2.1, first use a 1*1 convolution operation to adjust the multi-view features extracted by the VGG feature extraction module , , To dimension L, and then expand in the spatial dimension to obtain new features representing the left, middle, and right modalities respectively , , ,

[0096]

[0097] Among them, , , represents a 1*1 convolution, represents expansion in the spatial dimension;

[0098] Step S2.2.2, perform cross-attention calculation between the features of the left modality and the features of the middle modality, and between the features of the right modality and the features of the middle modality, to perform interaction between different modalities.

[0099] When performing cross-attention calculation, take the cross-attention calculation between the features of the left modality and the features of the middle modality as an example for illustration. When performing cross-attention calculation between the features of the left modality and the features of the middle modality, specifically:

[0100] Step S2.2.2.1, use the first fully connected layer to linearly transform the features of the middle modality and use it as the query matrix Q in the attention calculation. Use the second fully connected layer to linearly transform the features of the left modality and use it as the key matrix K and value matrix V in the attention calculation;

[0101]

[0102] Among them, represents linear transformation, , , all represent the parameters of linear transformation;

[0103] Step S2.2.2.2, after obtaining the query matrix Q and the key matrix K, calculate the product between the query matrix Q and the transpose of the key matrix K through a matrix algorithm to obtain the cross-attention weight matrix :

[0104]

[0105] Among them, represents the number of features, represents the feature dimension;

[0106] Since matrix multiplication can be regarded as vector multiplication, as shown in the above formula, each row vector in the query matrix Q is multiplied by each column vector in the key matrix K through matrix multiplication, and the result is the similarity relationship between the two vectors. The vectors here are actually the eigenvectors in the feature maps of the two modalities. Therefore, the result can represent the relationship between the two modalities.

[0107] Step S2.2.2.3, in order to eliminate the influence of vector dimension, normalize the cross-attention weight matrix as follows:

[0108]

[0109] where, represents the scaling factor, and represents the normalization operation on each row of the cross-attention weight matrix so that the sum of probabilities of each row is 1.

[0110] Step S2.2.2.4, taking the output of the nth element in the ith row of the attention weight matrix as an example, the output of the Softmax operation can be expressed as:

[0111]

[0112] Then multiply the attention weight matrix by the value matrix V, and update the features of the left modality with the relationship between the features of the left modality and the features of the middle modality to obtain the updated features of the left modality . At the same time, in order to accelerate the training and convergence process of the model, a residual connection mechanism is used therein, and in order to make the updated features of the left modality more abundant, a multi-layer perceptron is also used to improve the learning effect of the model, as Figure 4 shown. The specific principle of the residual connection is to sum the updated features of the left modality and the features of the left modality before update in an element-wise addition manner to achieve accelerating the convergence of the model, which is specifically expressed as:

[0113]

[0114] The multi-layer perceptron is implemented through two fully-connected layers with ReLU activation functions. The ReLU activation function is defined as:

[0115]

[0116] Therefore, the calculation method of the final left auxiliary modality features after interaction is:

[0117]

[0118] Among them, and represent the learnable parameters of the first fully connected layer, and represent the learnable parameters of the second fully connected layer, and respectively represent the multiplication operation of the attention weight matrix and the value matrix V, represents the left-modal feature before update, represents the left-modal feature after update.

[0119] The features of the right modality and the features of the middle modality are calculated by cross-attention in the same way to obtain the updated right feature map . For the features of the middle modality with more lesion features, continue to calculate the attention to obtain the new features of the middle modality , that is, the query, key, and value in the attention all come from the linear transformation of the same data.

[0120] Through the cross-attention operation, the redundant information between different modality feature maps is reduced, the consistency between different modalities is improved, the interference of background information is suppressed, and then the fusion is performed by using the feature connection method to avoid affecting the subsequent perception of the acne lesion area, and finally the fused feature is obtained.

[0121] Similar to the doctor's judgment of the patient's condition by observing the acne lesions on the patient's face, the acne lesion area also plays a crucial role in the process of automated diagnosis. Therefore, the model needs to focus on these areas and extract the features of these lesion areas. This point was ignored in previous deep learning methods, resulting in poor performance. At the same time, in order to improve the generality of the proposed method, the network structure needs to be improved to perform acne perception under the condition of only using image-level label supervision. To solve this problem, this embodiment adopts an attention-based acne feature perception module, and interacts with the fused features by setting perceptrons to perceive the acne area and extract the features centered on the acne area; at the same time, since the size and type of the acne area are variable, multiple perceptrons are set to improve the perception ability. When the acne feature perception module perceives the acne area, specifically:

[0122] Step S2.3.1, initialize a learnable perception matrix , ; Each row vector in the perception matrix is a perceptron, ;

[0123] Step S2.3.2: Use the perception matrix as the query matrix in attention calculation, and use the fused features as the key and value matrices; enhance the perception of acne features through the cross-attention mechanism, calculate the multiplication of the query matrix and the key matrix to obtain the perception weight matrix , multiply the perception weight matrix with the value matrix to extract the features centered on the lesion area ;

[0124] Among them, , , N represents the number of perceptron vectors, L represents the dimension of the perceptron vectors, and Y represents the perception results of different perceptrons.

[0125] After obtaining the features centered on the acne area , predict the severity of the patient based on these features , the row vectors in the matrix have a dimension of C, and Y represents the prediction result of the feature vector extracted by the nth perceptron.

[0126] The comprehensive grading module includes x fully connected layers. Each fully connected layer predicts the corresponding feature vector in the features . The output of the nth fully connected layer is expressed as:

[0127]

[0128] Among them, represents the learnable parameters of the nth fully connected layer;

[0129] At the same time, it should be noted that for the extracted feature matrix , each feature vector in it contributes differently to the final prediction. Therefore, it is necessary to evaluate the contribution degree of each feature vector . Specifically, input the set of feature vectors Z into a fully connected layer with a Sigmoid activation function to predict the final contribution degree . The definition of the Sigmoid activation function is as follows, mapping the calculated contribution degree to the range of 0-1:

[0130]

[0131] Therefore, the final contribution degree prediction formula is as follows:

[0132]

[0133] The final prediction result of acne severity It is obtained by weighted summation of the prediction results of each feature according to the contribution degree, specifically as follows:

[0134]

[0135] Among them, represents the contribution degree of the i-th feature vector extracted by the perceptron, represents the prediction result of the i-th feature vector extracted, as Figure 5 shown.

[0136] Step S3, train the facial acne grading model;

[0137] Use the face image samples and label data obtained in step S1 to train the facial acne grading model constructed in step S2 to obtain a mature facial acne grading model.

[0138] First is the input data part. Each patient contains pictures of three modalities (left, middle, right). For the pictures of each modality, they are converted into corresponding numerical values using the RGB three-channel encoding method. At the same time, in order to balance the three aspects of grading effect, consumption of computing resources, and grading speed, the size of each picture is processed to 512×512 pixels. For the severity level labels annotated in natural language, they are corresponding encoded using numerical numbers. The collected MVACNE dataset is divided into five levels, corresponding encoding is 0 - 4, and the public dataset ACNE04 is divided into four severity levels, corresponding encoding is 0 - 3. In addition, in order to enhance the generalization of the model, each picture is enhanced by random cropping, horizontal flipping, and random rotation to form different forms from the original picture. And normalization processing is performed in the channel dimension to accelerate the convergence process of the model.

[0139] The loss function is the goal of model optimization. Since the acne severity grading task is similar to the image classification task, the multi-class cross-entropy loss function is selected as the optimization goal. According to what is mentioned above, for each patient, its prediction result is: , where C represents the total number of severity levels. Therefore, each component represents the possibility that the model thinks this picture belongs to this category. Then, the output of the model is passed through , and is converted into a probability distribution through the Softmax function , where The calculation formula of is as follows:

[0140]

[0141] In the previous text, the severity levels were corresponded to numerical codes. Now, the numerical codes need to be converted into one-hot encodings. Taking the five types of labels in the HX-MMACNE dataset as an example, if the label is "1", then initialize a vector of length 5 and change the value at position 1 (counting from 0) to 1, that is, the label , which means the probability at level 1 is considered as 1 by the label, and the probabilities at other positions are 0. After conversion, the label vector obtained is: . Then the final loss function is:

[0142]

[0143] where C represents the total number of levels, G represents the total number of data used for training, j represents the data of the j-th patient, represents the probability that the predicted severity level of the j-th patient is z, represents the probability that the label of the j-th patient is at level z (either 0 or 1).

[0144] The loss function represents the gap between the results of the model and the actual labels. The gradients are calculated through the backpropagation of the loss function to update the parameters. Specifically, by taking the derivative of the loss function, the gradients of the loss function with respect to the model parameters are obtained, and then the parameters are updated by a certain proportion in the gradient direction using the learning rate. By continuously adjusting this process, the model reaches an optimal set of parameters. In this article, the initial learning rate is set to 0.001, and the Stochastic Gradient Descent (SGD) optimizer is used to optimize the model parameters. The weight decay is set to 0.0001, which is actually a regularization term used to prevent overfitting of the model and improve the generalization ability of the model. During training, the batch size is set to 16, that is, there are 16 cases in each batch, and the order is shuffled in each round of training. A total of 120 rounds of training are performed, and the learning rate is decayed by 50% every 20 rounds to prevent the model from getting stuck at a local optimum.

[0145] After the model training is completed, the test set is used to test the effect of the trained model. Only when the indicators meet certain requirements can it be considered that the model has achieved the desired effect. In this embodiment, when testing the model, the hierarchical accuracy (Accuracy), precision, sensitivity, specificity, and Youden index (YI) are mainly used for comprehensive evaluation, and the higher these indicators, the better. This test is carried out on two datasets, namely HX-MMACNE collected in cooperation with the hospital and the publicly available dataset ACNE04. The test results on the collected HX-MMACNE dataset are shown in the following table:

[0146]

[0147] The best results for the indicators are shown in bold black in the above table. It can be seen that the grading method of this application has achieved the best results on the collected test set.

[0148] In addition, in order to test the performance of the proposed method in the case of missing modalities and its adaptability to other dataset distributions, tests were carried out on the public dataset ACNE04. Since each case in the ACNE04 dataset only contains unimodal data of the patient's face, in this embodiment, through image flipping transformation, the generated images are used as other auxiliary modalities for input, which is also a measure of the method of this application in dealing with missing modalities. The test results are as follows:

[0149]

[0150] Similarly, it can be seen that the method of this application can achieve state-of-the-art performance even when other modality data is missing in the public dataset, reflecting the superiority of the method of this application and its robustness in the case of missing modalities.

[0151] In addition, in order to show the degree of attention of the method of this application to acne features, in this embodiment, the acne attention matrix in the model is also visualized. It can be seen that the proposed method can finely focus on the acne positions, which are exactly the evidence for doctors to diagnose patients, and also increase the credibility of the diagnosis grading results and improve the interpretability of the model. While the previous prediction models are completely black-box models and have low credibility in the field of medical auxiliary diagnosis, as shown in Figure 6 shown.

[0152] Step S4, facial acne grading;

[0153] Obtain acne images of the face to be graded under different modalities, and input them into the mature facial acne grading model obtained in step S3 to obtain the acne severity grading result.

[0154] Embodiment 2

[0155] This embodiment provides an acne grading system based on multi-modal fusion, including:

[0156] A facial image sample acquisition module, configured to acquire facial image samples under different modalities and label the facial image samples to obtain label data.

[0157] The facial image samples consist of two parts: The first part is from the publicly available facial acne severity grading dataset ACNE04, which contains unimodal pictures from each patient. There are a total of 1457 cases and 1457 patient pictures, and other modal pictures are generated through data augmentation. The second part is the diagnostic dataset HX-MMACNE of patients in different modalities taken by the right VISIA collected in cooperation with dermatologists. This dataset contains a total of 652 cases and 1956 pictures.

[0158] The above two parts of facial image samples both contain labeled data, which are all annotated by professional doctors.

[0159] During training, it is necessary to perform unified preprocessing and augmentation on the facial images so that the picture size and format of the facial images are suitable for computer processing, and the generalization of the network is improved through data augmentation.

[0160] Preprocess the facial image samples, and the specific method is as follows:

[0161] Step S1.1, uniformly use a 3D format to represent the facial image samples, and use the bilinear interpolation method to uniformly scale the facial image samples to a size of 3 * 520 * 520.

[0162] Step S1.2, center-crop the facial image samples to 3 * 512 * 512, and perform data augmentation in sequence through channel normalization, random flipping, and rotation.

[0163] Step S1.3, digitally number the severity levels of each case of acne in the facial image samples and process them into one-hot label form.

[0164] Divide the above image samples into a training set and a test set according to a ratio of 4:1, which are used to train and optimize the model and finally test and evaluate the model. To enhance the stability of the model results, both datasets are divided in the way of five-fold cross-validation, and the average value of the five results is taken as the final evaluation result.

[0165] The facial image sample acquisition module is used to construct a facial acne grading model. The facial acne grading model includes a VGG feature extraction module, a multi-modal feature fusion module, an acne feature perception module, and a comprehensive grading module. The VGG feature extraction module extracts facial features in different modalities. The multi-modal feature fusion module fuses the facial features in different modalities by cross-attention modeling the relationships between them. The acne feature perception module enhances the perception of acne features through the cross-attention mechanism. The comprehensive grading module makes a comprehensive and accurate acne severity grading result, as Figure 2 shown.

[0166] The VGG feature extraction module can use the existing and conventional VGG network without creative labor. The facial image is input into the VGG feature extraction module for feature extraction, and facial features in different modalities are extracted. , where respectively represent the height, width, and channel dimensions of the features.

[0167] Next, these multi-modal features are sent to the multi-modal feature fusion module, where the redundancy between different modal features is reduced through the interaction between features, the consistency between features is enhanced, and multi-modal feature fusion is performed.

[0168] In the prior art, the acne severity grading is usually performed on a single-modal image of the patient's face. However, a single-modal picture usually only contains a partial area of the face, ignoring the acne lesions in other parts of the face, resulting in inaccurate diagnosis results and inability to fully reflect the patient's condition. Therefore, in this embodiment, three pictures in different modalities are used to obtain the complete information of the patient's face, and the multi-modal feature fusion module is designed. By interacting between the features of different modalities, the redundancy between features is reduced, the consistency information is enhanced, and at the same time, the operation of splicing and fusing in the channel dimension avoids interference with the subsequent acne grading. When the multi-modal feature fusion module performs feature fusion, as Figure 3 shown, specifically:

[0169] Step S2.2.1, first use the 1*1 convolution operation to adjust the multi-view features extracted by the VGG feature extraction module , , to dimension L, and then expand in the spatial dimension to obtain new features representing the left, middle, and right three modalities respectively , , ,

[0170]

[0171] Among them, , , represents the 1*1 convolution, represents expanding in the spatial dimension;

[0172] Step S2.2.2, perform cross-attention calculation between the features of the left modality and the features of the middle modality, and between the features of the right modality and the features of the middle modality, to perform interaction between different modalities.

[0173] When performing cross-attention calculation, the cross-attention calculation between the features of the left modality and the features of the middle modality is taken as an example for illustration. When performing cross-attention calculation between the features of the left modality and the features of the middle modality, specifically:

[0174] Step S2.2.2.1, use the first fully connected layer to linearly transform the middle modality features and use it as the query matrix Q in the attention calculation. Use the second fully connected layer to linearly transform the left modality features and use it as the key matrix K and value matrix V in the attention calculation;

[0175]

[0176] Among them, represents linear transformation, , , all represent the parameters of the linear transformation;

[0177] Step S2.2.2.2, after obtaining the query matrix Q and the key matrix K, calculate the product between the query matrix Q and the transpose of the key matrix K through matrix algorithm to obtain the cross-attention weight matrix :

[0178]

[0179] Among them, represents the number of features, represents the feature dimension;

[0180] Because matrix multiplication can be regarded as vector multiplication, as shown in the above formula, each row vector in the query matrix Q is multiplied by each column vector in the key matrix K through matrix multiplication, and the result is the similarity relationship between the two vectors. The vectors here are actually the feature vectors in the feature maps of the two modalities. Therefore, the result can represent the relationship between the two modalities.

[0181] Step S2.2.2.3, in order to eliminate the influence of vector dimension, normalize the cross-attention weight matrix :

[0182]

[0183] Among them, represents the scaling factor, represents the normalization operation on each row of the cross-attention weight matrix so that the sum of probabilities of each row is 1.

[0184] Step S2.2.2.4, using the attention weight matrix Taking the output of the nth element in the ith row in [ ] as an example, the output of the Softmax operation can be expressed as:

[0185]

[0186] Next, multiply the attention weight matrix by the value matrix V, and update the features of the left modality using the relationship between the features of the left modality and the features of the middle modality to obtain the updated features of the left modality . At the same time, in order to accelerate the training and convergence process of the model, a residual connection mechanism is used therein, and in order to make the updated features of the left modality more abundant, a multi-layer perceptron is also used to improve the learning effect of the model, as shown in Figure 4 . The specific principle of the residual connection is to sum the updated features of the left modality and the features of the left modality before update in an element-wise addition manner to achieve accelerating the convergence of the model, which is specifically expressed as:

[0187]

[0188] The multi-layer perceptron is implemented through two fully connected layers with ReLU activation functions. The ReLU activation function is defined as:

[0189]

[0190] Therefore, the calculation method of the final left auxiliary modality features after interaction is:

[0191]

[0192] where , represent the learnable parameters of the first fully connected layer, , represent the learnable parameters of the second fully connected layer, , respectively represent the multiplication operation of the attention weight matrix and the value matrix V, represents the features of the left modality before update, represents the updated features of the left modality.

[0193] The features of the right modality and the features of the middle modality are calculated for cross-attention in the same way to obtain the updated right feature map . For the features of the middle modality with more lesion features, continue to calculate the attention to obtain the new features of the middle modality , that is, the query, key, and value in the attention all come from the linear transformation of the same data.

[0194] Through cross-attention operation, the redundant information between feature maps of different modalities is reduced, the consistency between different modalities is enhanced, and the interference of background information is suppressed. Then, fusion is performed by means of feature connection to avoid affecting the subsequent perception of acne lesion areas, and finally, fused features are obtained. 。

[0195] Similar to a doctor's judgment of a patient's condition by observing the acne lesions on the patient's face, the acne lesion area also plays a crucial role in the process of automated diagnosis. Therefore, it is necessary for the model to focus on these areas and extract the features of these lesion areas. However, this point was ignored in previous deep learning methods, resulting in poor performance. At the same time, in order to improve the generality of the proposed method, the network structure needs to be improved to perform acne perception under the supervision of only image-level labels. To solve this problem, this embodiment adopts an attention-based acne feature perception module, and interacts with the fused features by setting perceptrons to perceive the acne area and extract the features centered on the acne area. At the same time, since the size and type of acne areas vary, multiple perceptrons are set to improve the perception ability. When the acne feature perception module perceives the acne area, specifically:

[0196] Step S2.3.1, initialize a learnable perception matrix , ; Each row vector in the perception matrix is a perceptron, ;

[0197] Step S2.3.2, use the perception matrix as the query matrix in attention calculation, and use the fused features as the key and value matrices; enhance the perception of acne features through the cross-attention mechanism, calculate the multiplication of the query matrix and the key matrix to obtain the perception weight matrix , and multiply the perception weight matrix by the value matrix to extract the features centered on the lesion area ;

[0198] Among them, , , N represents the number of perceptron vectors, L represents the dimension of the perceptron vectors, and Y represents the perception results of different perceptrons.

[0199] After obtaining the features centered on the acne area , predict the severity of the patient based on these features . The dimension of the row vector in the matrix is C, and Y represents the prediction result of the feature vector extracted by the nth perceptron.

[0200] The comprehensive grading module includes x fully connected layers, and each fully connected layer makes predictions on the corresponding feature vectors in the features The output of the nth fully connected layer is expressed as:

[0201]

[0202] where represents the learnable parameters of the nth fully connected layer;

[0203] At the same time, it should be noted that the extracted feature matrix and each feature vector in it has different contributions to the final prediction. Therefore, it is necessary to evaluate the contribution degree of each feature vector Specifically, the feature vector set Z is input into a fully connected layer with a Sigmoid activation function to predict the final contribution degree where the definition of the Sigmoid activation function is as follows, mapping the calculated contribution degree to the range of 0-1:

[0204]

[0205] Therefore, the final contribution degree prediction formula is as follows:

[0206]

[0207] The final acne severity prediction result is obtained by weighted summation of the prediction results of each feature according to the contribution degree, specifically:

[0208]

[0209] where represents the contribution degree of the ith feature vector extracted by the perceptron, represents the prediction result of the ith feature vector extracted, as Figure 5 shown.

[0210] The facial acne grading model training module is used to train the facial acne grading model constructed by the facial acne grading model construction module with the facial image samples and label data obtained by the facial image sample acquisition module to obtain a mature facial acne grading model.

[0211] First is the input data part. Each patient contains pictures in three modalities (left, middle, and right). For the pictures in each modality, they are converted into corresponding numerical values using the RGB three-channel encoding method. At the same time, in order to balance the three aspects of grading effect, consumption of computing resources, and grading speed, the size of each picture is processed to 512×512 pixels. For the severity level labels with natural language annotations, they are corresponding numbered with digits. The collected MVACNE dataset is divided into five levels, which are correspondingly encoded as 0 - 4. The publicly available dataset ACNE04 is divided into four severity levels, which are correspondingly encoded as 0 - 3. In addition, in order to enhance the generalization of the model, each picture is augmented by random cropping, horizontal flipping, and random rotation to form different forms from the original picture. And normalization processing is performed in the channel dimension to accelerate the convergence process of the model.

[0212] The loss function is the optimization goal of the model. Since the acne severity grading task is similar to the image classification task, the multi-class cross-entropy loss function is selected as the optimization goal. As mentioned above, for each patient, its prediction result is: , where C represents the total number of severity levels. Therefore, each component represents the probability that the model thinks this picture belongs to this category. Then, the output of the model is passed through , and is converted into a probability distribution through the Softmax function , where The calculation formula is as follows:

[0213]

[0214] In the above text, the severity levels are corresponding numbered with digits. Now, the digits need to be converted into one-hot encoding. Taking the five-class labels in the HX-MMACNE dataset as an example, if the label is "1", then initialize a vector with a length of 5, and change the value at position 1 (counting from 0) to 1, that is, the label , which means the probability at level 1 is 1 and the probabilities at other levels are 0. After conversion, the label vector is obtained: . Then the final loss function is:

[0215]

[0216] Among them, C represents the total number of levels, G represents the total number of data used for training, j represents the data of the j-th patient, represents the probability that the predicted severity level of the j-th patient is z, represents the probability of the label of the j-th patient at level z (either 0 or 1).

[0217] The loss function represents the gap between the model's results and the actual labels. The gradient is calculated through the backpropagation of the loss function to update the parameters. Specifically, by taking the derivative of the loss function, the gradient of the loss function with respect to the model parameters is obtained, and then the parameters are updated by a certain proportion in the gradient direction using the learning rate. By continuously adjusting this process, the model reaches an optimal set of parameters. In this paper, the initial learning rate is set to 0.001, and the Stochastic Gradient Descent (SGD) optimizer is used to optimize the model's parameters. The weight decay is set to 0.0001, which is actually a regularization term used to prevent the model from overfitting and improve the model's generalization ability. During training, the batch size is set to 16, that is, there are 16 cases in each batch, and the order is shuffled in each round of training. A total of 120 rounds of training are performed, and the learning rate is decayed by 50% every 20 rounds to prevent the model from getting stuck at a local optimum.

[0218] After the model training is completed, the test set is used to test the effect of the trained model. Only when the indicators meet certain requirements can the model be considered to have achieved the desired effect. In this embodiment, when testing the model, the hierarchical accuracy (Accuracy), precision, sensitivity, specificity, and Youden index (YI) are mainly used for comprehensive evaluation, and the higher these indicators, the better. This test is carried out on two datasets, namely HX-MMACNE collected in cooperation with the hospital and the publicly available dataset ACNE04. The test results on the collected HX-MMACNE dataset are shown in the following table:

[0219]

[0220] The best results for each indicator in the above table are shown in bold black font. It can be seen that the classification method of this application has achieved the best results on the collected test set.

[0221] In addition, to verify the performance of the proposed method in the case of missing modalities and its adaptability to other dataset distributions, tests were carried out on the publicly available dataset ACNE04. Since each case in the ACNE04 dataset only contains unimodal data of the patient's face, in this embodiment, through image flipping transformation, the generated images are used as other auxiliary modalities for input, which is also a measure of the method of this application in dealing with missing modalities. The test results are as follows:

[0222]

[0223] Similarly, it can be seen that the method of this application can achieve state-of-the-art performance even when other modality data is missing in the publicly available dataset, demonstrating the superiority of the method of this application and its robustness in the case of missing modalities.

[0224] In addition, to demonstrate the degree of attention of the method of the present application to acne features, in this embodiment, the acne attention matrix in the model is also visualized. It can be seen that the proposed method can attentively focus on the acne positions in a refined manner, which are exactly the evidence for doctors to diagnose patients, increasing the credibility of the diagnosis grading result and improving the interpretability of the model. In contrast, previous prediction models are completely black-box models, with low credibility in the field of medical auxiliary diagnosis, as shown in Figure 6 the appendix.

[0225] Step S4, facial acne grading;

[0226] Obtain acne images of the face to be graded in different modalities, and input them into the facial acne grading model training module to obtain a mature facial acne grading model and the acne severity grading result.

[0227] Embodiment 3

[0228] A computer device includes a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor is caused to execute the steps of the acne grading method based on multi-modal fusion.

[0229] Among them, the computer device can be a computing device such as a desktop computer, a notebook, a palm computer, or a cloud server. The computer device can perform human-computer interaction with the user through a keyboard, a mouse, a remote control, a touchpad, or a voice control device, etc.

[0230] The memory at least includes one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or D interface display memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory may be an internal storage unit of the computer device, such as the hard disk or memory of the computer device. In other embodiments, the memory may also be an external storage device of the computer device, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device. Of course, the memory may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the memory is commonly used to store the operating system and various application software installed on the computer device, such as the program code of the acne grading method based on multi-modal fusion, etc. In addition, the memory can also be used to temporarily store various data that have been output or will be output.

[0231] In some embodiments, the processor may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor is generally used to control the overall operation of the computer device. In this embodiment, the processor is used to run the program code stored in the memory or process data, such as running the program code of the acne grading method based on multi-modal fusion.

[0232] Embodiment 4

[0233] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to execute the steps of the acne grading method based on multi-modal fusion.

[0234] Wherein, the computer-readable storage medium stores an interface display program, and the interface display program can be executed by at least one processor to cause the at least one processor to execute the steps of the acne grading method based on multi-modal fusion as described above.

[0235] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the acne grading method based on multi-modal fusion described in the embodiments of the present application.

Claims

1. A method for acne grading based on multimodal fusion, characterized in that It includes the following steps: Step S1, obtaining acne image sample data; Obtaining face image samples in different modalities, and annotating the face image samples to obtain label data; Step S2, constructing a facial acne grading model; Constructing a facial acne grading model, which includes a VGG feature extraction module, a multi-modal feature fusion module, an acne feature perception module, and a comprehensive grading module. The VGG feature extraction module extracts facial features in different modalities. The multi-modal feature fusion module fuses the facial features in different modalities by cross-attention to model the relationship between them. The acne feature perception module enhances the perception of acne features through the cross-attention mechanism. The comprehensive grading module makes a comprehensive and accurate acne severity grading result; Step S3, training the facial acne grading model; Using the face image samples and label data obtained in Step S1 to train the facial acne grading model constructed in Step S2 to obtain a mature facial acne grading model; Step S4, facial acne grading; Obtaining acne images of the face to be graded in different modalities, and inputting them into the mature facial acne grading model obtained in Step S3 to obtain an acne severity grading result; In Step S2, when the acne feature perception module perceives the acne area, specifically: Step S2.3.1, initialize a learnable perception matrix , ; Each row vector in the perception matrix is a perceptron, ; Step S2.3.2, take the perception matrix as the query matrix in attention calculation, and take the fused feature as the key and value matrices; enhance the perception of acne features through the cross-attention mechanism, and obtain the perception weight matrix by multiplying the query matrix and the key matrix , multiply the perception weight matrix by the value matrix to extract the features centered on the lesion area ; Among them, , , N represents the number of perceptron vectors, L represents the dimension of the perceptron vectors, and Y represents the perception results of different perceptrons.

2. The acne grading method based on multi-modal fusion according to claim 1, wherein, In Step S1, preprocessing the face image samples, and the specific method is: Step S1.1, uniformly representing the face image samples in a 3D format, and uniformly scaling the face image samples to a size of 3*520*520 using the bilinear interpolation method; Step S1.2, centrally cropping the face image samples to 3*512*512, and performing data augmentation in sequence through channel normalization, random flipping, and rotation; Step S1.3, digitally numbering the severity level of each case of acne in the face image samples and processing them into one-hot label form.

3. The acne grading method based on multimodal fusion according to claim 1, wherein In Step S2, when the multi-modal feature fusion module performs feature fusion, specifically: Step S2.2.1, first use a 1*1 convolution operation to adjust the multi-view features extracted by the VGG feature extraction module , , to dimension L, and then expand in the spatial dimension to obtain new features representing the left, middle, and right three modalities respectively , , , Among them, , , represents a 1*1 convolution, represents expansion in the spatial dimension; Step S2.2.2, performing cross-attention calculations respectively between the features of the left modality and the features of the middle modality, and between the features of the right modality and the features of the middle modality to perform interactions between different modalities.

4. The acne grading method based on multimodal fusion according to claim 1, wherein In Step S2.2.2, when performing cross-attention calculation between the features of the left modality and the features of the middle modality, specifically: Step S2.2.2.1, use the first fully connected layer to linearly transform the middle modal features and use them as the query matrix Q in attention calculation. Use the second fully connected layer to linearly transform the left modal features and use them as the key matrix K and value matrix V in attention calculation; Among them, represents a linear transformation, , , all represent the parameters of the linear transformation; Step S2.2.2.2, calculate the product between the query matrix Q and the transpose of the key matrix K through a matrix algorithm to obtain the cross-attention weight matrix : Among them, represents the number of features, represents the feature dimension; Step S2.2.2.3, normalize the cross-attention weight matrix : Among them, represents the scaling factor, represents the normalization operation on each row of the cross-attention weight matrix ; Step S2.2.2.4, multiply the attention weight matrix by the value matrix V, and update the features of the left modality with the relationship between the features of the left modality and the features of the middle modality to obtain the updated left auxiliary modality features ; Among them, , represent the learnable parameters of the first fully connected layer, , represent the learnable parameters of the second fully connected layer, , respectively represent the multiplication operations of the attention weight matrix and the value matrix V, represents the left modal feature before update, represents the left modal feature after update.

5. The acne grading method based on multimodal fusion according to claim 1, wherein In step S2, the comprehensive grading module includes x fully connected layers, and each fully connected layer makes predictions on the corresponding feature vectors in the features The output of the nth fully connected layer is expressed as: in the feature vectors Among them, represents the learnable parameters of the nth fully connected layer; Final acne severity prediction result It is obtained by weighted summation of the prediction results of each feature according to the contribution degree, specifically: Among them, represents the contribution degree of the i-th feature vector extracted by the perceptron, represents the prediction result of the i-th feature vector extracted.

6. The acne grading method based on multimodal fusion according to claim 1, wherein In Step S3, when training the facial acne grading model, the loss function is: Among them, C represents the total number of levels, G represents the total number of data for training, j represents the data of the j-th patient, represents the probability that the predicted severity level of the j-th patient is z, represents the probability that the label of the j-th patient is at level z.

7. An acne grading system based on multi-modal fusion, characterized in that, It includes: A face image sample acquisition module, which is used to obtain face image samples in different modalities, and annotate the face image samples to obtain label data; A facial acne grading model construction module, which is used to construct a facial acne grading model. The facial acne grading model includes a VGG feature extraction module, a multi-modal feature fusion module, an acne feature perception module, and a comprehensive grading module. The VGG feature extraction module extracts features in images in different modalities. The multi-modal feature fusion module fuses the features in images in different modalities by cross-attention to model the relationship between them. The acne feature perception module enhances the perception of acne features through the cross-attention mechanism. The comprehensive grading module makes a comprehensive and accurate acne severity grading result; The facial acne grading model training module is used to train the facial acne grading model constructed by the facial acne grading model construction module with the facial image samples and label data obtained by the facial image sample acquisition module, so as to obtain a mature facial acne grading model; The facial acne grading module is used to obtain acne images of the face to be graded in different modalities, and input them into the facial acne grading model training module to obtain a mature facial acne grading model, so as to obtain the acne severity grading result; In the facial acne grading model construction module, when the acne feature perception module perceives the acne area, specifically: Step S2.3.1, initialize a learnable perception matrix , ; Each row vector in the perception matrix is a perceptron, ; Step S2.3.2: Use the sensing matrix as the query matrix in attention calculation, and use the fused features as the key and value matrices; enhance the perception of acne features through the cross-attention mechanism, and obtain the perception weight matrix by calculating the multiplication of the query matrix and the key matrix. Multiply the perception weight matrix with the value matrix to extract the features centered on the lesion area ; Among them, , , N represents the number of perceptron vectors, L represents the dimension of the perceptron vectors, and Y represents the perception results of different perceptrons.

8. A computer device, characterized in that: It includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: A computer program is stored. When the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Acne grading method, device, equipment and medium

    CN116863522A

  • Acne grading method, device and equipment and storage medium

    CN117649683A

  • Acne grading method, system and equipment based on semi-supervised learning and storage medium

    CN115440346A

  • Edge enhanced medical image segmentation method based on double decoders

    CN117911423A