Lesion image enhancement method and device based on diffusion model, equipment and medium

By using a diffusion model-based skin lesion image enhancement method, the target skin lesion image is automatically processed, solving the problems of long generation time and susceptibility to intervention caused by manual operation, and achieving efficient and reliable image enhancement results.

CN122434754APending Publication Date: 2026-07-21JINHUA MUNICIPAL CENT HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JINHUA MUNICIPAL CENT HOSPITAL
Filing Date
2026-04-29
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In existing technologies, the enhancement process of target skin lesion images relies on manual operation, resulting in long generation time and susceptibility to human intervention, making it difficult to meet the needs of teaching or research.

Method used

A skin lesion image enhancement method based on a diffusion model is adopted. The original skin lesion image is obtained from the training samples, the region is divided and the feature vectors are concatenated, the diffusion model is used to generate the predicted enhanced image, and the model parameters are optimized by gradient descent to finally generate the enhanced target skin lesion image.

Benefits of technology

It eliminates the need for manual operation, reduces generation time, improves generation efficiency, and enhances the reliability and quality of the augmented images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122434754A_ABST
    Figure CN122434754A_ABST
Patent Text Reader

Abstract

The application relates to the image technical field and the artificial intelligence technical field, and discloses a skin lesion image enhancement method and device based on a diffusion model, equipment and a medium. The method comprises the following steps: inputting prompt text and extended description input text of each target skin lesion area into a text encoder, generating a text feature vector of each target skin lesion area through the text encoder, splicing an image feature vector of each target skin lesion area, the text feature vector of each target skin lesion area and corresponding enhancement parameters of each target skin lesion area, and generating a fusion feature vector of each target skin lesion area; inputting the fusion feature vector of each target skin lesion area into an updated diffusion model, generating an optimal enhanced image of each target skin lesion area through the updated diffusion model, and fusing the optimal enhanced image of each target skin lesion area to generate an enhanced target skin lesion image. The application can improve the generation efficiency of the enhanced target skin lesion image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of image technology and artificial intelligence technology, and in particular to a method, apparatus, device and medium for enhancing skin lesion images based on a diffusion model. Background Technology

[0002] Targeted lesion images have broad application prospects in dermatology teaching and research. By enhancing targeted lesion images, their visual recognition can be effectively improved, thereby providing higher-quality image data for teaching activities.

[0003] However, the enhancement process for target lesion images primarily relies on manual operation. This manual method requires operators to observe the target lesion images and manually adjust them one by one to achieve the enhancement effect needed for teaching or research. This increases the generation time of the enhanced target lesion images and is easily affected by human intervention, which is detrimental to improving the generation efficiency of enhanced target lesion images. Therefore, how to generate enhanced target lesion images is a technical problem that urgently needs to be solved. Summary of the Invention

[0004] This application provides a method, apparatus, device, and medium for enhancing skin lesion images based on a diffusion model, in order to solve the aforementioned technical problem of how to generate enhanced target skin lesion images.

[0005] In a first aspect, embodiments of this application provide a skin lesion image enhancement method based on a diffusion model, applied to an electronic device, the skin lesion image enhancement method comprising: The original skin lesion images in the training samples are obtained, and the original skin lesion images are divided into multiple original skin lesion regions. The image feature vector, text feature vector, and enhancement parameters corresponding to each original skin lesion region are concatenated to generate the fusion feature vector of each original skin lesion region. The original skin lesion images are skin lesion images captured by a multispectral imager in the past time period. The fused feature vector of each original skin lesion region is input into the diffusion model. The diffusion model generates a predicted enhanced image of each original skin lesion region. The influence coefficient corresponding to each time step is generated by the preset influence coefficient generation model. The divergence loss of the diffusion model is generated based on the influence coefficient corresponding to each time step, the predicted enhanced image of each original skin lesion region, and the preset divergence loss model. The mean squared error loss of the diffusion model is generated based on the preset mean squared error loss model. The divergence loss and mean squared error loss are added together to generate the total loss of the diffusion model. The parameters of the diffusion model are updated using gradient descent until the total loss of the diffusion model is less than the preset loss value. The updated diffusion model is then stored. The target skin lesion image is obtained and divided into multiple target skin lesion regions using an image segmentation model. The target skin lesion image is a skin lesion image captured by a multispectral imager in the current time period. Each target lesion region is input into an image feature extraction network, which generates an image feature vector for each target lesion region. The prompt text and extended description of each target lesion region are input into a text encoder, which generates a text feature vector for each target lesion region. The image feature vector, text feature vector, and enhancement parameters corresponding to each target lesion region are concatenated to generate a fused feature vector for each target lesion region. The fused feature vector of each target lesion region is input into the updated diffusion model. The updated diffusion model generates the optimal enhanced image of each target lesion region. The optimal enhanced images of each target lesion region are then fused together by the fusion network to generate the enhanced target lesion image.

[0006] In one possible implementation of the first aspect, the original skin lesion image from the training samples is acquired. Using an image segmentation model, the original skin lesion image is divided into multiple original skin lesion regions. The image feature vector, text feature vector, and enhancement parameters corresponding to each original skin lesion region are concatenated to generate a fused feature vector for each original skin lesion region, including: The original skin lesion images in the training samples are obtained and input into the image segmentation model. The image segmentation model divides the original skin lesion images into multiple original skin lesion regions. From the annotation information of each original skin lesion region, the prompt text, the enhancement parameters corresponding to each original skin lesion region, and the real enhancement image corresponding to each original skin lesion region are obtained. Each original skin lesion area is input into an image feature extraction network, which generates an image feature vector for each original skin lesion area. The prompt text for each original skin lesion area is input into a large language model, which generates an extended description for each original skin lesion area. The prompt text and the extended description for each original skin lesion area are input into a text encoder, which generates a feature vector for the prompt text and a feature vector for the extended description for each original skin lesion area. The feature vectors of the prompt text and the extended description for each original skin lesion area are concatenated to generate a text feature vector for each original skin lesion area. Principal component analysis was used to reduce the dimensionality of the image feature vector and the text feature vector of each original lesion region. The dimensionality-reduced image feature vector and the dimensionality-reduced text feature vector of each original lesion region were generated. The dimensionality-reduced image feature vector, the dimensionality-reduced text feature vector of each original lesion region, and the enhancement parameters corresponding to each original lesion region were concatenated to generate the fused feature vector of each original lesion region.

[0007] In one possible implementation of the first aspect, each target lesion region is input into an image feature extraction network to generate an image feature vector for each target lesion region. The prompt text and extended description of each target lesion region are input into a text encoder to generate a text feature vector for each target lesion region. The image feature vector, text feature vector, and enhancement parameters corresponding to each target lesion region are concatenated to generate a fused feature vector for each target lesion region, including: Each target lesion region is input into an image feature extraction network, which generates an image feature vector for each target lesion region. The prompt text for each target lesion region is input into a large language model, which generates an extended description for each target lesion region. The prompt text and the extended description for each target lesion region are then input into a text encoder, which generates feature vectors for the prompt text and the extended description for each target lesion region. Finally, the feature vectors for the prompt text and the extended description for each target lesion region are concatenated to generate a text feature vector for each target lesion region. Principal component analysis is used to reduce the dimensionality of the image feature vector and the text feature vector of each target lesion region. The dimensionality-reduced image feature vector and the dimensionality-reduced text feature vector of each target lesion region are generated. The dimensionality-reduced image feature vector, the dimensionality-reduced text feature vector of each target lesion region and the corresponding enhancement parameters of each target lesion region are concatenated to generate the fused feature vector of each target lesion region.

[0008] In one possible implementation of the first aspect, the enhancement parameters corresponding to the original lesion region include the texture enhancement parameters corresponding to the original lesion region, the local contrast corresponding to the original lesion region, the edge enhancement parameters corresponding to the original lesion region, the gamma correction value corresponding to the original lesion region, and the spatial noise reduction intensity corresponding to the original lesion region; the enhancement parameters corresponding to the target lesion region include the texture enhancement parameters corresponding to the target lesion region, the local contrast corresponding to the target lesion region, the edge enhancement parameters corresponding to the target lesion region, the gamma correction value corresponding to the target lesion region, and the spatial noise reduction intensity corresponding to the target lesion region.

[0009] In one possible implementation of the first aspect, the influence coefficient generation model is defined as follows: ; Indicates the first The influence coefficient corresponding to the first time step; The larger the influence coefficient at the time step, the better the diffusion model performs at the _ ... The stronger the influence of the fused feature vector of the original lesion region in each time step; the stronger the influence of the fusion feature vector of the original lesion region in the first time step; The smaller the influence coefficient at the time step, the better the diffusion model performs at the _ ... The weaker the influence of the fused feature vector of the original lesion region on each time step; Indicates control parameters, The value range is from 0 to 1; Indicates the sequence number of the time step; This represents the total number of time steps.

[0010] In one possible implementation of the first aspect, the divergence loss model is defined as follows: ; ; This represents the divergence loss of the diffusion model. The larger the divergence loss of the diffusion model, the weaker its fitting ability on the original skin lesion area; the smaller the divergence loss of the diffusion model, the stronger its fitting ability on the original skin lesion area. This represents the total number of original lesion areas in the original lesion image; Indicates the serial number of the original skin lesion area; Indicates the first A true enhanced image of the original skin lesion area; Indicates the first The original skin lesion area during the spread process Predicted augmented images at each time step; Indicates the first The original skin lesion area during the diffusion process The predicted augmented image of the previous time step at each time step; Indicates the first The fused feature vector of each original skin lesion region; Indicates the first The fused feature vector at each time step; Indicates the first The influence coefficient corresponding to each time step; This represents the true probability distribution of the diffusion model in the forward process; This represents the expected value under the true probability distribution; These represent the model parameters of the diffusion model; This represents the predicted probability distribution generated by the diffusion model based on the model parameters.

[0011] In one possible implementation of the first aspect, the mean squared error loss model is defined as follows: ; ; The mean squared error loss of the diffusion model indicates that the diffusion model has a weaker ability to restore details in the original lesion area; the smaller the mean squared error loss of the diffusion model indicates that the diffusion model has a stronger ability to restore details in the original lesion area. Let represent the random sampling noise at time step t; This represents the expectation of the diffusion model under random sampled noise at time step t; Represents a diffusion model; These represent the model parameters of the diffusion model; Indicates the serial number of the original skin lesion area; Indicates the first The original skin lesion area during the diffusion process Predicted augmented images at each time step; Indicates the first The influence coefficient corresponding to each time step; Indicates the first The fused feature vector of each original skin lesion region; Indicates the first The fused feature vector at each time step; The diffusion model is represented in the first... Predictive noise at each time step output; This represents the square of the L2 norm.

[0012] Secondly, embodiments of this application provide a skin lesion image enhancement device based on a diffusion model, applied to an electronic device, comprising: The first acquisition module is used to acquire the original skin lesion images in the training samples, divide the original skin lesion images into multiple original skin lesion regions, and concatenate the image feature vector, text feature vector, and enhancement parameters corresponding to each original skin lesion region to generate a fusion feature vector for each original skin lesion region. The original skin lesion images are skin lesion images captured by the multispectral imager in the past time period. The first generation module is used to input the fused feature vector of each original skin lesion region into the diffusion model, generate a predicted enhanced image of each original skin lesion region through the diffusion model, generate the influence coefficient corresponding to each time step through the preset influence coefficient generation model, generate the divergence loss of the diffusion model according to the influence coefficient corresponding to each time step, the predicted enhanced image of each original skin lesion region and the preset divergence loss model, and generate the mean squared error loss of the diffusion model according to the preset mean squared error loss model. The second acquisition module is used to add the divergence loss and the mean squared error loss to generate the total loss of the diffusion model. It uses gradient descent to update the parameters of the diffusion model until the total loss of the diffusion model is less than the preset loss value. Then it stops updating the parameters of the diffusion model, stores the updated diffusion model, acquires the target skin lesion image, and divides the target skin lesion image into multiple target skin lesion regions through the image segmentation model. The target skin lesion image is the skin lesion image captured by the multispectral imager in the current time period. The second generation module is used to input each target lesion region into the image feature extraction network, generate an image feature vector for each target lesion region through the image feature extraction network, input the prompt text and extended description of each target lesion region into the text encoder, generate a text feature vector for each target lesion region through the text encoder, and concatenate the image feature vector, text feature vector and enhancement parameters corresponding to each target lesion region to generate a fused feature vector for each target lesion region. The third generation module is used to input the fused feature vector of each target lesion region into the updated diffusion model, generate the optimal enhanced image of each target lesion region through the updated diffusion model, and fuse the optimal enhanced images of each target lesion region through the fusion network to generate the enhanced target lesion image.

[0013] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the skin lesion image enhancement method described in the first aspect above.

[0014] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the skin lesion image enhancement method described in the first aspect above.

[0015] Fifthly, embodiments of this application provide a computer program product that, when run on an electronic device, causes the electronic device to execute the skin lesion image enhancement method described in the first aspect.

[0016] The beneficial effects of the embodiments of this application are as follows: Firstly, the fusion feature vector of each target lesion region is input into the updated diffusion model. The updated diffusion model generates the optimal enhanced image of each target lesion region. The optimal enhanced images of each target lesion region are then fused together by the fusion network to generate the enhanced target lesion image. Since no manual operation is required, the generation time of the enhanced target lesion image is reduced, which is beneficial to improving the generation efficiency of the enhanced target lesion image. Secondly, since the updated diffusion model is not affected by human intervention, it helps to improve the reliability of the enhanced target lesion image. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is an application scenario diagram of the skin lesion image enhancement method provided in the embodiments of this application; Figure 2 This is a schematic flowchart of the skin lesion image enhancement method provided in the embodiments of this application; Figure 3 A flowchart illustrating the implementation of S204 provided in this application embodiment; Figure 4 A schematic block diagram of the skin lesion image enhancement device provided in the embodiments of this application; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.

[0020] The skin lesion image enhancement method provided in this application embodiment can be applied to electronic devices, including but not limited to servers, mobile phones, and tablet computers. This application embodiment does not impose any restrictions on the specific type of electronic device.

[0021] Please see Figure 1 , Figure 1The application scenario diagram of the skin lesion image enhancement method provided in the embodiments of this application is described in detail below: The electronic device connects to the file storage system, accesses the file storage system, retrieves the dataset from the file storage system, the dataset contains multiple training samples, and the original skin lesion images in the training samples are obtained.

[0022] In this embodiment of the application, the electronic device can automatically extract training samples in batches from the file storage system according to a preset time interval or condition trigger, without manual intervention, which can reduce the time for obtaining training samples and improve the efficiency of obtaining training samples.

[0023] Please see Figure 2 , Figure 2 This is a schematic flowchart of a skin lesion image enhancement method provided in an embodiment of this application, which can be applied to electronic devices.

[0024] like Figure 2 As shown, the skin lesion image enhancement method provided in this application includes the following steps, which are detailed below: S201, Obtain the original skin lesion image from the training sample, divide the original skin lesion image into multiple original skin lesion regions, and concatenate the image feature vector, text feature vector, and enhancement parameters corresponding to each original skin lesion region to generate the fusion feature vector of each original skin lesion region. The original skin lesion image is a skin lesion image captured by a multispectral imager in the past time period. Among them, a multispectral imager is an optical acquisition device that can simultaneously acquire multiple continuous narrow-band spectral images of a target scene.

[0025] Among them, skin lesion images are digital images that primarily capture abnormal areas of the skin. Their content clearly presents the morphological outline, boundary features, color distribution, and surface texture structure of the skin lesions, and is used to record, analyze, or display the visual information of abnormal areas of the skin.

[0026] The original skin lesion image is divided into multiple original skin lesion regions, which do not overlap with each other. Each original skin lesion region is a local sub-image of the original skin lesion image, and the union of all original skin lesion regions is equal to the original skin lesion image.

[0027] This process involves acquiring original skin lesion images from the training samples, dividing them into multiple original skin lesion regions using an image segmentation model, and concatenating the image feature vector, text feature vector, and enhancement parameters corresponding to each original skin lesion region to generate a fused feature vector for each region. This fusion vector includes: The original skin lesion images in the training samples are obtained and input into the image segmentation model. The image segmentation model divides the original skin lesion images into multiple original skin lesion regions. From the annotation information of each original skin lesion region, the prompt text, the enhancement parameters corresponding to each original skin lesion region, and the real enhancement image corresponding to each original skin lesion region are obtained. Each original skin lesion area is input into an image feature extraction network, which generates an image feature vector for each original skin lesion area. The prompt text for each original skin lesion area is input into a large language model, which generates an extended description for each original skin lesion area. The prompt text and the extended description for each original skin lesion area are input into a text encoder, which generates a feature vector for the prompt text and a feature vector for the extended description for each original skin lesion area. The feature vectors of the prompt text and the extended description for each original skin lesion area are concatenated to generate a text feature vector for each original skin lesion area. Principal component analysis was used to reduce the dimensionality of the image feature vector and the text feature vector of each original lesion region. The dimensionality-reduced image feature vector and the dimensionality-reduced text feature vector of each original lesion region were generated. The dimensionality-reduced image feature vector, the dimensionality-reduced text feature vector of each original lesion region, and the enhancement parameters corresponding to each original lesion region were concatenated to generate the fused feature vector of each original lesion region.

[0028] The prompt text for the preset lesion area and the extended description of the preset lesion area are both textual descriptions used to guide the updated diffusion model to perform enhancement operations on the preset lesion area.

[0029] For ease of explanation, the following example is provided: For example, there are multiple preset skin lesion areas, which are the first preset area, the second preset area, the third preset area, and the fourth preset area; The prompt text for the first preset area is: This area is an erythematous lesion. It is necessary to enhance the edge sharpness, improve the color difference between the erythema and normal skin, and suppress the surface oily shine.

[0030] The prompt text for the second preset area is: This area contains silvery-white scales. The texture contrast needs to be enhanced to highlight the scale layers and boundaries.

[0031] The prompt text for the third preset area is: This area is a mole. The original boundary shape should be maintained, the color difference should be enhanced appropriately, and over-sharpening should be avoided.

[0032] The prompt text for the fourth preset area is: This area has strong reflections. Highlights need to be suppressed to expose the pores and texture structure underneath.

[0033] The extended description of the first preset area is as follows: This area is an erythematous lesion located on the cheekbone of the face. It is roughly circular with partially clear and partially blurred borders. The color distribution in the first preset area is uneven, with a darker color in the center and gradually lighter color at the edges. There is slight oily sheen on the surface, which obscures some of the capillary texture. The surrounding normal skin has no obvious abnormalities. It is necessary to enhance the sharpness of the borders to make the blurred borders clearer, increase the color difference between the erythema and the normal skin to highlight the extent of the lesion, and at the same time suppress the oily sheen on the surface to expose the underlying capillary texture.

[0034] The extended description of the second preset area is as follows: This area contains silvery-white scales, which are layered and distributed. In some areas, the scales are thicker and the boundaries are irregular. The contrast between the scales and the red spots below is low, and the edges of some scales are blurred. The surface of the second preset area is rough and fine cracks are visible. It is necessary to enhance the texture contrast to highlight the layering of the scales, strengthen the boundary between the scales and the red spots, and at the same time maintain the natural shape of the scale edges to avoid over-sharpening and artifacts.

[0035] The extended description of the third preset area is as follows: This area is a pigmented nevus, oval in shape, with clear and regular borders; the pigment is evenly distributed, ranging from light brown to dark brown, with a moderate color difference from the surrounding normal skin; the surface is free of scales, exudate, and obvious reflection; the original border shape should be maintained, and the color difference should be moderately enhanced to highlight the distinction between the third preset area and normal skin, avoiding over-sharpening that causes jagged artifacts at the edges, while maintaining a smooth transition of the third preset area.

[0036] The extended description of the fourth preset area is as follows: This area has strong reflection, is located in the center of the lesion, is raised, and appears as a bright white patch with an irregular shape; The fourth preset area covers the underlying pore structure and texture details. Normal skin texture can be seen at the edges of some areas, but it is cut off by the highlight. It is necessary to suppress the highlight intensity, restore the pore and texture structure of the highlighted area, and at the same time maintain a natural transition between the fourth preset area and the surrounding skin to avoid brightness cutoff or artifacts.

[0037] S202, input the fused feature vector of each original skin lesion region into the diffusion model, generate the predicted enhanced image of each original skin lesion region through the diffusion model, generate the influence coefficient corresponding to each time step through the preset influence coefficient generation model, generate the divergence loss of the diffusion model according to the influence coefficient corresponding to each time step, the predicted enhanced image of each original skin lesion region and the preset divergence loss model, and generate the mean square error loss of the diffusion model according to the preset mean square error loss model. The influence coefficient generation model is defined as follows: ; Indicates the first The influence coefficient corresponding to the first time step; The larger the influence coefficient at the time step, the better the diffusion model performs at the _ ... The stronger the influence of the fused feature vector of the original lesion region in each time step; the stronger the influence of the fusion feature vector of the original lesion region in the first time step; The smaller the influence coefficient at the time step, the better the diffusion model performs at the _ ... The weaker the influence of the fused feature vector of the original lesion region on each time step; Indicates control parameters, The value range is from 0 to 1; Indicates the sequence number of the time step; This represents the total number of time steps.

[0038] In the influence coefficient generation model, It increases with increasing t, that is, as t increases, the th The influence coefficient corresponding to each time step also increases. This is because the fused feature vectors in the early time steps have more noise and lower quality, so smaller coefficients need to be assigned to suppress interference; the fused feature vectors in the later time steps gradually become clearer and of higher quality, so larger coefficients need to be assigned to highlight their contribution.

[0039] Preferably, The value is 0.3. When the value of is 0.3, the influence coefficient corresponding to the time step increases slowly, which is suitable for tasks that require the diffusion model to learn step by step.

[0040] The divergence loss model is defined as follows: ; ; This represents the divergence loss of the diffusion model. The larger the divergence loss of the diffusion model, the weaker its fitting ability on the original skin lesion area; the smaller the divergence loss of the diffusion model, the stronger its fitting ability on the original skin lesion area. This represents the total number of original lesion areas in the original lesion image; Indicates the serial number of the original skin lesion area; Indicates the first A true enhanced image of the original skin lesion area; Indicates the first The original skin lesion area during the diffusion process Predicted augmented images at each time step; Indicates the first The original skin lesion area during the diffusion process The predicted augmented image of the previous time step at each time step; Indicates the first The fused feature vector of each original skin lesion region; Indicates the first The fused feature vector at each time step; Indicates the first The influence coefficient corresponding to each time step; This represents the true probability distribution of the diffusion model in the forward process; This represents the expected value under the true probability distribution; These represent the model parameters of the diffusion model; This represents the predicted probability distribution generated by the diffusion model based on the model parameters.

[0041] The mean squared error loss model is defined as follows: ; ; The mean squared error loss of the diffusion model indicates that the diffusion model has a weaker ability to restore details in the original lesion area; the smaller the mean squared error loss of the diffusion model indicates that the diffusion model has a stronger ability to restore details in the original lesion area. Let represent the random sampling noise at time step t; This represents the expectation of the diffusion model under random sampled noise at time step t; Represents a diffusion model; These represent the model parameters of the diffusion model; Indicates the serial number of the original skin lesion area; Indicates the first The original skin lesion area during the diffusion process Predicted augmented images at each time step; Indicates the first The influence coefficient corresponding to each time step; Indicates the first The fused feature vector of each original skin lesion region; Indicates the first The fused feature vector at each time step; The diffusion model is represented in the first... Predictive noise at each time step output; This represents the square of the L2 norm.

[0042] Since the fused feature vectors at different time steps contribute differently to the final reconstruction target, the influence coefficients corresponding to each time step can adaptively assess the importance of each time step. Time steps with larger weight coefficients indicate that the fused feature vectors at that step contribute significantly to the current task and should be retained and enhanced; time steps with weight coefficients close to zero indicate that the fused feature vectors at that step have low correlation with the task target or contain interference and should be suppressed or filtered out, thus reducing the interference of noise on the final output.

[0043] S203, the divergence loss and mean squared error loss are added together to generate the total loss of the diffusion model. The parameters of the diffusion model are updated using gradient descent until the total loss of the diffusion model is less than the preset loss value. The updated diffusion model is then stored. The target skin lesion image is obtained. The target skin lesion image is divided into multiple target skin lesion regions by the image segmentation model. The target skin lesion image is the skin lesion image captured by the multispectral imager in the current time period. The target lesion image is divided into multiple target lesion regions, which do not overlap with each other. Each target lesion region is a local sub-image of the target lesion image, and the union of all target lesion regions is equal to the target lesion image.

[0044] The process involves summing the divergence loss and the mean squared error loss to generate the total loss of the diffusion model. The goal is to minimize this total loss. Gradient descent is used to update the diffusion model's parameters until the total loss falls below a preset loss value. Once the parameters are updated, the model is stored. Since a larger total loss indicates weaker performance and a smaller total loss indicates stronger performance, using the total loss falling below the preset loss value as the termination condition essentially means that achieving the desired performance level is the direct basis for ending training. This ensures that the updated diffusion model meets the accuracy requirements of practical applications.

[0045] The past time period and the current time period are different time periods.

[0046] Past time period: refers to a continuous historical period preceding the current moment.

[0047] Current time period: refers to a continuous time interval including the current moment.

[0048] For ease of explanation, the following example is provided: For example, taking April 8, 2026 as the current time, the past time period was from April 1, 2026 to April 7, 2026. The current time period is from April 8, 2026 to April 15, 2026.

[0049] The skin lesion images captured by the multispectral imager between April 1, 2026 and April 7, 2026 are the preset skin lesion images.

[0050] The images of the skin lesions captured by the multispectral imager between April 8, 2026 and April 15, 2026 are the target skin lesion images.

[0051] S204, each target lesion region is input into the image feature extraction network, and the image feature extraction network generates an image feature vector for each target lesion region. The prompt text and extended description of each target lesion region are input into the text encoder, and the text encoder generates a text feature vector for each target lesion region. The image feature vector, the text feature vector, and the enhancement parameters corresponding to each target lesion region are concatenated to generate a fused feature vector for each target lesion region. The prompt text for the target lesion area and the extended description of the target lesion area are both textual descriptions used to guide the updated diffusion model to perform enhancement operations on the target lesion area.

[0052] For ease of explanation, the following example is provided: For example, there are multiple target lesion areas, which are the first target area, the second target area, the third target area, and the fourth target area; The prompt text for the first target area is: This area is an erythematous lesion. It is necessary to enhance the edge sharpness, improve the color difference between the erythema and normal skin, and suppress the surface oily shine.

[0053] The prompt text for the second target area is: This area contains silvery-white scales. The texture contrast needs to be enhanced to highlight the scale layers and boundaries.

[0054] The prompt text for the third target area is: This area is a mole. The original boundary shape should be maintained, the color difference should be enhanced appropriately, and over-sharpening should be avoided.

[0055] The prompt text for the fourth target area is: This area has strong reflections. Highlights need to be suppressed to expose the pores and texture structure underneath.

[0056] The extended description of the first target area is as follows: This area is an erythematous lesion located on the cheekbone of the face. It is roughly circular with partially clear and partially blurred borders. The color distribution of the first target area is uneven, with a darker color in the center and gradually lighter color at the edges. There is a slight oily sheen on the surface, which obscures some of the capillary texture. The surrounding normal skin has no obvious abnormalities. It is necessary to enhance the sharpness of the borders to clarify the blurred borders, increase the color difference between the erythema and the normal skin to highlight the extent of the lesion, and at the same time suppress the oily sheen on the surface to expose the underlying capillary texture.

[0057] The extended description of the second target area is as follows: This area contains silvery-white scales, which are layered and distributed. The scales are thicker in some areas and have irregular boundaries. The contrast between the scales and the erythema below is low, and the edges of some scales are blurred. The surface of the second target area is rough and has visible fine cracks. It is necessary to enhance the texture contrast to highlight the layering of the scales, strengthen the boundary between the scales and the erythema, and at the same time maintain the natural shape of the scale edges to avoid over-sharpening and artifacts.

[0058] The extended description of the third target area is as follows: This area is a pigmented nevus, oval in shape, with clear and regular borders; the pigment is evenly distributed, ranging from light brown to dark brown, with a moderate color difference from the surrounding normal skin; the surface is free of scales, exudate, and obvious reflection; the original border shape should be maintained, and the color difference should be moderately enhanced to highlight the distinction between the third target area and normal skin, avoiding over-sharpening that would cause jagged artifacts at the edges, while maintaining a smooth transition of the third target area.

[0059] The extended description of the fourth target area is as follows: This area has strong reflection, is located in the central raised part of the lesion, and appears as a bright white patch with an irregular shape; The fourth target area covers the underlying pore structure and texture details. Normal skin texture can be seen at the edge of some areas, but it is cut off by the highlight. It is necessary to suppress the highlight intensity, restore the pore and texture structure of the highlighted area, and at the same time maintain a natural transition between the fourth target area and the surrounding skin to avoid brightness cutoff or artifacts.

[0060] Among them, the enhancement parameters corresponding to the original skin lesion area include the texture enhancement parameters corresponding to the original skin lesion area, the local contrast corresponding to the original skin lesion area, the edge enhancement parameters corresponding to the original skin lesion area, the gamma correction value corresponding to the original skin lesion area, and the spatial noise reduction intensity corresponding to the original skin lesion area. The enhancement parameters corresponding to the target lesion area include the texture enhancement parameters, local contrast, edge enhancement parameters, gamma correction value, and spatial noise reduction intensity.

[0061] In the target lesion area enhancement process, the texture enhancement parameter is used to highlight the fine structure of the surface of the target lesion area, such as scale layers, pore openings and pigment particle distribution, so that the micro texture of the target lesion area is clearer and more distinguishable. Among them, the local contrast parameter is used to enhance the difference in brightness between the target lesion area and the surrounding normal skin, making the edge of the lesion sharper and facilitating accurate localization of the skin abnormality. Among them, the edge enhancement parameter further strengthens the gradient change of the boundary of the target lesion area and improves the contour integrity; Among them, the gamma correction value is used to adjust the overall brightness distribution of the target skin lesion area and improve the problem of local overexposure or underexposure caused by uneven illumination. Among them, the spatial noise reduction intensity is used to suppress sensor noise generated in the target skin lesion area during the acquisition process. The synergistic effect of the above parameters can significantly improve the amount of information in the target skin lesion area.

[0062] Specifically, the dimensionality-reduced image feature vector, the dimensionality-reduced text feature vector, and the enhancement parameters corresponding to each target lesion region are concatenated to generate a fused feature vector for each target lesion region. The fused feature vector of each target lesion region contains both visual and semantic information, which can make up for the lack of expressive power of a single modality.

[0063] S205, the fused feature vector of each target lesion region is input into the updated diffusion model. The updated diffusion model generates the optimal enhanced image of each target lesion region. The optimal enhanced images of each target lesion region are fused by the fusion network to generate the enhanced target lesion image.

[0064] The fused feature vector of each target lesion region is input into the updated diffusion model. The updated diffusion model generates the optimal enhanced image of each target lesion region. The optimal enhanced image of each target lesion region is then input into the fusion network. The fusion network extracts the shallow and deep texture features of each target lesion region step by step. The optimal enhanced images of each target lesion region are then fused to generate the enhanced target lesion image.

[0065] Among them, the fusion network includes, but is not limited to, U-shaped fusion network and multi-scale feature fusion network.

[0066] Among them, the U-shaped fusion network is a U-shaped encoding and decoding feature fusion network.

[0067] In this process, the optimal enhanced images of each target lesion area are fused together by a fusion network to generate an enhanced target lesion image. The enhanced target lesion image can effectively improve the edge sharpness, color contrast and texture level of the lesion area, enhance the recognizability of key visual features, and make the originally blurry lesion structure clearly presented.

[0068] For ease of explanation, the following example is provided: For example, a research center used a multispectral imager to acquire several images of target skin lesions. Due to limitations in the imaging principles and multi-band photosensitive characteristics of multispectral hardware, the photosensitivity of different spectral channels varies, leading to unclear lesion boundaries in the target skin lesion images. Using the skin lesion image enhancement method of this application, the originally blurred lesion edges become clear and sharp, the color difference between erythema and normal skin becomes more prominent, and subtle texture changes are also revealed. The enhanced target skin lesion images can be used to create easily understandable research materials, facilitating comprehension for non-specialists.

[0069] The skin lesion image enhancement method further includes the following steps: First, the fused feature vector of each target skin lesion region is input into the updated diffusion model. Then, the updated diffusion model generates the optimal enhanced image for each target skin lesion region. Finally, a fusion network is used to fuse these optimal enhanced images to generate the enhanced target skin lesion image. The enhanced images of the target skin lesions are uploaded to the resource system of the teaching platform.

[0070] The enhanced images of the target skin lesions are uploaded to the resource system of the teaching platform, supporting online learning and remote training, so that high-quality teaching resources can break through geographical limitations and benefit more institutions.

[0071] For ease of explanation, the following example is provided: For example, the boundaries between scales and erythema in the target lesion image are relatively blurred, making it difficult for medical students to accurately distinguish the edge of the lesion and the distribution range of scales. The electronic device uses the lesion image enhancement method of this application to enhance the target lesion image. Since the edge sharpness of the lesion area in the enhanced target lesion image is significantly improved, the color difference between erythema and normal skin is more prominent, and the layering of silvery-white scales is clearer, the enhanced target lesion image allows medical students to observe the morphological characteristics of the skin more intuitively.

[0072] The beneficial effects of the embodiments of this application are as follows: Firstly, the fusion feature vector of each target lesion region is input into the updated diffusion model. The updated diffusion model generates the optimal enhanced image of each target lesion region. The optimal enhanced images of each target lesion region are then fused together by the fusion network to generate the enhanced target lesion image. Since no manual operation is required, the generation time of the enhanced target lesion image is reduced, which is beneficial to improving the generation efficiency of the enhanced target lesion image. Secondly, since the updated diffusion model is not affected by human intervention, it helps to improve the reliability of the enhanced target lesion image.

[0073] Please see Figure 3 , Figure 3 The implementation flowchart of S204 provided in the embodiments of this application is described in detail below: S301, each target lesion region is input into an image feature extraction network, which generates an image feature vector for each target lesion region. The prompt text for each target lesion region is input into a large language model, which generates an extended description for each target lesion region. The prompt text and the extended description for each target lesion region are input into a text encoder, which generates a feature vector for the prompt text and a feature vector for the extended description for each target lesion region. The feature vectors for the prompt text and the extended description for each target lesion region are concatenated to generate a text feature vector for each target lesion region. S302, using principal component analysis, performs dimensionality reduction on the image feature vector and text feature vector of each target lesion region, generating dimensionality-reduced image feature vector and text feature vector for each target lesion region. The dimensionality-reduced image feature vector, text feature vector, and enhancement parameters corresponding to each target lesion region are then concatenated to generate a fused feature vector for each target lesion region.

[0074] The dimensionality-reduced image feature vector, the dimensionality-reduced text feature vector, and the enhancement parameters corresponding to each target lesion region are concatenated to generate a fused feature vector for each target lesion region, including: Each target lesion region and its corresponding prompt text are input into a multimodal large model. The multimodal large model generates enhancement parameters for each target lesion region. The dimensionality-reduced image feature vector, the dimensionality-reduced text feature vector, and the enhancement parameters for each target lesion region are then concatenated to generate a fused feature vector for each target lesion region.

[0075] The system generates enhancement parameters for each target lesion region using a multimodal large model, enabling rapid output of these parameters and a highly efficient and concise overall processing flow. The multimodal large model can directly utilize pre-trained models, which possess general visual understanding and cross-modal alignment capabilities, eliminating the need for downstream tasks to be trained from scratch and significantly saving computational resources and training time.

[0076] In this embodiment, the image feature vector and text feature vector of each target lesion region are dimensionality reduced, which significantly reduces the vector dimension, thereby reducing the amount of computation, reducing the occupation of GPU memory and RAM, and improving the training and inference speed of the diffusion model.

[0077] For the skin lesion image enhancement method described in the above embodiments, please refer to [link / reference]. Figure 4 , Figure 4 This is a schematic block diagram of the skin lesion image enhancement device provided in the embodiments of this application. Figure 4 The skin lesion image enhancement device 400 shown can be applied to, for example... Figure 1 The application scenario diagram shows electronic devices. The following section uses electronic devices as an example to illustrate this. Figure 4 The skin lesion image enhancement device 400 shown will be described in detail. The skin lesion image enhancement device 400 may include a first acquisition module 401, a first generation module 402, a second acquisition module 403, a second generation module 404, and a third generation module 405.

[0078] The first acquisition module 401 is used to acquire the original skin lesion image in the training sample, divide the original skin lesion image into multiple original skin lesion regions, and concatenate the image feature vector, text feature vector, and enhancement parameters corresponding to each original skin lesion region to generate a fusion feature vector for each original skin lesion region. The original skin lesion image is a skin lesion image captured by a multispectral imager in the past time period. The first generation module 402 is used to input the fused feature vector of each original skin lesion region into the diffusion model, generate a predicted enhanced image of each original skin lesion region through the diffusion model, generate the influence coefficient corresponding to each time step through the preset influence coefficient generation model, generate the divergence loss of the diffusion model according to the influence coefficient corresponding to each time step, the predicted enhanced image of each original skin lesion region and the preset divergence loss model, and generate the mean square error loss of the diffusion model according to the preset mean square error loss model. The second acquisition module 403 is used to add the divergence loss and the mean square error loss to generate the total loss of the diffusion model. The parameters of the diffusion model are updated using gradient descent until the total loss of the diffusion model is less than the preset loss value. The updated diffusion model is then stored. The target skin lesion image is acquired and divided into multiple target skin lesion regions using an image segmentation model. The target skin lesion image is a skin lesion image captured by a multispectral imager in the current time period. The second generation module 404 is used to input each target lesion region into the image feature extraction network, generate an image feature vector for each target lesion region through the image feature extraction network, input the prompt text and extended description of each target lesion region into the text encoder, generate a text feature vector for each target lesion region through the text encoder, and concatenate the image feature vector, text feature vector and enhancement parameters corresponding to each target lesion region to generate a fused feature vector for each target lesion region. The third generation module 405 is used to input the fused feature vector of each target lesion region into the updated diffusion model, generate the optimal enhanced image of each target lesion region through the updated diffusion model, and fuse the optimal enhanced images of each target lesion region through the fusion network to generate the enhanced target lesion image.

[0079] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0080] The beneficial effects of the embodiments of this application are as follows: Firstly, the fusion feature vector of each target lesion region is input into the updated diffusion model. The updated diffusion model generates the optimal enhanced image of each target lesion region. The optimal enhanced images of each target lesion region are then fused together by the fusion network to generate the enhanced target lesion image. Since no manual operation is required, the generation time of the enhanced target lesion image is reduced, which is beneficial to improving the generation efficiency of the enhanced target lesion image. Secondly, since the updated diffusion model is not affected by human intervention, it helps to improve the reliability of the enhanced target lesion image.

[0081] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0082] like Figure 5 As shown, Figure 5 The electronic device includes: at least one processor 20, a memory 21, and a computer program 22 stored in the memory 21 and executable on the at least one processor 20, wherein the processor 20 executes the computer program 22 to implement the steps in any of the above method embodiments.

[0083] The electronic device may include, but is not limited to, processor 20 and memory 21. Those skilled in the art will understand that... Figure 5 This is merely an example of an electronic device and does not constitute a limitation on electronic devices. It may include more or fewer components than shown in the illustration, or combinations of certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0084] The processor 20 is used to run a computer program 22 stored in the memory 21, and performs the following steps when executing the computer program 22: The original skin lesion images in the training samples are obtained, and the original skin lesion images are divided into multiple original skin lesion regions. The image feature vector, text feature vector, and enhancement parameters corresponding to each original skin lesion region are concatenated to generate the fusion feature vector of each original skin lesion region. The original skin lesion images are skin lesion images captured by a multispectral imager in the past time period. The fused feature vector of each original skin lesion region is input into the diffusion model. The diffusion model generates a predicted enhanced image of each original skin lesion region. The influence coefficient corresponding to each time step is generated by the preset influence coefficient generation model. The divergence loss of the diffusion model is generated based on the influence coefficient corresponding to each time step, the predicted enhanced image of each original skin lesion region, and the preset divergence loss model. The mean squared error loss of the diffusion model is generated based on the preset mean squared error loss model. The divergence loss and mean squared error loss are added together to generate the total loss of the diffusion model. The parameters of the diffusion model are updated using gradient descent until the total loss of the diffusion model is less than the preset loss value. The updated diffusion model is then stored. The target skin lesion image is obtained and divided into multiple target skin lesion regions using an image segmentation model. The target skin lesion image is a skin lesion image captured by a multispectral imager in the current time period. Each target lesion region is input into an image feature extraction network, which generates an image feature vector for each target lesion region. The prompt text and extended description of each target lesion region are input into a text encoder, which generates a text feature vector for each target lesion region. The image feature vector, text feature vector, and enhancement parameters corresponding to each target lesion region are concatenated to generate a fused feature vector for each target lesion region. The fused feature vector of each target lesion region is input into the updated diffusion model. The updated diffusion model generates the optimal enhanced image of each target lesion region. The optimal enhanced images of each target lesion region are then fused together by the fusion network to generate the enhanced target lesion image.

[0085] The processor 20 may be a central processing unit, or it may be other general-purpose processors, digital signal processors, field-programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0086] In some embodiments, the memory 21 may be an internal storage unit of the electronic device, such as a hard disk or memory of the electronic device.

[0087] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.

[0088] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A method for enhancing skin lesion images based on a diffusion model, characterized in that, The skin lesion image enhancement method, applied to electronic devices, includes: The original skin lesion images in the training samples are obtained, and the original skin lesion images are divided into multiple original skin lesion regions. The image feature vector, text feature vector, and enhancement parameters corresponding to each original skin lesion region are concatenated to generate the fusion feature vector of each original skin lesion region. The original skin lesion images are skin lesion images captured by a multispectral imager in the past time period. The fused feature vector of each original skin lesion region is input into the diffusion model. The diffusion model generates a predicted enhanced image of each original skin lesion region. The influence coefficient corresponding to each time step is generated by the preset influence coefficient generation model. The divergence loss of the diffusion model is generated based on the influence coefficient corresponding to each time step, the predicted enhanced image of each original skin lesion region, and the preset divergence loss model. The mean squared error loss of the diffusion model is generated based on the preset mean squared error loss model. The divergence loss and mean squared error loss are added together to generate the total loss of the diffusion model. The parameters of the diffusion model are updated using gradient descent until the total loss of the diffusion model is less than the preset loss value. The updated diffusion model is then stored. The target skin lesion image is obtained and divided into multiple target skin lesion regions using an image segmentation model. The target skin lesion image is a skin lesion image captured by a multispectral imager in the current time period. Each target lesion region is input into an image feature extraction network, which generates an image feature vector for each target lesion region. The prompt text and extended description of each target lesion region are input into a text encoder, which generates a text feature vector for each target lesion region. The image feature vector, text feature vector, and enhancement parameters corresponding to each target lesion region are concatenated to generate a fused feature vector for each target lesion region. The fused feature vector of each target lesion region is input into the updated diffusion model. The updated diffusion model generates the optimal enhanced image of each target lesion region. The optimal enhanced images of each target lesion region are then fused together by the fusion network to generate the enhanced target lesion image.

2. The method for enhancing skin lesion images according to claim 1, characterized in that, The original skin lesion images from the training samples are obtained. Using an image segmentation model, the original skin lesion images are divided into multiple original skin lesion regions. The image feature vector, text feature vector, and enhancement parameters corresponding to each original skin lesion region are concatenated to generate a fused feature vector for each original skin lesion region, including: The original skin lesion images in the training samples are obtained and input into the image segmentation model. The image segmentation model divides the original skin lesion images into multiple original skin lesion regions. From the annotation information of each original skin lesion region, the prompt text, the enhancement parameters corresponding to each original skin lesion region, and the real enhancement image corresponding to each original skin lesion region are obtained. Each original skin lesion area is input into an image feature extraction network, which generates an image feature vector for each original skin lesion area. The prompt text for each original skin lesion area is input into a large language model, which generates an extended description for each original skin lesion area. The prompt text and the extended description for each original skin lesion area are input into a text encoder, which generates a feature vector for the prompt text and a feature vector for the extended description for each original skin lesion area. The feature vectors of the prompt text and the extended description for each original skin lesion area are concatenated to generate a text feature vector for each original skin lesion area. Principal component analysis was used to reduce the dimensionality of the image feature vector and the text feature vector of each original lesion region. The dimensionality-reduced image feature vector and the dimensionality-reduced text feature vector of each original lesion region were generated. The dimensionality-reduced image feature vector, the dimensionality-reduced text feature vector of each original lesion region, and the enhancement parameters corresponding to each original lesion region were concatenated to generate the fused feature vector of each original lesion region.

3. The method for enhancing skin lesion images according to claim 1, characterized in that, Each target lesion region is input into an image feature extraction network, which generates an image feature vector for each target lesion region. The prompt text and extended description for each target lesion region are input into a text encoder, which generates a text feature vector for each target lesion region. The image feature vector, text feature vector, and enhancement parameters for each target lesion region are concatenated to generate a fused feature vector for each target lesion region, including: Each target lesion region is input into an image feature extraction network, which generates an image feature vector for each target lesion region. The prompt text for each target lesion region is input into a large language model, which generates an extended description for each target lesion region. The prompt text and the extended description for each target lesion region are then input into a text encoder, which generates feature vectors for the prompt text and the extended description for each target lesion region. Finally, the feature vectors for the prompt text and the extended description for each target lesion region are concatenated to generate a text feature vector for each target lesion region. Principal component analysis is used to reduce the dimensionality of the image feature vector and the text feature vector of each target lesion region. The dimensionality-reduced image feature vector and the dimensionality-reduced text feature vector of each target lesion region are generated. The dimensionality-reduced image feature vector, the dimensionality-reduced text feature vector of each target lesion region and the corresponding enhancement parameters of each target lesion region are concatenated to generate the fused feature vector of each target lesion region.

4. The method for enhancing skin lesion images according to claim 1, characterized in that, The enhancement parameters corresponding to the original lesion area include the texture enhancement parameters, local contrast, edge enhancement parameters, gamma correction value, and spatial noise reduction intensity. The enhancement parameters corresponding to the target lesion area include the texture enhancement parameters, local contrast, edge enhancement parameters, gamma correction value, and spatial noise reduction intensity.

5. The method for enhancing skin lesion images according to claim 1, characterized in that, The influence coefficient generation model is defined as follows: ; Indicates the first The influence coefficient corresponding to the first time step; The larger the influence coefficient at the time step, the better the diffusion model performs at the _ ... The stronger the influence of the fused feature vector of the original lesion region in each time step; the stronger the influence of the fusion feature vector of the original lesion region in the first time step; The smaller the influence coefficient at the time step, the better the diffusion model performs at the _ ... The weaker the influence of the fused feature vector of the original lesion region on each time step; Indicates control parameters, The value range is from 0 to 1; Indicates the sequence number of the time step; This represents the total number of time steps.

6. The method for enhancing skin lesion images according to claim 1, characterized in that, The divergence loss model is defined as follows: ; ; This represents the divergence loss of the diffusion model. The larger the divergence loss of the diffusion model, the weaker its fitting ability on the original skin lesion area; the smaller the divergence loss of the diffusion model, the stronger its fitting ability on the original skin lesion area. This represents the total number of original lesion areas in the original lesion image; Indicates the serial number of the original skin lesion area; Indicates the first A true enhanced image of the original skin lesion area; Indicates the first The original skin lesion area during the spread process Predicted augmented images at each time step; Indicates the first The original skin lesion area during the spread process The predicted augmented image of the previous time step at each time step; Indicates the first The fused feature vector of each original skin lesion region; Indicates the first The fused feature vector at each time step; Indicates the first The influence coefficient corresponding to each time step; This represents the true probability distribution of the diffusion model in the forward process; This represents the expected value under the true probability distribution; These represent the model parameters of the diffusion model; This represents the predicted probability distribution generated by the diffusion model based on the model parameters.

7. The method for enhancing skin lesion images according to claim 1, characterized in that, The mean squared error loss model is defined as follows: ; ; The mean squared error loss of the diffusion model indicates that the diffusion model has a weaker ability to restore details in the original lesion area; the smaller the mean squared error loss of the diffusion model indicates that the diffusion model has a stronger ability to restore details in the original lesion area. Let represent the random sampling noise at time step t; This represents the expectation of the diffusion model under random sampled noise at time step t; Represents a diffusion model; These represent the model parameters of the diffusion model; Indicates the serial number of the original skin lesion area; Indicates the first The original skin lesion area during the diffusion process Predicted augmented images at each time step; Indicates the first The influence coefficient corresponding to each time step; Indicates the first The fused feature vector of each original skin lesion region; Indicates the first The fused feature vector at each time step; The diffusion model is represented in the first... Predictive noise output at each time step; This represents the square of the L2 norm.

8. A skin lesion image enhancement device based on a diffusion model, characterized in that, Applied to electronic devices, including: The first acquisition module is used to acquire the original skin lesion images in the training samples, divide the original skin lesion images into multiple original skin lesion regions, and concatenate the image feature vector, text feature vector, and enhancement parameters corresponding to each original skin lesion region to generate a fusion feature vector for each original skin lesion region. The original skin lesion images are skin lesion images captured by the multispectral imager in the past time period. The first generation module is used to input the fused feature vector of each original skin lesion region into the diffusion model, generate a predicted enhanced image of each original skin lesion region through the diffusion model, generate the influence coefficient corresponding to each time step through the preset influence coefficient generation model, generate the divergence loss of the diffusion model according to the influence coefficient corresponding to each time step, the predicted enhanced image of each original skin lesion region and the preset divergence loss model, and generate the mean squared error loss of the diffusion model according to the preset mean squared error loss model. The second acquisition module is used to add the divergence loss and the mean squared error loss to generate the total loss of the diffusion model. It uses gradient descent to update the parameters of the diffusion model until the total loss of the diffusion model is less than the preset loss value. Then it stops updating the parameters of the diffusion model, stores the updated diffusion model, acquires the target skin lesion image, and divides the target skin lesion image into multiple target skin lesion regions through the image segmentation model. The target skin lesion image is the skin lesion image captured by the multispectral imager in the current time period. The second generation module is used to input each target lesion region into the image feature extraction network, generate an image feature vector for each target lesion region through the image feature extraction network, input the prompt text and extended description of each target lesion region into the text encoder, generate a text feature vector for each target lesion region through the text encoder, and concatenate the image feature vector, text feature vector and enhancement parameters corresponding to each target lesion region to generate a fused feature vector for each target lesion region. The third generation module is used to input the fused feature vector of each target lesion region into the updated diffusion model, generate the optimal enhanced image of each target lesion region through the updated diffusion model, and fuse the optimal enhanced images of each target lesion region through the fusion network to generate the enhanced target lesion image.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the lesion image enhancement method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the lesion image enhancement method as described in any one of claims 1 to 7.