An auxiliary diagnosis method and device based on a mixed burn wound image generation
Patent Information
- Application Number
- CN202611009978.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-07-08
AI Technical Summary
现有的图像生成手段生成复合烧伤创面图像无法精准建模混合深浅度分布,传统图像掩码生成仅支持单一深度区域生成,无法模拟真实混合深浅度烧伤特征
[0031]1. 本发明通过对真实烧伤区域掩码进行指标评价,提取不同烧伤深度与整体皮肤的占比、烧伤创面形态评估指标(BSI)以及烧伤创面人体部位类别三种先验知识,再通过多阶段人体有效解剖区域约束实现由深到浅的复合烧伤创面掩码生成,使生成的多样化掩码在空间分布、深度层次和边缘形态上更加符合真实烧伤病灶的复合深度形态特征,既保证了医学合理性又提高了合成样本的多样性。
Smart Images

Figure CN122530595B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical image processing, and in particular to an auxiliary diagnostic method and device based on the generation of mixed burn wound images. Background Technology
[0002] With the development of burn medicine, accurate assessment of burn wound depth is crucial for clinical treatment planning and patient prognosis. Traditional burn depth assessment relies primarily on clinicians' visual observation and experience, which suffers from high subjectivity, poor consistency, and a high rate of clinical misjudgment. Because burn image annotation requires meticulous pixel-by-pixel drawing by professional physicians and is constrained by privacy ethics, data acquisition is costly and data volume is scarce, thus limiting the generalization ability of existing models. While traditional data augmentation methods can expand the data volume, they cannot effectively increase sample diversity, especially the diversity of burn wound morphology.
[0003] Burn wounds typically exhibit high heterogeneity, irregular shape, mixed depths, and the coexistence of multiple injuries. Real-world burns often contain superficial second-degree, deep second-degree, and third-degree burns, or even normal skin, forming areas of mixed superficial and deep burns. Existing image generation methods cannot accurately model the mixed depth distribution of burn wounds, and traditional image masking only supports the generation of single-depth regions, failing to simulate the characteristics of real mixed superficial and deep burns. Summary of the Invention
[0004] To address the problems existing in the prior art, the present invention aims to provide an auxiliary diagnostic method and device based on the generation of hybrid burn wound images. This method can generate burn images that conform to the actual burn wound based on the characteristics of real burn images, effectively alleviating the problem of scarce burn medical image annotation data, thereby improving the accuracy of auxiliary diagnosis.
[0005] To achieve the above objectives, the present invention provides an auxiliary diagnostic method based on mixed burn wound images, comprising the following steps:
[0006] Step S110: Obtain a real burn wound image, and a real burn area mask and corresponding real diagnostic text annotation information based on the real burn wound image;
[0007] Step S120: Evaluate the actual burn area mask to obtain evaluation indicators including the proportion of different burn depths to the overall skin, burn wound morphology evaluation indicators, and burn wound human body part categories; determine the effective anatomical area of the human body as the initial constraint boundary according to the burn wound human body part categories; generate burn area masks of different burn depths in stages within the initial constraint boundary according to the principle of deep to shallow; after each stage is generated, the generated burn area is removed from the current anatomical constraint boundary, and the remaining area is used as the constraint boundary of the next stage; spatial mutual exclusion between stages is achieved through pixel distance thresholds to form a diversified burn area mask that conforms to the mixed characteristics of deep and shallow burn wounds.
[0008] Step S130: Using the diverse burn area mask and the real diagnostic text annotation information as input, the trained image generation model generates a synthetic burn wound image;
[0009] Step S140: Construct a hybrid burn dataset based on the real burn wound image and the synthetic burn wound image; perform feature modulation on the hybrid burn dataset to suppress the feature distribution difference between the synthetic burn wound image and the real burn wound image; train a semantic segmentation model using the modulated hybrid burn dataset; input the burn wound image to be diagnosed into the trained semantic segmentation model and output auxiliary diagnostic results.
[0010] Further, step S120 specifically includes:
[0011] Step S121: Evaluate the indicators of the real burn area mask to obtain: (1) the proportion of different burn depths to the overall skin, that is, the pixel area proportion of each burn depth category is counted from the real burn area mask; (2) burn wound morphology evaluation indicators. ,in, The theoretical pixel perimeter is the same as that of a circle with the same area as the burn wound. λ is the actual pixel perimeter of the burn wound; λ is the edge defect correction coefficient; (3) the human body part category of the burn wound, that is, the human body part where the burn wound is located is numbered;
[0012] Step S122: Extract the corresponding human body part mask according to the human body part category of the burn wound, and establish the initial constraint boundary by using the largest connected component as the effective anatomical region through connected component analysis;
[0013] Step S123: First, generate a burn region mask for the first burn depth within the initial constraint boundary, and determine the total number of restricted random walks based on the proportion of the current burn depth. The edge morphology of the generated area is controlled by the aforementioned burn wound morphology evaluation index; where α is the ratio of burn depth to total skin area. The total number of pixels in the effective anatomical region. This is a floor function; if the number of pixels in the generated burn area is less than... If so, then regenerate the mask;
[0014] Remove the generated burn area from the current constraint boundary, use the remaining area as the constraint boundary for the next stage, and generate a burn area mask for the next burn depth. Proceed step by step from deep to shallow until the burn depth areas are generated. Spatial mutual exclusion between adjacent stages is achieved through pixel distance thresholds.
[0015] Furthermore, the edge defect correction coefficient λ; no crack defects: Superficial small cracks and fine burrs: Extensive periphery ulceration and skin edge tearing: .
[0016] Furthermore, the feature modulation in step S140 specifically includes extracting semantic priors common to each scale using semantic anchors with cross-scale shared weights, and using the semantic priors to generate channel-level modulation parameters to dynamically modulate scale-specific features in order to suppress high-frequency artifacts introduced by synthesizing burn wound images.
[0017] Furthermore, in step S130, a burn depth grading prompting word library is built based on the real diagnostic text annotation information. Positive pathological description prompting words and negative rejection prompting words are configured according to superficial second-degree burns, deep second-degree burns, and third-degree burns, respectively, to guide the generation of synthetic burn wound images with burn depth annotation information.
[0018] Furthermore, it also includes step S150, performing a diagnostic assessment based on the auxiliary diagnostic results, and adjusting one or more of the sampling ranges of the proportion of different burn depths to the overall skin and the sampling ranges of the burn wound morphology assessment indicators in step S120 according to the feedback of the assessment results.
[0019] Furthermore, it also includes step S150: performing a diagnostic assessment based on the auxiliary diagnostic results, and adjusting the edge defect correction coefficient in step S120 according to the assessment results.
[0020] The present invention also provides an auxiliary diagnostic device based on mixed burn wound images, comprising:
[0021] The data acquisition module is used to acquire real burn wound images and their corresponding real burn area masks and real diagnostic text annotation information;
[0022] A diversified mask generation module is used to evaluate the real burn area mask and obtain evaluation indicators such as the proportion of different burn depths to the overall skin, burn wound morphology evaluation indicators, and burn wound human body part category. The module determines the initial constraint boundary based on the burn wound human body part category and generates burn area masks of different burn depths in stages according to the principle of deep to shallow. Each stage uses the remaining area after removing the generated area as the constraint boundary of the next stage. The stages are spatially mutually exclusive through pixel distance thresholds to form diversified burn area masks that conform to the mixed characteristics of real burn wounds with varying depths.
[0023] The burn image generation module is used to generate synthetic burn wound images by taking the diverse burn area masks and the real diagnostic text annotation information as inputs and using a trained image generation model.
[0024] The diagnostic output module is used to perform feature modulation on the synthetic burn wound image to suppress the feature distribution difference between it and the real burn wound image, so as to obtain an enhanced synthetic burn wound image; the real burn wound image and the enhanced synthetic burn wound image are fused to construct a hybrid burn dataset, a semantic segmentation model is trained, and auxiliary diagnostic results are output.
[0025] Furthermore, the mask generation module includes:
[0026] The indicator evaluation submodule is used to obtain three evaluation indicators: (1) the proportion of different burn depths to the overall skin; (2) the burn wound morphology evaluation indicators; and (3) the human body part categories of burn wounds.
[0027] The anatomical region determination submodule is used to determine the initial constraint boundary based on the human body part category of the burn wound. After generating a mask for each burn depth, the generated region is removed from the current constraint boundary to form the constraint boundary for the next stage.
[0028] A diversified mask generation submodule is used to determine the number of random walk steps within the constraint boundaries of each stage based on the proportion, control the edge morphology based on the morphology evaluation index, and generate burn area masks of different burn depths in stages from deep to shallow. Spatial mutual exclusion between adjacent stages is achieved through pixel distance thresholds.
[0029] Furthermore, the diagnostic output module is implemented by embedding a feature modulation module between the feature encoding module and the prediction decoding module of the semantic segmentation model. The feature modulation module uses cross-scale semantic anchors to dynamically modulate the encoded features at each scale to suppress high-frequency artifacts in the synthesized image.
[0030] The beneficial effects of this invention are as follows:
[0031] 1. This invention evaluates the masks of real burn areas by extracting three prior knowledge points: the proportion of different burn depths to the overall skin, the burn wound morphology assessment index (BSI), and the human body part category of the burn wound. Then, through multi-stage constraints on the effective anatomical region of the human body, it generates composite burn wound masks from deep to shallow. This makes the generated diverse masks more consistent with the composite depth and morphological characteristics of real burn lesions in terms of spatial distribution, depth level, and edge morphology, thus ensuring medical rationality and improving the diversity of synthetic samples.
[0032] 2. This invention uses an image generation model to generate controlled images from real burn images and diverse masks, synthesizing high-fidelity burn wound images with diagnostic text annotations. This effectively expands the scale and diversity of training data and alleviates the problem of scarce medical image annotation data for burns.
[0033] 3. This invention employs a semantic segmentation model consisting of three modules: feature encoding, feature modulation, and decoding prediction. At the feature level, it suppresses high-frequency artifacts in the synthesized image through cross-scale feature modulation, alleviating the distribution difference between the real image and the generated image, and finally outputs a pixel-level semantic segmentation mask, which significantly improves the segmentation accuracy and generalization ability under mixed data training.
[0034] 4. This invention employs a multi-category semantic segmentation approach to diagnose burn depth. The semantic segmentation model categorizes each pixel in the input image, with different pixel categories directly corresponding to clinically defined burn depths (superficial second-degree, deep second-degree, and third-degree). The pixel-level prediction results output by the segmentation model represent the burn depth diagnosis results corresponding to the wound area. The segmentation mask output by the model not only delineates the burn wound boundary but also directly assigns a severity level to each wound, achieving integrated segmentation and diagnosis. No additional classification module is required; the model outputs the wound location, wound percentage, and burn depth of each wound in a single step, significantly improving the completeness and efficiency of clinical diagnosis. Attached Figure Description
[0035] Figure 1 This is a flowchart of an auxiliary diagnostic method for burn wounds based on generative data augmentation, according to Embodiment 1 of the present invention.
[0036] Figure 2 for Figure 1 Flowchart for generating diverse burn area masks in step S120;
[0037] Figure 3 for Figure 1 A schematic diagram of the three-module architecture of the semantic segmentation model in step S140;
[0038] Figure 4 This is a schematic diagram of a burn wound auxiliary diagnostic device based on generative data augmentation according to Embodiment 2 of the present invention;
[0039] Figure 5 A visualization of t-SNE before and after modulation of a deep second-dimensional dataset;
[0040] Figure 6 A heatmap visualization analysis of deep features, shallow features, and fused features after decoding in a semantic segmentation model of burn images. Detailed Implementation
[0041] To enable those skilled in the art to better understand the present application, the technical solutions in this embodiment will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.
[0042] Example 1
[0043] An auxiliary diagnostic method based on hybrid burn wound image generation is applicable to burn wound depth assessment and auxiliary diagnosis. It can be executed by a burn wound auxiliary diagnostic device, which can be implemented in hardware and / or software and configured in a computer device, such as a medical terminal or server. The method includes the following steps:
[0044] S110. Obtain a real burn wound image, and a real burn area mask and corresponding real diagnostic text annotation information based on the real burn wound image.
[0045] The real burn wound image refers to the wound image data of burn patients obtained by medical institutions through image acquisition equipment, which is used to describe the appearance and structural characteristics of the burn wound, and may include at least one of the following: color information, texture information and morphological information of the wound.
[0046] The aforementioned realistic burn area mask refers to a semantic segmentation mask obtained by professional burn surgeons through pixel-by-pixel annotation of real burn wound images. Different pixel values represent different burn depth categories, including at least one of the following: normal skin area, superficial second-degree burn area, deep second-degree burn area, and third-degree burn area. This mask is used both for subsequent metric evaluation to generate diverse masks and as spatial conditional input during the training of the image generation model.
[0047] The aforementioned authentic diagnostic text annotation information refers to structured text prompts obtained by professional burn surgeons annotating real burn wound images. These prompts are used to characterize the spatial distribution and pathological features of different burn depth regions within the wound. Specifically, each burn depth region includes positive and negative prompts: positive prompts include keywords related to the burn depth grade and descriptive terms of typical pathological morphology, such as superficial second-degree burns and significant local swelling; negative prompts are used to exclude non-authentic pathological features that may appear during the formation process, such as eschar and carbonization in second-degree burn areas, which are negative prompts.
[0048] S120. Evaluate the actual burn area mask using the following three evaluation indicators: (1) the proportion of different burn depths to the overall skin; (2) the evaluation index of burn wound morphology; (3) the human body part category of the burn wound; generate diverse burn area masks based on the three evaluation indicators through a controlled random generation process; specifically, the following steps are included:
[0049] S121. The mask of the real burn area is evaluated by the following three evaluation indicators: (1) the proportion of different burn depths to the whole skin; (2) the evaluation index of burn wound morphology; (3) the human body part category of burn wound.
[0050] The evaluation of the ratio of different burn depths to the total skin area refers to the calculation of the pixel area ratio of each burn depth category (superficial second degree, deep second degree, and third degree) to the total skin area from the actual burn area mask based on clinical burn statistics. Specifically, the area of each burn depth category is determined by counting the number of pixels of each category in the burn area mask, and the pixel area ratio of each burn depth category to the total skin area is calculated by comparing it with the number of pixels of the total skin area.
[0051] The burn wound morphology assessment index (BSI, Border Smoothness Index) is mentioned. ,in, The theoretical pixel perimeter is the same as that of a circle with the same area as the burn wound. λ represents the actual pixel perimeter of the burn wound; λ is the edge defect correction coefficient, reflecting the broken or fragmented state of the burn wound edge, and the absence of broken edges. Superficial small cracks and fine burrs: Extensive periphery ulceration and skin edge tearing: The actual perimeter of the burn wound is determined by counting the number of pixels at the boundary of the burn area mask.
[0052] The category of human body part of the burn wound refers to the numbering of the human body part where the burn wound is located, such as limbs or trunk;
[0053] S122. Determine the initial mask generation boundary for human body parts: Perform connected component analysis on the mask of the real burn area, use the 8-neighbor connected component algorithm to identify the human body part region, select the connected component with the largest area as the effective anatomical region, extract the human body part category number, establish the position constraint range, lock the initial mask generation boundary of the human body part, form the first effective anatomical region, and avoid the generated mask from going out of bounds to non-effective anatomical regions.
[0054] Specifically, connected component analysis (CFI) is performed on the human body part mask to obtain the effective part regions. CFI refers to the image processing operation of grouping and labeling spatially adjacent pixels with the same pixel value in a binary mask. By selecting the largest connected region as the effective part region through CFI, noise or fragmented regions that may exist during segmentation can be removed, ensuring that the generation process is only carried out within a reasonable anatomical range. The aforementioned effective part regions represent the pixel range of the human target region after separation from the background, constituting the positional constraint parameters.
[0055] The specific method for analyzing the 8-neighbor connected components is as follows:
[0056] Masking of human body parts Perform connected component analysis to obtain the effective region. Connected component analysis (CBI) refers to an image processing operation that groups and labels spatially adjacent pixels with the same pixel value in a binary mask. For example, the effective region can be obtained using the following set of formulas. :
[0057] ,
[0058] ,
[0059] ,
[0060] ,
[0061] ;
[0062] in, This is the input binary mask for the human body part. The labeling results obtained from connected component analysis. For the index of the connected region, For pixel coordinates, Represents pixels Belongs to the A connected region, For the first A set of pixels in a connected region This indicates the index of the connected region with the largest area. This represents the largest connected region selected. The above-mentioned effective region... This represents the pixel range of the human target area after separation from the background, constituting the positional constraint parameters to ensure that the generated burn area is only within a reasonable human anatomical range.
[0063] S123. Generation of Diverse Burn Region Masks: Within the first effective anatomical area determined in step S122, based on the three real burn image mask evaluation indicators obtained in step S121, a burn region mask for the first burn depth, such as deep second-degree burns, is generated through restricted random walks. Specifically, the total number of walk steps is determined according to the ratio of deep second-degree burn depth to the total skin area, and the total number of restricted random walk steps is... Where α represents the burn depth as a percentage of the total skin area. The total number of pixels in the effective anatomical region. This is a floor function; if the number of pixels in the generated burn area is less than... If so, then regenerate the mask;
[0064] Next, the boundaries for generating the human body part mask are redefined: after removing the burn area of the first burn depth within the first effective anatomical region, it becomes the second effective anatomical region; specifically, secondary seeds are randomly sampled within the second effective anatomical region to form a burn area mask of the second burn depth, which is shallower than the first burn depth, such as using a shallow second degree burn depth, to generate a composite burn wound mask; spatial mutual exclusion constraints are achieved through pixel distance thresholds to suppress the overlap between the first and second effective anatomical regions; the generation of the composite burn wound mask generally follows the principle of from deep to shallow, forming a diverse burn mask that conforms to the real burn wound;
[0065] The pixel distance is a geometric metric (such as Euclidean distance) that describes the spatial proximity of two pixels in the image coordinate system. In this application, the pixel distance threshold for achieving spatial mutual exclusion constraint specifically means that the pixel distance threshold between the burn area of the first burn depth and the burn area of the second burn depth is greater than or equal to 1. That is, spatial mutual exclusion constraint is achieved through the pixel value of the pixel distance threshold to suppress the overlap between the first effective anatomical area and the second effective anatomical area.
[0066] The aforementioned diverse burn mask generation extracts evaluation indicators from real burn area masks, using these as prior knowledge to guide mask generation. Through multi-stage constraints on the effective anatomical region of the human body, it avoids medically unreasonable situations such as positional drift or exceeding the human body contour in the generated burn mask. It also better matches the complex depth and morphological characteristics of burn lesions in real data, ensuring high fidelity of the generated images for subsequent burn image generation, making them more consistent with real burn images, and significantly improving the diversity of synthetic samples.
[0067] S130. Based on the real burn wound image and its corresponding real diagnostic text annotation information, train the image generation model to obtain the training weight parameters of the image generation model; using the diverse burn area mask obtained in step S120 and the real diagnostic text annotation information as input, use the image generation model to generate a synthetic burn wound image.
[0068] Specifically, the image generation model employs a conditional generation architecture based on spatial adaptive normalization. This architecture dynamically adjusts the spatial distribution of generated features according to the input semantic segmentation mask through a spatial adaptive normalization mechanism, thereby achieving controlled image generation based on the semantic segmentation mask. For example, the image generation model uses a SPADE (Spatially-Adaptive Denormalization) conditional generation architecture, which consists of multiple residual blocks, each containing a spatial adaptive normalization layer and a convolutional layer.
[0069] Training an image generation model includes the following steps:
[0070] Training data pairs are constructed based on real burn wound images and their corresponding real diagnostic text annotations. The real diagnostic text annotations are transformed to obtain a semantic segmentation mask, which serves as the spatial conditional input to the conditional generator. Specifically, different burn depth regions in the real diagnostic text annotations are mapped to different label values, forming a semantic segmentation mask of the same size as the real burn wound images. Each pixel value in this mask represents the burn depth category at the corresponding location, providing spatial structure information for the generator.
[0071] The training data is used to train the input conditional generator model, resulting in training weight parameters. The generator employs a spatial adaptive normalization mechanism, dynamically adjusting the mean and variance of features at each layer of the network based on diverse burn region masks to ensure strict alignment of the spatial structure of the generated image with the diverse burn region masks. During this process, a weighted overall loss function consisting of adversarial loss, feature matching loss, perceptual loss, and KL divergence loss is used to jointly optimize the generator and encoder, thereby improving the realism of the generated image while maintaining the stability of the training process.
[0072] The adversarial loss employs a Hinge loss form combined with a multi-scale discriminator structure to distinguish between real and fake images in both the original resolution image and its multi-level downsampled versions, thereby improving training stability and the quality of generated high-resolution images. The feature matching loss guides the generator to learn a more stable multi-scale feature distribution by constraining the consistency of feature representations in the intermediate layers of the discriminator between real and generated images. The perceptual loss, based on a pre-trained VGG visual feature encoding network, constrains the consistency of high-level semantic features between generated and real images, further enhancing the realism of the images. To achieve multimodal synthesis and style-guided image generation, an image encoder is introduced to map the input image into latent features, and KL divergence constraints are applied to make the latent space approximate a standard Gaussian prior distribution. The above loss terms are weighted and summed according to empirical weights to form the joint optimization objective of the generator and encoder. The training process continues until the model converges, ultimately yielding training weight parameters with image generation capabilities.
[0073] In the image generation stage, a graded prompt word library is built based on real diagnostic text annotation information. Positive pathological description prompt words and negative rejection prompt words are configured according to superficial second-degree burns, deep second-degree burns, and third-degree burns, respectively. The diverse burn area mask generated in step S120 is used as spatial condition input, and the prompt words corresponding to the burn depth are used to guide the image generation model to generate a synthetic burn wound image with diagnostic text annotation information.
[0074] S140. Construct a hybrid burn dataset based on the real burn wound image and the synthetic burn wound image; perform feature modulation on the hybrid burn dataset to suppress the feature distribution difference between the synthetic burn wound image and the real burn wound image, obtaining a modulated hybrid burn dataset; train a semantic segmentation model using the modulated hybrid burn dataset; input the burn wound image to be diagnosed into the trained semantic segmentation model, and output auxiliary diagnostic results. This step includes a hybrid burn dataset construction sub-step, a feature modulation sub-step, a semantic segmentation model training sub-step, and a burn intelligent diagnosis sub-step.
[0075] Step S141: Construction of Hybrid Burn Dataset: Real burn images and synthetic burn wound images are fused to construct a hybrid burn dataset. The training dataset contains both real and synthetic burn wound images and is used to train the semantic segmentation model. The validation dataset contains only real burn wound images and is used to evaluate the performance of the semantic segmentation model.
[0076] Optionally, the validation dataset comprises 20% of the total real data in the mixed burn dataset, the training dataset comprises 80% of the total real data, and synthetic burn wound images are used only to supplement the training dataset. This division ratio can be adjusted according to the actual data scale and diagnostic accuracy requirements, for example, it can be set to 15% / 85% or 25% / 75%, etc., and this application does not impose specific limitations. It should be noted that synthetic burn wound images only participate in the construction of the training dataset and do not participate in the construction of the validation dataset to ensure the objectivity and clinical reliability of the validation results. All samples in the validation dataset are real burn wound images annotated by professional physicians, which can objectively reflect the diagnostic performance of the model in actual clinical scenarios.
[0077] Step S142, Feature Modulation, refers to feature modulation of the constructed hybrid burn dataset. This is achieved by embedding a feature modulation module between the feature encoding and decoding of the semantic segmentation model. The feature modulation module extracts common semantic priors from the encoded features at each scale through semantic anchors with shared weights across scales. It then uses these semantic priors to generate channel-level modulation parameters to dynamically modulate the scale-specific features at each scale, ensuring that the modulated features are unaffected by specific-scale artifacts in the synthesized image, resulting in a modulated hybrid burn dataset. Although the generative model trained in step S130 can synthesize high-quality composite burn wound images at the visual level, the statistical distribution of the generated image and the real image in the feature space still differs, inevitably containing high-frequency artifacts in specific frequency bands, leading to a significant distribution difference between the real and generated data. To eliminate this difference, this invention introduces a burn image semantic segmentation network with feature modulation. By explicitly aligning and adaptively modulating the features of the real and generated images, the generated image more closely approximates the distribution characteristics of the real image at the feature level.
[0078] The feature modulation module is used to calculate the feature distribution difference between the multi-scale feature set of the real burn wound image and the multi-scale feature set of the generated burn wound image, and generate channel-by-channel modulation parameters based on the feature distribution difference. The modulation parameters are then applied to the multi-scale feature set of the generated burn wound image, and the mean and variance of the features are adjusted layer by layer through spatial adaptive normalization. The modulation process starts from the lowest resolution level and gradually progresses to higher resolution levels, finally outputting the enhanced synthetic burn wound image via the decoder. Figure 5As shown, the t-SNE visualizations of the deep 2D dataset before and after modulation verify that the feature modulation module can effectively bring the feature distributions of the two domains closer together. t-SNE, as a nonlinear dimensionality reduction method, can fully reveal the relative distance and overlap between the two classes of samples in the feature space. Before modulation, the feature points of the real and generated images show a clear separation trend in the t-SNE embedding space. Real samples cluster in the left region of the coordinate system, while generated samples are mainly distributed on the right, with a large inter-class distance and a small overlap between the two distributions. This directly verifies that there are indeed significant domain differences between the generated and real images in terms of statistical distribution and deep semantic features. After the proposed feature modulation module, the feature point distributions of the real and generated images change significantly. The two classes of samples are no longer clearly separated clusters, but rather intertwined and uniformly mixed, making it difficult for the domain classifier to effectively distinguish them. This indicates that the feature modulation module successfully suppresses high-frequency artifacts in the generated image and brings the feature representations of the real and generated data closer to the same embedding space, achieving effective domain alignment.
[0079] Step S143: Semantic segmentation model training
[0080] This application employs a three-module semantic segmentation architecture based on self-supervised visual representation learning to train a hybrid burn dataset. For example... Figure 3 As shown, the burn image semantic segmentation model includes three core modules:
[0081] Feature encoding module: used to extract feature maps of the real burn wound image and the synthetic burn wound image at different semantic levels, and output the multi-scale feature set of the real burn wound image and the multi-scale feature set of the generated burn wound image.
[0082] Feature modulation module: used to calculate the feature distribution difference between the multi-scale feature set of the real burn wound image and the multi-scale feature set of the generated burn wound image, and generate channel-by-channel modulation parameters based on the feature distribution difference;
[0083] (3) Decoding and Prediction Module: The modulated multi-scale features are decoded step by step to generate a pixel-level segmentation mask. Decoding starts with the lowest resolution features and gradually recovers the spatial size through upsampling. At each decoding level, the current deep features are fused with the corresponding resolution of the modulated shallow features. During fusion, a spatial gating mask is generated using deep semantic features, and each spatial location in the shallow features is adaptively filtered—regions with high gating values retain details, while regions with low gating values suppress artifacts. After multi-level decoding and fusion, the resolution and semantic details are gradually refined to the size of the input image. Finally, a classification convolutional layer is used to predict the burn depth category of each pixel, outputting a pixel-level semantic segmentation mask.
[0084] To visually verify the effectiveness of the semantic segmentation model for burn images in suppressing background noise and extracting clean detail features, a heatmap visualization analysis was performed on the deep features, shallow features, and fused features after decoding in the model. The results are as follows: Figure 6 As shown in the visualization, while deep features can locate the core area of burn lesions and possess strong high-level semantic information, their spatial resolution is low, and the lesion edges exhibit significant blurring. Conversely, directly extracted shallow features, although containing rich texture and edge details, are filled with background noise and irrelevant artifacts. Directly using traditional indiscriminate stitching and fusion will inevitably interfere with the final prediction results. In contrast, the feature map output by the gated decoder shows a significant optimization effect. The processed features not only maintain consistency with deep features in lesion localization but also have clearer and sharper boundaries, presenting a highly pure multi-scale semantic representation. With the assistance of the gated decoder, the network can accurately fit various complex and varied burn boundaries.
[0085] Through the above three-module architecture, the semantic segmentation model suppresses artifact generation at the feature level through cross-scale modulation and weakens background noise at the decoding level through spatial gating. It works together to alleviate the distribution difference between real and generated images from two levels, providing a reliable technical guarantee for the accurate segmentation of burn wounds.
[0086] The overall process of semantic segmentation model mask prediction is as follows: The input burn wound image is fed into a frozen self-supervised pre-trained backbone network to extract multi-scale encoded features; then, the feature modulation module performs cross-scale modulation on the features at each scale to suppress high-frequency artifacts introduced by the synthetic image and outputs the modulated multi-scale features; finally, the decoding prediction module decodes step by step from the lowest resolution features, and at each level, the current deep features are upsampled and spatially gated and fused with the modulated shallow features of the corresponding resolution to gradually restore the spatial resolution to the size of the input image; finally, the classification convolutional layer predicts the burn depth category of each pixel and outputs a pixel-level semantic segmentation mask with the same resolution as the input image.
[0087] New, real-world burn wound images to be diagnosed are input into a trained semantic segmentation model for semantic segmentation and depth diagnosis of the burn area. The diagnostic result refers to the pixel-level burn depth category prediction map output by the semantic segmentation model, where each pixel is classified into one of the following categories: normal skin, superficial second-degree burn, deep second-degree burn, or third-degree burn. This segmentation mask simultaneously delineates the burn wound boundary and assigns the injury severity level, achieving integrated segmentation and diagnosis—pixel-level segmentation results are directly converted into burn depth diagnostic conclusions without the need for additional classification modules, outputting the wound location, wound area proportion, and burn depth of each wound in one go.
[0088] S144. Intelligent burn diagnosis: Construct a hybrid burn dataset based on the real burn wound image and the synthetic burn wound image; train the semantic segmentation model using the hybrid burn dataset; input the burn wound image to be diagnosed into the trained semantic segmentation model, and output auxiliary diagnostic results.
[0089] A hybrid burn dataset is constructed by fusing real and synthetic burn wound images. This hybrid burn dataset is then divided into a training dataset and a validation dataset. The training dataset contains both real and synthetic burn wound images and is used to train the semantic segmentation model. The validation dataset contains only real burn wound images and is used to evaluate the performance of the semantic segmentation model.
[0090] Optionally, the validation dataset comprises 20% of the total real data in the mixed burn dataset, the training dataset comprises 80% of the total real data, and synthetic burn wound images are used only to supplement the training dataset. This division ratio can be adjusted according to the actual data scale and diagnostic accuracy requirements. It should be noted that synthetic burn wound images are only used in the construction of the training dataset and not in the construction of the validation dataset to ensure the objectivity and clinical reliability of the validation results.
[0091] The overall process of semantic segmentation model mask prediction is as follows: A new, real burn wound image to be diagnosed is input into the trained semantic segmentation model, which performs semantic segmentation and depth classification of the burn region, and outputs auxiliary diagnostic results. Here, the new, real burn wound image to be diagnosed refers to a burn patient's wound image that was not used during training and is intended for burn depth assessment. The diagnostic result refers to the pixel-level burn depth category prediction map output by the semantic segmentation model.
[0092] S150. Perform a diagnostic assessment based on the auxiliary diagnostic results and output the assessment results; adjust the parameters based on the assessment results.
[0093] The evaluation results refer to the performance metrics obtained by quantifying the consistency between the diagnostic results and the ground truth annotations. These metrics may include at least one of the following: mean intersection-over-union (mIoU), Dice index coefficient, pixel accuracy, and F1 score for each category. The evaluation process uses real burn wound images from the validation dataset as evaluation samples. By calculating the degree of overlap between the model's prediction results and the annotation results from professional physicians, the diagnostic performance of the model is objectively quantified.
[0094] In one optional implementation, the evaluation results can be used to optimize and adjust the mask generation parameters in step S120 and the parameters of the semantic segmentation model in step S140. Specifically, based on the evaluation feedback, the sampling range of different burn depths and their proportions to the overall skin, the sampling range of the burn wound morphology assessment index BSI, the edge defect correction coefficient λ, and the modulation intensity parameters in the feature modulation module are adjusted to make the generated synthetic burn wound image more closely resemble the actual clinical scenario in terms of structural features and feature distribution, thereby further improving the diagnostic accuracy of the semantic segmentation model.
[0095] For example, if the evaluation results show that the model's segmentation accuracy in the deep second-degree burn area is low, the generation ratio of deep second-degree burn masks can be increased in the burn mask generation step. By adjusting the sampling distribution of different burn depths and the proportion of the overall skin, the synthetic dataset can contain more deep second-degree burn samples. At the same time, the modulation intensity of the feature modulation module corresponding to the deep second-degree feature channel can be appropriately increased to enhance the model's ability to recognize this category.
[0096] The overall technical solution of this embodiment can be summarized as follows: Obtain real burn area masks and real diagnostic text annotations from real burn wound images; evaluate the real masks using metrics, and after obtaining three evaluation metrics, generate diverse masks for composite burn wounds from deep to shallow through multi-stage constraints on effective anatomical regions of the human body; train an image generation model using real data, and generate synthetic burn images using the diverse masks and text annotations as input; employ a semantic segmentation model consisting of three modules: feature encoding, feature modulation, and decoding prediction, suppressing generated artifacts through cross-scale feature modulation, and progressively restoring resolution and outputting pixel-level semantic segmentation masks through a spatially gated decoder; fuse the synthetic images with real images to construct a hybrid dataset to train the model; and perform feedback optimization based on diagnostic evaluation results. This technical solution, through a generative data augmentation end-to-end approach encompassing metric evaluation, multi-stage constraint mask generation, image synthesis, and three-module semantic segmentation, effectively alleviates the scarcity of burn medical image annotation data, suppresses the distribution shift of generated data through the adaptive mechanism of feature modulation, and significantly improves the accuracy and robustness of burn wound semantic segmentation.
[0097] Example 2: An auxiliary diagnostic device based on generative data augmentation, corresponding to the method described in Example 1, such as... Figure 4 As shown, it includes the following functional modules:
[0098] The data acquisition module performs the function of step S110 in Embodiment 1, acquiring a real burn wound image, as well as a real burn area mask and corresponding real diagnostic text annotation information based on the real burn wound image. This module receives burn wound image data from a medical image acquisition device and reads the corresponding mask and text annotation from the annotation database.
[0099] The diversified burn mask generation module performs the function of step S120 in Embodiment 1, evaluates the actual burn area mask, and obtains three evaluation indicators: the proportion of different burn depths to the overall skin, burn wound morphology assessment indicators, and burn wound human body part categories. Based on these three evaluation indicators, it generates diversified burn area masks through a multi-stage effective human anatomical region constraint and controlled random generation process. This module internally includes an indicator evaluation submodule, an anatomical region determination submodule, and a mask generation submodule.
[0100] The burn image generation module performs the function of step S130 in Embodiment 1. It trains the image generation model based on the real burn wound image and its corresponding real diagnostic text annotation information to obtain the training weight parameters of the image generation model. Using the diverse burn region mask and the real diagnostic text annotation information as input, it generates a synthetic burn wound image using the image generation model. This module internally includes a model training submodule, a prompt word library submodule, and an image reasoning submodule.
[0101] The diagnostic output module performs the function of step S140 in Embodiment 1, constructing a hybrid burn dataset based on the real burn wound image and the synthetic burn wound image; performing feature modulation on the hybrid burn dataset to suppress the feature distribution difference between the synthetic burn wound image and the real burn wound image, obtaining a modulated hybrid burn dataset; training a semantic segmentation model using the modulated hybrid burn dataset and outputting auxiliary diagnostic results. The semantic segmentation model adopts a three-module architecture consisting of a feature encoding module, a feature modulation module, and a decoding prediction module. The feature modulation module is embedded between feature encoding and decoding, and the decoding prediction module outputs pixel-level burn depth diagnostic results, achieving integrated segmentation and diagnosis. It performs the function of step S140 in Embodiment 1, constructing a hybrid burn dataset based on the real burn wound image and the synthetic burn wound image; performing feature modulation on the hybrid burn dataset to suppress the feature distribution difference between the synthetic burn wound image and the real burn wound image, obtaining a modulated hybrid burn dataset; training a semantic segmentation model using the modulated hybrid burn dataset and outputting auxiliary diagnostic results. The semantic segmentation model adopts a three-module architecture consisting of a feature encoding module, a feature modulation module, and a decoding prediction module. The feature modulation module is embedded between the feature encoding and decoding modules, and the decoding prediction module outputs pixel-level burn depth diagnostic results, realizing integrated segmentation and diagnosis.
[0102] The parameter feedback optimization module is used to perform the function of step S150 in Embodiment 1, perform diagnostic evaluation based on the auxiliary diagnostic results and output the evaluation results, and adjust one or more of the following parameters in the mask generation module: the sampling range of different burn depths and the proportion of the whole skin, the sampling range of the burn wound morphology evaluation index BSI, the edge defect correction coefficient λ, and the modulation intensity parameter in the feature modulation module, according to the evaluation results.
[0103] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. An auxiliary diagnostic method based on mixed burn wound image generation, characterized in that, The steps include the following: Step S110: Obtain a real burn wound image, and a real burn area mask and corresponding real diagnostic text annotation information based on the real burn wound image; Step S120: Evaluate the actual burn area mask to obtain evaluation indicators including the proportion of different burn depths to the overall skin, burn wound morphology evaluation indicators, and burn wound human body part categories; determine the effective anatomical area of the human body as the initial constraint boundary according to the burn wound human body part categories; generate burn area masks of different burn depths in stages within the initial constraint boundary according to the principle of deep to shallow; after each stage is generated, the generated burn area is removed from the current anatomical constraint boundary, and the remaining area is used as the constraint boundary of the next stage; spatial mutual exclusion between stages is achieved through pixel distance thresholds to form a diversified burn area mask that conforms to the mixed characteristics of deep and shallow burn wounds. Step S130: Using the diverse burn area mask and the real diagnostic text annotation information as input, the trained image generation model generates a synthetic burn wound image; Step S140: Construct a hybrid burn dataset based on the real burn wound image and the synthetic burn wound image; perform feature modulation on the hybrid burn dataset to suppress the feature distribution difference between the synthetic burn wound image and the real burn wound image; train a semantic segmentation model using the modulated hybrid burn dataset; input the burn wound image to be diagnosed into the trained semantic segmentation model and output auxiliary diagnostic results; Specifically, step S120 includes: Step S121: Evaluate the indicators of the real burn area mask to obtain: (1) the proportion of different burn depths to the overall skin, that is, the pixel area proportion of each burn depth category is counted from the real burn area mask; (2) burn wound morphology evaluation indicators. ,in, The theoretical pixel perimeter is the same as that of a circle with the same area as the burn wound. λ is the actual pixel perimeter of the burn wound; λ is the edge defect correction coefficient; (3) the human body part category of the burn wound, that is, the human body part where the burn wound is located is numbered; Step S122: Extract the corresponding human body part mask according to the human body part category of the burn wound, and establish the initial constraint boundary by using the largest connected component as the effective anatomical region through connected component analysis; Step S123: First, generate a burn region mask for the first burn depth within the initial constraint boundary, and determine the total number of restricted random walks based on the proportion of the current burn depth. The edge morphology of the generated area is controlled by the aforementioned burn wound morphology evaluation index; where α is the ratio of burn depth to total skin area. The total number of pixels in the effective anatomical region. This is a floor function; if the number of pixels in the generated burn area is less than... If so, then regenerate the mask; Remove the generated burn area from the current constraint boundary, use the remaining area as the constraint boundary for the next stage, and generate a burn area mask for the next burn depth. Proceed step by step from deep to shallow until the burn depth area is generated. Spatial mutual exclusion between adjacent stages is achieved through pixel distance threshold. Wherein, the edge defect correction coefficient λ; no crack defect: Superficial small cracks and fine burrs: Extensive periphery ulceration and skin edge tearing: ; Specifically, the feature modulation in step S140 includes extracting semantic priors common to each scale using semantic anchors with cross-scale shared weights, and using the semantic priors to generate channel-level modulation parameters to dynamically modulate scale-specific features in order to suppress high-frequency artifacts introduced by synthesizing burn wound images.
2. The diagnostic method according to claim 1, characterized in that, In step S130, a burn depth grading prompting word library is built based on the real diagnostic text annotation information. Positive pathological description prompting words and negative rejection prompting words are configured for superficial second-degree burns, deep second-degree burns, and third-degree burns, respectively, to guide the generation of synthetic burn wound images with burn depth annotation information.
3. The diagnostic method according to claim 1, characterized in that, It also includes step S150, performing a diagnostic assessment based on the auxiliary diagnostic results, and adjusting one or more of the sampling ranges of the proportion of different burn depths to the overall skin and the sampling ranges of the burn wound morphology assessment indicators in step S120 according to the feedback of the assessment results.
4. The diagnostic method according to claim 1, characterized in that, It also includes step S150, performing a diagnostic assessment based on the auxiliary diagnostic results, and adjusting the edge defect correction coefficient in step S120 based on the assessment results.
5. An auxiliary diagnostic device based on mixed burn wound image generation, used to implement the auxiliary diagnostic method based on mixed burn wound image generation as described in any one of claims 1 to 4, characterized in that, include: The data acquisition module is used to acquire real burn wound images and their corresponding real burn area masks and real diagnostic text annotation information; A diversified mask generation module is used to evaluate the mask of the real burn area and obtain evaluation indicators of the proportion of different burn depths to the whole skin, burn wound morphology evaluation indicators, and burn wound body part category. The initial constraint boundary is determined according to the category of human body part of the burn wound. Burn area masks of different burn depths are generated in stages according to the principle of deep to shallow. The remaining area after removing the generated area is used as the constraint boundary of the next stage. The stages are spatially mutually exclusive through the pixel distance threshold, forming a diversified burn area mask that conforms to the characteristics of mixed deep and shallow burn wounds. The burn image generation module is used to generate synthetic burn wound images by taking the diverse burn area masks and the real diagnostic text annotation information as inputs and using a trained image generation model. The diagnostic output module is used to perform feature modulation on the synthetic burn wound image to suppress the feature distribution difference between it and the real burn wound image, so as to obtain an enhanced synthetic burn wound image; the real burn wound image and the enhanced synthetic burn wound image are fused to construct a hybrid burn dataset, a semantic segmentation model is trained, and auxiliary diagnostic results are output.
6. The auxiliary diagnostic device according to claim 5, characterized in that, The mask generation module includes: The indicator evaluation submodule is used to obtain three evaluation indicators: (1) the proportion of different burn depths to the overall skin; (2) the burn wound morphology evaluation indicators; and (3) the human body part categories of burn wounds. The anatomical region determination submodule is used to determine the initial constraint boundary based on the human body part category of the burn wound. After generating a mask for each burn depth, the generated region is removed from the current constraint boundary to form the constraint boundary for the next stage. A diversified mask generation submodule is used to determine the number of random walk steps within the constraint boundaries of each stage based on the proportion, control the edge morphology based on the morphology evaluation index, and generate burn area masks of different burn depths in stages from deep to shallow. Spatial mutual exclusion between adjacent stages is achieved through pixel distance thresholds.
7. The auxiliary diagnostic device according to claim 5, characterized in that, The diagnostic output module is implemented by embedding a feature modulation module between the feature encoding module and the prediction decoding module of the semantic segmentation model. The feature modulation module uses cross-scale semantic anchors to dynamically modulate the encoded features at each scale to suppress high-frequency artifacts in the synthesized image.
Citation Information
Patent Citations
Generating images using one or more neural networks
CN114365185A
Multi-task image segmentation method and system for burn injury assessment
CN118229708A