Pulmonary tuberculosis lesion map segmentation system based on symmetric similarity and dynamic sample amplification

Through the sternum inhibition network and the entropy-driven sample amplification module, combined with symmetric similarity detection, the problem of high model generalization and misjudgment rate in tuberculosis lesions is solved, and efficient and accurate lesion segmentation effect is achieved.

CN120495179APending Publication Date: 2025-08-15YANGTZE DELTA REGION INST OF UNIV OF ELECTRONICS SCI & TECH OF CHINE (HUZHOU)
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510479822.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The prior art has problems with poor generalization of the model and high misjudgment rate in the segmentation of tuberculosis lesions, especially due to poor segmentation effect caused by rib occlusion and image equipment differences, and the traditional self-attention mechanism fails to effectively utilize human structural characteristics.

Method used

The sternum inhibition network module is used to remove sternum images, combined with entropy-driven dynamic amplification of sample number and the tuberculosis lesion detection module based on symmetric similarity, the learning ability and applicability of the model are improved by dynamically adjusting the sample number and amplification method.

Benefits of technology

It significantly improves the accuracy and recall of tuberculosis lesions, enhances the generalization and scope of application of the model, and can accurately segment the lesion areas under different equipment and imaging conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495179A_ABST
    Figure CN120495179A_ABST
Patent Text Reader

Abstract

The invention discloses a pulmonary tuberculosis focus map segmentation system based on symmetric similarity and dynamic sample amplification, which comprises a sternum suppression network module, an entropy-driven sample number dynamic amplification module, a pulmonary tuberculosis focus detection module and a sternum suppression network module, and is used for removing a sternum image in a chest X-ray map; the entropy-driven sample dynamic amplification module is used for carrying out dynamic amplification on the number of sample images based on the complexity of the image and the existing training state of the model for a single image; specifically, for each image, an information entropy is calculated for probability distribution output by a model of each training round; in the early stage of training, the number of training image samples is amplified slightly, and the number of amplified samples is increased along with model training; and the pulmonary tuberculosis focus detection module is used for designing a pulmonary tuberculosis focus detection network model by using the target detector and the symmetric similarity of the pulmonary tuberculosis images, and segmenting a tuberculosis focus map from the X-ray map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image segmentation, and in particular relates to a tuberculosis lesion image segmentation system based on symmetric similarity and dynamic sample amplification. Background Art

[0002] Tuberculosis (TB) is an infectious disease caused by Mycobacterium tuberculosis. Its main clinical symptoms include cough, low-grade fever in the afternoon, night sweats, and weight loss. TB is preventable and controllable, but early detection and response are crucial. Therefore, timely and rapid analysis of TB patients and their close contacts is crucial for accurately identifying at-risk populations, developing effective response strategies, limiting transmission, and reducing the cost of subsequent prevention and control.

[0003] Imaging-assisted analysis technologies for pulmonary tuberculosis mainly include CT and X-rays. Currently, X-rays remain the preferred conventional technology for pulmonary tuberculosis analysis, with advantages such as low cost and low radiation exposure. Because manual reading of X-rays is time-consuming, prone to missed detections, and requires consideration of the reader's experience, the development of automatic segmentation and detection technology for pulmonary tuberculosis lesions on chest X-rays can quickly understand the specific conditions of pulmonary tuberculosis lesions and provide more auxiliary information for the next step of personalized response plans and precise prevention and control. However, based on the imaging principle of X-rays, multiple tissues or organs in the body will appear superimposed on chest X-rays, and the presence of ribs can mask or mimic lung nodules, masses, or infiltrates, especially in the axillary area, thereby affecting the analysis of tuberculosis.

[0004] Traditional computer-aided segmentation methods for pulmonary tuberculosis lesions primarily utilize sliding windows of varying sizes to select specific regions within X-ray images as candidate frames. These frames are then extracted and analyzed to determine if they represent the desired lesion. A drawback of these traditional methods is their poor adaptability to image changes, making it difficult to accurately segment and detect complex lesions. With the rapid development of deep learning technology in recent years, computer-assisted analysis of tuberculosis based on medical imaging has become a hot topic. Tuberculosis at different stages often manifests differently on images, making efficient and accurate detection and segmentation of pulmonary tuberculosis lesions crucial for pulmonary tuberculosis analysis. However, X-ray images generated by different devices can vary significantly, and existing models employ simple, static, random sample data augmentation methods, making it difficult to ensure the generalization of the trained models. Furthermore, segmentation performance is significantly compromised when data from different devices is used. The self-attention mechanism has been widely used in various fields. Its benefit is that it allows the model to better focus on certain specific areas. However, existing attention mechanisms model global relationships in images, integrating features from all locations without taking into account the unique structure of the human body. This approach is not only time-consuming and labor-intensive, but may not be the optimal solution for segmenting tuberculosis lesions. Summary of the Invention

[0005] To address the shortcomings of the existing technology, while taking into account efficiency, fully consider the tuberculosis lesions that are ignored due to imaging factors, improve the accuracy and recall rate of pulmonary tuberculosis lesion segmentation, and take into account the applicability of the model, the present invention adopts the following technical solutions:

[0006] A pulmonary tuberculosis lesion image segmentation system based on symmetric similarity and dynamic sample amplification, including a sternum suppression network module, an entropy-driven sample quantity dynamic amplification module, and a pulmonary tuberculosis lesion detection module;

[0007] The sternum suppression network module is used to remove the sternum image in the chest X-ray image, reduce the misjudgment of tuberculosis caused by the presence of ribs, and improve the performance index of the model lesion detection;

[0008] The entropy-driven dynamic sample expansion module dynamically expands the number of sample images for a single image based on the image's complexity and the model's current training state. Specifically, for each image, the information entropy of the model output probability distribution for each training round is calculated using the softmax function. In the early stages of training, the number of training image samples is slightly expanded. As the model trains, the number of expanded samples is increased to improve the model's ability to learn from different samples.

[0009] The tuberculosis lesion detection module uses an anchor-based target detector and utilizes the symmetric similarity of tuberculosis images to design a tuberculosis lesion detection network model to segment the tuberculosis lesion image from the X-ray image.

[0010] Furthermore, the sternum suppression network module is constructed based on the U-NetSharp architecture to build a three-level hierarchical structure including multi-layer encoders and decoders, introduces a cross-layer multi-scale jump connection mechanism, and adopts a feature fusion strategy of collaborative basic convolution blocks and void convolution blocks while retaining the original U-NetSharp feature extraction capability. The first level consists of four void convolution blocks and five basic blocks, the second level consists of four void convolution blocks and basic blocks, and the third level contains only one void convolution block, where the basic convolution block contains a 3×3 standard convolution kernel (step 1, zero padding) and an xUnit activation module (ReLU-depth separable convolution-Sigmoid complex The proposed network is composed of a dilated convolutional neural network (DCNN) with a progressive dilation rate design. The number of dilations increases gradually from the shallow layer to the bottleneck layer (2→16). By cascading multiple basic blocks with different dilation rates and supplemented by inter-block skip connections, the exponential expansion of the receptive field can be achieved while maintaining a constant number of parameters. The network architecture uses the dilated convolutional block as the core feature extractor, and fuses the extracted multi-scale features with the basic convolutional features through cross-level skip connections. At the same time, to address the semantic gap problem of multi-scale features, an intermediate fusion module composed of basic convolutional blocks is introduced after feature splicing to generate an intermediate tensor representing multi-level contextual associations, ultimately resulting in the suppression of sternal structure and enhancement of soft tissue in chest X-ray images.

[0011] Furthermore, the formula used by the sternum suppression network module is as follows:

[0012] F=D 1,5 (B 1,4 (B 1,3 (B 1,2 (D 1,1 (IM)))))+D 2,4 (B 2,3 (B 2,2 (D 2,1 (D 1,1 ))))+D 3,3 (B 3,2 (D 3,1 (D 2,1 )))+

[0013] D 4,2 (D 4,1 (D 3,1 ))+D 5,1 (D 4,1 )

[0014] B i,j (x)=InstanceNorm(xUnit(3x3Conv(x)))xUnit(x)=x·σ(DWConv(ReLU(x)))

[0015] Di,j (x)=InstanceNorm(xUnit(DilatedConv(x)))

[0016]

[0017] Among them, IM represents the input chest X-ray image, B i,j represents the basic convolutional block, D i,j Represents a dilated convolution block, x represents the image features extracted from the previous branch, InstanceNorm represents the normalization operation, σ represents the Sigmoid function, DilatedConv represents a dilated convolution, k represents the radius of the convolution kernel, d represents the dilation rate (Dilation Rate), which controls the interval between convolution kernel elements, Input(x+d·i,y+d·j) represents the value of the input feature map at position (x+d·i,y+d·j), Kernel(i,j) represents the weight of the convolution kernel at coordinate (i,j), and F represents the final output image after sternum suppression.

[0018] Furthermore, the sternum suppression network module performs network optimization by minimizing the structural similarity loss function SSIM and the mean absolute error loss function MAE, decomposing SSIM into the product of brightness l, contrast c, and structure s, and the formula is expressed as:

[0019]

[0020] Among them, X and Y represent the predicted image and the true label respectively, μ x 、μ y Represents the pixel mean of the local window in the predicted image X and the true label Y, σ x , σ x Denotes the standard deviation of the local window in the predicted image X and the true label Y, σ xy Represents the covariance of the local window in the predicted image X and the true label Y, C1=(K1·L) 2 , C2=(K2·L) 2 , represents the constant term, K1=0.01 and K2=0.03 are constants, L is the pixel dynamic range (255 for 8-bit grayscale images), n and m represent the height and width of the image respectively, and x ij 、y ij Represent the predicted image pixel value and the real image pixel value respectively;

[0021] The overall loss function is:

[0022] Loss(X,Y)=α*SSIM(X,Y)+(1-α)·MAE(X,Y)

[0023] Here, α represents the hyperparameter weight and is set to 0.84.

[0024] Furthermore, the entropy-driven dynamic sample expansion module is based on an adaptive image sample expansion framework based on model training feedback. During the training phase, the cognitive state of each sample is evaluated through a quantitative model, which serves as the basis for dynamically expanding the sample number. Specifically, an initial enhancement strength (mag, initially 0) is first set to expand the number of images according to the current training state of the model during the training process. The softmax function of the output of each image in the previous training round of the model is calculated, and the information entropy of the obtained probability distribution is calculated to represent the model's confidence level in the current sample. When the sample output result is poor, the prediction result is uncertain and the model output has a high entropy value. When the sample output result is high, the model confidence is high and the entropy value is low. The sample expansion degree (mag) of the current round is calculated based on the current entropy value, scaled to the interval [0, 1], and inversely proportional to the entropy value. When mag→1, the samples after the expansion have higher variability, allowing the model to learn more diverse samples. When mag→0, only a small sample expansion is applied to ensure that the model can correctly learn effective tuberculosis features. By dynamically adjusting mag during the training process, adaptive sample size expansion is achieved, the generalization of the model to different samples is improved, and model overfitting is prevented.

[0025] Furthermore, in order to maintain the symmetry of the tuberculosis image, the entropy-driven dynamic sample expansion module does not use rotation, cropping, displacement, or other expansion operations that destroy the image symmetry when expanding the sample quantity. The present invention divides the intensity of the sample quantity expansion operation into 31 levels (0-30). The greater the intensity, the greater the image expansion amount and the greater the difference from the original image. The expansion methods used include no operation, contrast adjustment Contrast, histogram equalization Equalize, exposure inversion Solarize, color depth reduction Posterize, brightness adjustment Brightness, and sharpness adjustment Sharpness. To avoid premature attenuation of the expansion degree under the guidance of traditional cross-entropy loss, the entropy-driven dynamic sample quantity expansion module introduces an entropy regularization term to compress the prediction distribution. The loss function constraint forces the model to continuously perceive the expansion disturbance during training, thereby ensuring that it does not fall into a local optimal solution. Specifically, if the model shows high uncertainty for some amplified samples, it may indicate that these expansion methods are too aggressive, making it difficult for the model to correctly segment. In this case, the number of times these expansion methods are used is reduced during training to ensure the stability and performance of the model. The specific formula used is as follows:

[0026] P c=softmax(Output net )

[0027]

[0028] a contrast ,a brightness ,a sharpness =0.1+mag(x)·1.8

[0029]

[0030] Among them, Output net Represents the output of the model, C represents the number of categories, P c Represents the probability distribution of the model output after passing through the softmax function, ε=10 -8 is a small value used for numerical stability, Ent represents the rounding function, mag represents the function for finding the vector modulus, and a Contrast 、a Equalize 、a Solarize 、a Posterize 、a Brightness 、a Sharpness They correspond to different image sample amplification methods, level indicates the intensity level of the current amplification operation, IM eq represents the image after histogram equalization, and IM represents the original image;

[0031] The overall loss function is:

[0032]

[0033] Among them, L CE represents the category cross entropy loss, L Dice Represents the segmentation loss of the lesion area, Loss Ent represents the dynamic entropy-augmented regularization loss.

[0034] Furthermore, the tuberculosis lesion detection module is improved from the YOLO v9 target detector. In order to keep the network lightweight and improve the computational efficiency of the network, the present invention reconstructs the corresponding module (RepNCSPELAN4) of the YOLOv9 target detector based on the depthwise separable convolution technology of MobileNet v4, adopts an inverted residual structure to replace the standard convolution layer, and constructs a cheap feature reuse path by introducing a double-layer depthwise convolution (DWC); at the target regression optimization level, the CIouLoss loss function is used to constrain the accuracy of the network prediction box regression; and due to the high degree of target overlap, in order to improve the accuracy and efficiency of non-maximum suppression (NMS) in the inference stage, the present invention adopts the non-maximum suppression algorithm Matrix NMS, which is improved from Soft NMS and improves the efficiency of algorithm execution while retaining more potential targets.

[0035] Furthermore, in the tuberculosis lesion detection module, in terms of feature fusion architecture, the Gold Yolo feature aggregation-distribution mechanism is introduced. First, a four-level feature alignment module (FAM) is constructed, and the third and fifth layer feature maps output by the backbone network are down-sampled to the same resolution, that is, the seventh layer resolution, through bilinear interpolation, and the ninth layer feature maps are up-sampled to the same resolution; then, a cascade splicing and channel compression strategy is adopted for cross-scale feature fusion, and a parameterized convolution group including 3×3 depth convolution, 1×1 point-by-point convolution and SE attention gating is used to generate global context features; finally, a multi-branch distribution structure is designed, and the enhanced global features are injected into the feature maps of each level through channel splitting and spatial point attention weighting, which can effectively overcome the information attenuation problem of the traditional FPN feature pyramid; the formula used is as follows:

[0036] F align =FAM([B3+B5+B7+B9])

[0037] F fuse =RepBlock(F align ) F inj_Pi , F inj_Pj =Split(F fuse )

[0038] F action_Pi = resize(σ(Conv 1x1 (F inj_Pi )) F global_Pi =resize(Conv 1x1 (F inj_Pi ))

[0039] F out =RepBlock(Conv 1x1 (B i)*F action_Pi +F global_Pi )

[0040] Among them, FAM represents the feature alignment module, B i represents the feature map of layer i, F align Represents a concatenated feature map of uniform size, RepBlock represents a multi-layer parameterized convolution module, and F fuse Indicates that the information fusion module fuses features, Split indicates that the channel splitting operation splits the fused features, and F inj_Pi 、F inj_Pj Indicates that the information injection module transmits features of different levels, Conv 1×1 represents the convolution operation, σ represents the Sigmoid activation function, resize represents the double upsampling using bilinear interpolation, F action_Pi represents the weighted coefficient calculated by simple attention, F global_Pi Represents an injected feature that contains global information.

[0041] Furthermore, in the tuberculosis lesion detection module, the chest X-ray image only images the chest area from a single perspective, and the differences between the images are usually limited to the abnormal areas that are difficult to detect. Because the human body is roughly symmetrical along the central axis, and the lung lesion area of pulmonary tuberculosis patients is significantly different from the normal human tissue structure, when a lesion appears in one lung area, the feature similarity between the area and the horizontally symmetrical area can be calculated to determine whether the area has a lesion, so that the model will focus on the area with large left and right differences, namely the tuberculosis lesion area; specifically, the multi-scale feature map obtained by the aggregation-distribution mechanism is horizontally flipped to obtain a symmetrical reference feature; the symmetry difference is calculated by the difference metric, and then the difference map is passed through the Sigm The oid activation function is normalized to the attention weight, and the attention weight is applied to the original feature to enhance the features of the asymmetric area. In order to deal with the deviation of the image structure caused by imaging or the offset of the lesion symmetry center caused by the non-rigid deformation of the lesion, an additional convolution layer is used to predict the offset of each sampling point with reference to the variable convolution method, and dynamic symmetry features are generated through bilinear interpolation. A learnable offset is introduced to allow the model to dynamically adjust the position of the symmetry reference point within a certain range. The GroundTruth true value box is mapped to the feature map scale to generate a binary mask for introducing symmetry contrast loss, which constrains the model to focus on the asymmetry caused by the lesion, encourages the model to generate greater symmetry differences in the lesion area, and maintains symmetry in the normal area.

[0042] Furthermore, the symmetric contrast loss introduced The formula is as follows:

[0043] Δx, Δy=Conv(F l )

[0044]

[0045] Among them, F l Represents the feature map of the first layer, Δx and Δy represent the symmetric offsets, and the symmetric contrast loss is used to encourage the model to find the tuberculosis lesion area with the largest difference; BaseGrid represents the coordinate matrix after horizontal flipping, and GridSample represents the bilinear sampling operation based on the coordinate matrix to obtain the dynamically adjusted symmetric feature map of the first layer. C l represents the number of channels of the feature map of layer l, represents the feature map of the cth channel in the lth layer, Represents the element-by-element multiplication operation, D l represents the symmetric difference of the feature map of the lth layer, σ represents the Sigmoid activation function, Represents the feature map after symmetric similarity amplification; Represents the symmetric difference value of the feature map position (i, j) of the lth layer, Represents the lesion area mask of the feature map position (i, j) of the lth layer (the lesion area is 1 and the normal area is 0), N p Represents the total number of pixels in the lesion area in the lth layer mask, N n Represents the total number of pixels in the normal area of the l-th layer mask.

[0046] The overall loss function is:

[0047]

[0048] Among them, L CE represents the category cross entropy loss, L Dice Represents the segmentation loss of the lesion area, Loss Ent represents the dynamic entropy-augmented regularization loss.

[0049] The advantages and beneficial effects of the present invention are:

[0050] This invention is based on a supervised deep learning approach, involving image sample amplification, target detection, and a self-attention mechanism, combining the advantages of both method-driven and data-driven approaches. Compared to traditional algorithms, the self-attention mechanism based on symmetric similarity establishes a bidirectional correlation mapping in the feature space of the anatomical structure of the lung lobe, which can greatly improve the model's feature extraction efficiency and learning ability. Removing rib features from the image can reduce the model's misjudgment rate due to rib obscuration. Dynamic sample size amplification can help establish a nonlinear mapping model between image complexity and amplification degree, thereby adaptively enhancing the model's generalization and expressiveness, thereby increasing the model's scope of application. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 It is a block diagram of a system in an embodiment of the present invention. DETAILED DESCRIPTION

[0052] The following describes the specific embodiments of the present invention in detail with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present invention and are not intended to limit the present invention.

[0053] The training and test data of the network are as follows: In the embodiment of the present invention, the TBX11K standardized medical imaging dataset is used for model training and verification. The dataset contains 800 chest X-ray images of 512×512 pixels, and the accompanying lesion annotation JSON file covers two pathological types of pulmonary tuberculosis: 630 cases of active tuberculosis, 140 cases of chronic tuberculosis and 30 mixed lesions (double-label independent annotation, that is, including active tuberculosis and chronic tuberculosis). A training set (480 cases), a validation set (160 cases) and a test set (160 cases) are constructed in a 3:1:1 ratio through a stratified sampling strategy. Dynamic mask loading technology is implemented for mixed cases during the training process to ensure that multiple lesion areas independently participate in the feature learning of the corresponding categories.

[0054] like Figure 1 As shown in the figure, the pulmonary tuberculosis lesion image segmentation system based on symmetric similarity and dynamic sample amplification includes a sternum suppression network module, an entropy-driven sample quantity dynamic amplification module, and a pulmonary tuberculosis lesion detection module. The specific process is as follows:

[0055] The sternum suppression network module (xU-NetFullSharp model) is an improvement on the U-NetSharp architecture. By constructing a three-level hierarchical structure consisting of a 5-layer encoder and a 5-layer decoder, a cross-layer multi-scale jump connection mechanism is introduced. While retaining the original U-NetSharp feature extraction capability, a feature fusion strategy of collaborative basic convolution blocks and dilated convolution blocks is adopted: the first level consists of four dilated convolution blocks and five basic blocks, the second level consists of four dilated convolution blocks and a basic block, and the third level contains only one dilated convolution block. The basic convolution block consists of a 3×3 standard convolution kernel (stride 1, zero padding) and an xUnit activation module (ReLU-depthwise separable convolution-Sigmoid composite structure), while the dilated convolution block adopts a progressive dilation rate design (the number of dilations from the shallow layer to the bottleneck layer is 2→16). By cascading multiple basic blocks with different dilation rates and supplementing them with inter-block jump connections, the receptive field can be exponentially expanded while maintaining a constant number of parameters. The network architecture uses the dilated convolution block as the core feature extractor, fusing the extracted multi-scale features with the basic convolution features through cross-level jump connections. At the same time, to address the semantic gap of multi-scale features, an intermediate fusion module composed of basic convolution blocks is introduced after feature splicing to generate an intermediate tensor representing multi-level contextual associations, ultimately achieving precise suppression of sternal structure and soft tissue enhancement in chest X-ray images. The network function is described as follows:

[0056] F=D 1,5 (B 1,4 (B 1,3 (B 1,2 (D 1,1 (IM)))))+D 2,4 (B 2,3 (B 2,2 (D 2,1 (D 1,1 ))))+D 3,3 (B 3,2 (D 3,1 (D 2,1 )))+

[0057] D 4,2 (D 4,1 (D 3,1 ))+D 5,1 (D 4,1 )

[0058] B i,j (x)=InstanceNorm(xUnit(3x3Conv(x)))xUnit(x)=x·σ(DWConv(ReLU(x)))

[0059] D i,j(x)=InstanceNorm(xUnit(DilatedConv(x)))

[0060]

[0061] Among them, IM represents the input chest X-ray image, B i,j represents the basic convolutional block, D i,j Represents a dilated convolution block, x represents the image features extracted from the previous branch, InstanceNorm represents the normalization operation, σ represents the Sigmoid function, DilatedConv represents a dilated convolution, k represents the radius of the convolution kernel, d represents the dilation rate (Dilation Rate), which controls the interval between convolution kernel elements, Input(x+d·i,y+d·j) represents the value of the input feature map at position (x+d·i,y+d·j), Kernel(i,j) represents the weight of the convolution kernel at coordinate (i,j), and F represents the final output image after sternum suppression.

[0062] The network optimization goal is to minimize the structural similarity loss function (SSIM) and the mean absolute error loss function (MAE). SSIM is decomposed into the product of brightness (l), contrast (c), and structure (s), and the formula is expressed as:

[0063]

[0064] Among them, X and Y represent the predicted image and the true label respectively, μ x 、μ y Represents the pixel mean of the local window in image X and Y, σ x , σ x Represents the standard deviation of the local window in image X and Y, σ xy Represents the covariance of the local window in images X and Y. The constant term C1=(K1·L) 2 , C2=(K2·L) 2 , K1=0.01, K2=0.03, L is the pixel dynamic range (255 for 8-bit grayscale images), n and m represent the height and width of the image respectively, x ij 、y ij Represent the predicted image pixel value and the real image pixel value respectively.

[0065] The overall loss function is:

[0066] Loss(X,Y)=α*SSIM(X,Y)+(1-α)·MAE(X,Y)

[0067] Among them, α is the hyperparameter weight, which is set to 0.84.

[0068] The entropy-driven dynamic sample augmentation module, an adaptive image sample augmentation framework based on model training feedback, focuses on evaluating the cognitive state of each sample during the training phase using a quantitative model, which serves as the basis for dynamic sample augmentation. Specifically, an initial augmentation strength (mag, initially 0) is set to augment the sample size based on the model's current training state during training. The softmax function is calculated for each image output from the previous training round, and the resulting probability distribution is then used to calculate the information entropy, which represents the model's confidence level in the current sample. Poor sample output indicates high uncertainty in the prediction, resulting in a high model output entropy. High sample output indicates high model confidence and a low entropy. The sample augmentation degree (mag) for the current round is calculated based on the current entropy value and scaled to the [0, 1] interval, inversely proportional to the entropy value. When mag increases to 1, the augmented samples exhibit greater variability, allowing the model to learn a wider variety of samples. When mag decreases to 0, only a small amount of sample augmentation is applied to ensure the model correctly learns valid nodule features. By dynamically adjusting the mag during training, adaptive sample augmentation is achieved, improving the model's generalization to diverse samples and preventing overfitting. To maintain the symmetry of tuberculosis images, augmentation operations that disrupt image symmetry, such as rotation, cropping, and translation, are avoided during sample augmentation. The intensity of the augmentation operation is categorized into 31 levels (0-30), with higher intensity indicating greater image augmentation and a greater difference from the original image. Seven augmentation methods are used: no operation, contrast adjustment, histogram equalization, solarize, posterize, brightness adjustment, and sharpness adjustment. To prevent premature attenuation of the augmentation under traditional cross-entropy loss guidance, the module introduces an entropy regularization term to compress the prediction distribution. This loss function constraint forces the model to continuously perceive augmentation perturbations during training, thus preventing it from falling into local optima. Specifically, if the model exhibits high uncertainty for certain augmented samples, this may indicate that the augmentation methods were overly aggressive, making it difficult for the model to correctly segment the image. The number of times these augmentation methods are used will be reduced during training to ensure the stability and performance of the model. The function description of this module is:

[0069] P c =softmax(Output net )

[0070]

[0071] a contrast ,a brightness ,a sharpness =0.1+mag(x)·1.8

[0072]

[0073] Output net represents the output of the model, C is the number of categories, ε=10 -8 For numerical stability, a x Corresponding to different image sample amplification methods, IM eq represents the image after histogram equalization, and IM represents the original image.

[0074] The overall loss function is:

[0075]

[0076] The tuberculosis lesion detection module is improved from the YOLO v9 target detector. In order to keep the network lightweight and improve the computational efficiency of the network, the present invention reconstructs the RepNCSPELAN4 module of YOLOv9 based on the depth-separable convolution technology of MobileNet v4: an inverted residual structure is used to replace the standard convolution layer, and a cheap feature reuse path is constructed by introducing a double-layer depth convolution (DWC). At the target regression optimization level, CIou Loss is used to constrain the network prediction frame regression accuracy. Due to the high degree of target overlap, in order to improve the accuracy and efficiency of non-maximum suppression (NMS) in the inference stage, the present invention adopts Matrix NMS, which is improved from Soft NMS. While retaining more potential targets, it improves the efficiency of algorithm execution. In terms of feature fusion architecture, the Gold Yolo feature aggregation-distribution mechanism is introduced. First, a four-level feature alignment module (FAM) is constructed to unify the third and fifth layer feature maps output by the backbone network to the same resolution (seventh layer resolution) through bilinear interpolation downsampling, and the ninth layer feature maps are upsampled to the same resolution; then, a cascade splicing and channel compression strategy is used for cross-scale feature fusion, and a parameterized convolution group (including 3×3 depth convolution, 1×1 point-by-point convolution and SE attention gating) is used to generate global context features; finally, a multi-branch distribution structure is designed to inject the enhanced global features into the feature maps of each level through channel splitting and spatial point attention weighting, which can effectively overcome the information attenuation problem of the traditional FPN feature pyramid. The function of this module is described as follows:

[0077] F align =FAM([B3+B5+B7+B9])

[0078] F fuse=RepBlock(F align ) F inj_P15 , F inj_P12 =Split(F fuse )

[0079] F action_Pi = resize(σ(Conv 1x1 (F inj_Pi )) F global_Pi =resize(Conv 1x1 (F inj_Pi ))

[0080] F out =RepBlock(Conv 1x1 (B i )*F action_Pi +F global_Pi )

[0081] Where FAM represents the feature alignment module, B i represents the feature map of layer i, F align is a concatenated feature map of uniform size, RepBlock represents a multi-layer parameterized convolution module, F fuse Indicates that the information fusion module fuses features, Split indicates that the channel splitting operation splits the fused features, and F inj_Pi Indicates that the information injection module transfers features of different levels. resize indicates that bilinear interpolation is used for double upsampling. action_Pi represents the weighted coefficient calculated by simple attention, F global_Pi Represents an injected feature that contains global information.

[0082] Chest X-ray images capture the chest region from a single perspective, and differences between images are typically limited to barely perceptible abnormal areas. Because the human body is roughly bilaterally symmetrical along its central axis, the lung lesions of tuberculosis patients differ significantly from normal tissue structure. When a lesion appears in one lung region, the presence of a lesion can be determined by calculating the feature similarity between that region and horizontally symmetric regions. This allows the model to focus on areas with significant left-right differences, namely the tuberculosis lesions. Specifically, the multi-scale feature map generated by the aggregation-distribution mechanism is horizontally flipped to generate symmetric reference features. Symmetry differences are then calculated using a dissimilarity metric. The dissimilarity map is then normalized using a sigmoid activation function to generate attention weights. The attention weights are then applied to the original features to enhance features in asymmetric regions. To address shifts in the symmetry center caused by imaging-induced image structural deviations or non-rigid lesion deformation, an additional convolutional layer is used, modeled after a variable convolution approach, to predict the offset for each sample point. Dynamic symmetry features are generated through bilinear interpolation. A learnable offset is introduced, allowing the model to dynamically adjust the position of the symmetry reference point within a certain range. The GroundTruth box is mapped to the feature map scale to generate a binary mask, which is used to introduce symmetric contrast loss. This constrains the model to focus on the asymmetry caused by the lesion, encourages the model to generate greater symmetric differences in the lesion area, and maintain symmetry in the normal area. The function of this module is described as follows:

[0083] Δx, Δy=Conv(F l )

[0084]

[0085] Among them F l Represents the feature map of the first layer, Δx and Δy represent the symmetric offsets, and the symmetric contrast loss encourages the model to find the tuberculosis lesion area with the largest difference. BaseGrid represents the coordinate matrix after horizontal flipping, and GridSample represents the bilinear sampling operation based on the coordinate matrix to obtain the dynamically adjusted symmetric feature map. C l represents the number of channels of the feature map of layer l, shows the feature map of the c-th channel, Represents an element-wise multiplication operation. Represents the symmetric difference value of the feature map position (i, j) of the lth layer, Represents the lesion area mask of the feature map position (i, j) of the lth layer (the lesion area is 1 and the normal area is 0), N p Represents the total number of pixels in the lesion area in the lth layer mask, N n Represents the total number of pixels in the normal area of the l-th layer mask.

[0086] The overall loss function is:

[0087]

[0088] We found that even when multiple target lesions are present in a single X-ray image, the network model can still effectively segment the target lesions, with no overlap in the segmented regions. The network model can easily learn our intent and distinguish between tuberculosis lesions that differ between the left and right lobes. Experiments show that even when data contains both old and active tuberculosis lesions, this method can still effectively segment the corresponding lesion regions. Furthermore, our method performs well for segmenting tuberculosis lesions of varying structures and sizes, and the results on the test set also demonstrate significant improvements in the model's generalization.

[0089] The system of the present invention covers a sternal suppression network module, an entropy-driven dynamic amplification module for the number of image samples, and a tuberculosis lesion detection module based on symmetric similarity, and is trained on the TBX11K dataset. Compared with traditional algorithms, the method of the present invention greatly improves the efficiency of data processing; due to the addition of dynamic entropy amplification, the generalization of the model is enhanced, making the model more applicable, and it can achieve good segmentation of tuberculosis with different morphological structures. Multi-scale feature fusion enables the model to pay good attention to small calcified tuberculosis and larger infected lesion areas; at the same time, the mechanism based on symmetric similarity increases the model's attention to the lesion area. For imaging data of both active tuberculosis lesions and old tuberculosis lesions, different categories of lesion areas can still be accurately segmented; for imaging data containing multiple lesions with different structures at the same time, efficient and accurate segmentation can also be achieved. The above-mentioned multiple methods greatly improve the accuracy of tuberculosis lesion segmentation.

[0090] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A pulmonary tuberculosis lesion image segmentation system based on symmetric similarity and dynamic sample amplification, including a sternum suppression network module, an entropy-driven sample quantity dynamic amplification module, and a pulmonary tuberculosis lesion detection module, characterized by: The sternum suppression network module is used to remove the sternum image in the chest X-ray image; The entropy-driven dynamic sample expansion module dynamically expands the number of sample images for a single image based on the image's complexity and the model's current training state. Specifically, for each image, the information entropy is calculated based on the probability distribution of the model output for each training round. In the early stages of training, the number of training image samples is slightly expanded, and the number of expanded samples is increased as the model trains. The tuberculosis lesion detection module uses a target detector and utilizes the symmetric similarity of tuberculosis images to design a tuberculosis lesion detection network model to segment the tuberculosis lesion image from the X-ray image.

2. The pulmonary tuberculosis lesion image segmentation system based on symmetric similarity and dynamic sample amplification according to claim 1, characterized in that: The sternum suppression network module includes a three-level hierarchical structure of multi-layer encoders and decoders, introduces a cross-layer multi-scale jump connection mechanism, and adopts a feature fusion strategy of collaborative basic convolution blocks and void convolution blocks. The first level is composed of four void convolution blocks and five basic blocks, the second level is composed of four void convolution blocks and basic blocks, and the third level contains only one void convolution block, wherein the basic convolution block contains a standard convolution kernel and an xUnit activation module. The void convolution block adopts a progressive void rate design, and the number of voids gradually increases from the shallow layer to the bottleneck layer. By cascading multiple basic blocks with differentiated void rates and supplemented by inter-block jump connections; the network architecture uses the void convolution block as the core feature extractor, and fuses the extracted multi-scale features with the basic convolution features through cross-layer jump connections. At the same time, after feature splicing, an intermediate fusion module composed of basic convolution blocks is introduced to generate an intermediate tensor representing multi-level contextual associations, ultimately suppressing the sternum structure and enhancing the soft tissue in the chest X-ray image.

3. The pulmonary tuberculosis lesion image segmentation system based on symmetric similarity and dynamic sample amplification according to claim 2, characterized in that: The formula used by the sternal suppression network module is as follows: F=D 1,5 (B 1,4 (B 1,3 (B 1,2 (D 1,1 (IM)))))+D 2,4 (B 2,3 (B 2,2 (D 2,1 (D 1,1 ))))+D 3,3 (B 3,2 (D 3,1 (D 2,1 )))+ D 4,2 (D 4,1 (D 3,1 ))+D 5,1 (D 4,1 ) B i,j (x)=InstanceNorm(xUnit(3x3Conv(x)))xUnit(x)=x·σ(DWConv(ReLU(x))) D i,j (x)=InstanceNorm(xUnit(DilatedConv(x))) Among them, IM represents the input chest X-ray image, B i,j represents the basic convolutional block, D i,j Represents a dilated convolution block, x represents the image features extracted from the previous branch, InstanceNorm represents the normalization operation, σ represents the Sigmoid function, DilatedConv represents a dilated convolution, k represents the radius of the convolution kernel, d represents the dilation rate, and controls the interval between convolution kernel elements. Input(x+d·i,y+d·j) represents the value of the input feature map at the position (x+d·i,y+d·j), Kernel(i,j) represents the weight of the convolution kernel at the coordinate (i,j), and F represents the final output image after sternum suppression.

4. The pulmonary tuberculosis lesion image segmentation system based on symmetric similarity and dynamic sample amplification according to claim 2, characterized in that: The sternum suppression network module is optimized by minimizing the structural similarity loss function SSIM and the mean absolute error loss function MAE. SSIM is decomposed into the product of brightness l, contrast c, and structure s. The formula is expressed as: Among them, X and Y represent the predicted image and the true label respectively, μ x 、μ y Represents the pixel mean of the local window in the predicted image X and the true label Y, σ x , σ x Denotes the standard deviation of the local window in the predicted image X and the true label Y, σ xy Represents the covariance of the local window in the predicted image X and the true label Y, C1=(K1·L) 2 , C2=(K2·L) 2 , represents the constant term, K1 and K2 are constants, L is the pixel dynamic range, n and m represent the height and width of the image respectively, x ij 、y ij Represent the predicted image pixel value and the real image pixel value respectively; The overall loss function is: Loss(X,Y)=α*SSIM(X,Y)+(1-α)·MAE(X,Y) Among them, α represents the hyperparameter weight.

5. The pulmonary tuberculosis lesion image segmentation system based on symmetric similarity and dynamic sample amplification according to claim 1, characterized in that: The entropy-driven dynamic sample expansion module is an adaptive image sample expansion framework based on model training feedback. During the training phase, the cognitive state of each sample is evaluated through a quantitative model, and this is used as the basis for dynamically expanding the number of samples. Specifically, the initial enhancement strength is first set to expand the number of images according to the current training state of the model during the training process. The softmax function of the output of each image in the previous training round of the model is calculated, and the information entropy of the obtained probability distribution is calculated to represent the model's confidence level in the current sample. When the sample output result is poor, the prediction result is uncertain and the entropy value of the model output is high. When the sample output result is high, it means that the model confidence is high and the entropy value is low; the degree of sample expansion in the current round is calculated based on the current entropy value, scaled to the interval [0,1], and inversely proportional to the entropy value; when mag→1, the sample after the expansion has higher variability; and when mag→0, only a small sample expansion is applied.

6. The pulmonary tuberculosis lesion image segmentation system based on symmetric similarity and dynamic sample amplification according to claim 5, characterized in that: The entropy-driven dynamic sample expansion module uses expansion methods including no operation, contrast adjustment Contrast, histogram equalization Equalize, exposure inversion Solarize, color depth reduction Posterize, brightness adjustment Brightness, and sharpness adjustment. The entropy-driven dynamic sample quantity expansion module introduces an entropy regularization term to compress the prediction distribution and forces the model to continuously perceive the expansion disturbance during training through the loss function constraint. Specifically, if the model shows high uncertainty for certain amplified samples, the number of times these expansion methods are used is reduced during training. The specific formula used is as follows: P c =softmax(Output net ) a contrast ,a brightness ,a sharpness =0.1+mag(x)·1.8 Among them, Output net Represents the output of the model, C represents the number of categories, P c It represents the probability distribution of the model output after passing the softmax function, ε is a small value used for numerical stability, Ent represents the rounding function, mag represents the function for finding the vector modulus, and a Contrast 、a Equalize 、a Solarize 、a Posterize 、a Brightness 、a Sharpness They correspond to different image sample amplification methods, level indicates the intensity level of the current amplification operation, IM eq represents the image after histogram equalization, and IM represents the original image; The overall loss function is: Among them, L CE represents the category cross entropy loss, L Dice Represents the segmentation loss of the lesion area, Loss Ent represents the dynamic entropy-augmented regularization loss.

7. The pulmonary tuberculosis lesion image segmentation system based on symmetric similarity and dynamic sample amplification according to claim 1, characterized in that: The tuberculosis lesion detection module reconstructs the corresponding module of the target detector based on the depthwise separable convolution technology, adopts the inverted residual structure to replace the standard convolution layer, and constructs a cheap feature reuse path by introducing a double-layer deep convolution; at the target regression optimization level, a loss function is used to constrain the accuracy of the network prediction box regression; and a non-maximum suppression algorithm is adopted.

8. The pulmonary tuberculosis lesion image segmentation system based on symmetric similarity and dynamic sample amplification according to claim 1, characterized in that: In the tuberculosis lesion detection module, a feature aggregation-distribution mechanism is introduced in the feature fusion architecture. First, a four-level feature alignment module is constructed. The third and fifth layer feature maps output by the backbone network are downsampled to the same resolution, that is, the seventh layer resolution, through bilinear interpolation, and the ninth layer feature maps are upsampled to the same resolution. Then, a cascade splicing and channel compression strategy is used to perform cross-scale feature fusion, and a global context feature is generated through a parameterized convolution group. Finally, a multi-branch distribution structure is designed to inject the enhanced global features into the feature maps of each level through channel splitting and spatial point attention weighting; the formula used is as follows: <h2 style=";text-align:left;direction:ltr">F<h2 style=";text-align:left;direction:ltr"> align <h2 style=";text-align:left;direction:ltr"> =FAM([B3+B5+B7+B9]) F fuse =RepBlock(F align ) F inj_Pi ,F inj_Pj =Split(F fuse ) F action_Pi =resize(σ(Conv 1x1 (F inj_Pi ))F global_Pi =resize(Conv 1x1 (F inj_Pi )) F out =RepBlock(Conv 1x1 (B i )*F action_Pi +F global_Pi ) Among them, FAM represents the feature alignment module, B i represents the feature map of layer i, F align Represents a concatenated feature map of uniform size, RepBlock represents a multi-layer parameterized convolution module, and F fuse Indicates that the information fusion module fuses features, Split indicates that the channel splitting operation splits the fused features, and F inj_Pi 、F inj_Pj Indicates that the information injection module transmits features of different levels, Conv 1×1 represents the convolution operation, σ represents the Sigmoid activation function, resize represents upsampling using bilinear interpolation, F action_Pi represents the weighted coefficient calculated by simple attention, F global_Pi Represents an injected feature that contains global information.

9. The pulmonary tuberculosis lesion image segmentation system based on symmetric similarity and dynamic sample amplification according to claim 1, characterized in that: In the tuberculosis lesion detection module, the multi-scale feature map obtained by the aggregation-distribution mechanism is horizontally flipped to obtain a symmetrical reference feature; The symmetry difference is calculated through the difference metric, and then the difference map is normalized into attention weights through the Sigmoid activation function. The attention weights are applied to the original features to enhance the features of asymmetric areas. An additional convolutional layer is used to predict the offset of each sampling point, and dynamic symmetry features are generated through bilinear interpolation. A learnable offset is introduced to allow the model to dynamically adjust the position of the symmetry reference point within a certain range. The true value box is mapped to the feature map scale to generate a binary mask for introducing symmetry contrast loss, which constrains the model to focus on the asymmetry caused by the lesion, encourages the model to generate larger symmetry differences in the lesion area, and maintains symmetry in the normal area.

10. The pulmonary tuberculosis lesion image segmentation system based on symmetric similarity and dynamic sample amplification according to claim 9, characterized in that: Introduced symmetric contrast loss The formula is as follows: Δx,Δy=Conv(F l ) Among them, F l represents the feature map of the first layer, Δx and Δy represent the symmetric offsets, and the symmetric contrast loss is used to encourage the model to find the tuberculosis lesion area with the largest difference; BaseGrid represents the coordinate matrix after horizontal flipping, and GridSample represents the bilinear sampling operation based on the coordinate matrix to obtain the dynamically adjusted symmetric feature map F of the first layer. l sym ; C l Indicates the number of channels of the l-th layer feature map, F l (c) represents the feature map of the cth channel in the lth layer, Represents the element-by-element multiplication operation, D l represents the symmetric difference of the feature map of the lth layer, σ represents the Sigmoid activation function, Represents the feature map after symmetric similarity amplification; Represents the symmetric difference value of the feature map position (i, j) of the lth layer, Represents the lesion area mask at the position (i, j) of the feature map of the lth layer, N p Represents the total number of pixels in the lesion area in the lth layer mask, N n Represents the total number of pixels in the normal area of the l-th layer mask. The overall loss function is: Among them, L CE represents the category cross entropy loss, L Dice Represents the segmentation loss of the lesion area, Loss Ent represents the dynamic entropy-augmented regularization loss.

Citation Information

Cited By

  • Medical image X-ray pulmonary tuberculosis medical information management system based on AI

    CN121354936A