Label noise-oriented pneumoconiosis chest X-ray image classification system
By introducing a dual dynamic threshold sample selection mechanism with progressive feature bias and multi-branch auxiliary modules, the problem of blurred decision boundaries caused by label noise in the classification of chest X-ray images of pneumoconiosis was solved, and accurate staging diagnosis of pneumoconiosis was achieved.
Patent Information
- Application Number
- CN202510898095.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-12-09
AI Technical Summary
Existing technologies for classifying chest X-ray images of pneumoconiosis suffer from label noise, which blurs the decision boundaries between adjacent stages, making it difficult to accurately distinguish between different stages and affecting the reliability and generalization ability of deep learning models.
A pneumoconiosis chest X-ray image classification system based on dual-threshold sample selection is adopted. By introducing the concept of progressive feature tendency and multi-branch auxiliary module, combined with dynamic threshold sample selection mechanism, the robustness of the model to label noise is improved.
It effectively simulates the progressive judgment mode of clinicians, accurately distinguishes between clean and difficult samples, reduces the interference of noise labels on model training, and improves the accuracy and robustness of pneumoconiosis diagnosis.
Smart Images

Figure CN121095613A_ABST
Abstract
Description
TECHNICAL FIELD
[0002] The application belongs to the field of medical image processing, and particularly relates to a pneumoconiosis chest X-ray image classification system based on double-threshold sample selection, which is used for solving the problem of adjacent stage decision boundary ambiguity caused by label noise in pneumoconiosis diagnosis, and realizing accurate staging diagnosis of pneumoconiosis. BACKGROUND
[0004] Pneumoconiosis is a serious occupational lung disease, and clinical diagnosis mainly relies on chest X-ray (CXR) images. However, due to the small size and low contrast of pneumoconiosis lesions, and the lack of objective quantitative indicators in the diagnostic criteria, the subjective judgment difference of different doctors leads to label noise in the data set, which seriously affects the reliability and generalization ability of the deep learning model. When dealing with pneumoconiosis CXR data, the existing method faces the problems of subtle differences in imaging features between adjacent stages, ambiguous decision boundaries, and label noise easily leading to model overfitting, and it is difficult to accurately distinguish different stages of pneumoconiosis. Therefore, there is an urgent need for a pneumoconiosis classification method that can effectively handle label noise and accurately define the decision boundary between adjacent stages. SUMMARY
[0006] To solve the above problems, the application provides a pneumoconiosis chest X-ray image classification system based on double-threshold sample selection, which introduces the concept of gradual feature tendency, designs a multi-branch auxiliary module and a double-dynamic threshold sample selection mechanism, and improves the robustness of the model to label noise and diagnostic accuracy.
[0007] The application adopts the following technical solutions:
[0008] A pneumoconiosis chest X-ray image classification system for label noise, which considers the class condition and instance dependence of PPLN by introducing gradual feature tendency analysis and calculating potential tendency difference based on the tendency, and combines probability-based dynamic threshold to construct a deep neural network model for sample division and classification, including three modules:
[0009] Main branch module: similar to a general deep neural network for classification tasks, input the pneumoconiosis chest X-ray image, and used for calculating and outputting the probability-based confidence score;
[0010] Multi-branch auxiliary module: contains auxiliary branches equal to the number of pneumoconiosis stages, and works in parallel with the main branch, through two stages of potential adjacent label mining and gradual feature tendency evaluation, the former stage inputs adjacent class images to each branch, and the latter stage inputs corresponding class images, respectively learns and predicts the gradual feature tendency, and outputs the confidence score reflecting the sample's tendency in the "mild" or "severe" direction;
[0011] Dual dynamic threshold sample selection module: calculate the potential tendency difference according to the confidence score output by the multi-branch auxiliary module, combine the probability-based dynamic threshold, form a dual dynamic threshold sample selection mechanism, and divide the samples into clean set and difficult set.
[0012] The system, the main structure of the multi-branch auxiliary module contains a plurality of auxiliary branches equal in number and category, which work in parallel with the main branch; the auxiliary branch diverges from the tail of the main network (for example, ResNet), and the head keeps sharing parameters; the module contains two stages, each of which starts with feature splitting, selects matching samples from a small batch of sample sets , respectively used for potential neighboring label mining (PNLM) and progressive feature tendency determination (PFTD);
[0013] Potential neighboring label: for a given label , its potential neighboring label (PNL) can be expressed as:
[0014]
[0015] The potential neighboring label set is defined as: ;
[0016] The generalized potential neighboring label is defined as:
[0017]
[0018] These definitions describe the situation that pneumoconiosis neighboring period labels are easily confused with each other and the form in the data set, and the progressive feature tendency can be explained as the degree of confusion of the label with any one of the GPNs in , and quantified by a confidence score.
[0019] The system, the feature splitting, the multi-branch auxiliary module is divided into two stages, which learn and predict the progressive feature tendency through different processes, PNLM accepts samples of neighboring categories, while PFTD accepts samples of corresponding categories, formally, for a given sample , all possible feature representations matching any stage in different branches are expressed as
[0020]
[0021] Among them, and respectively represent the features matched in the two stages of potential neighboring label mining and progressive feature tendency evaluation; represent the features in the main branch; the probability vectors corresponding to these features are denoted as:
[0022]
[0023] wherein, and are the probabilities in the PNLM and PFTD stages respectively, is the probability in the main branch.
[0024] In the PNLM stage of the multi-branch auxiliary module, a binary classifier is trained using the auxiliary branch, the sample is input into the branch corresponding to the generalized potential neighboring label set of the given label, the progressive feature tendency is learned, the branch-consistent focal loss (BCFL) is introduced to ensure the consistency of the learning of each branch, and is defined as:
[0025]
[0026] wherein, ; in the progressive feature tendency evaluation stage, the sample is input into the branch corresponding to the category of the given label, a prediction vector containing two components is output, and a confidence score is obtained through Sharpen regularization.
[0027] In the double dynamic threshold sample selection module, the absolute difference value of the feature tendencies predicted in two directions of the generalized potential neighboring label of the sample is defined as the latent tendency difference (LT-Diff), and for the sample the LT-Diff is defined as:
[0028]
[0029] wherein, only contains two elements and , which respectively represent the two generalized potential neighboring labels of the sample, two corresponding components and in respectively represent and directly reflect the feature tendencies predicted in two directions.
[0030] The system, based on the LT-Diff, calculates a dynamic threshold value for each sample in each iteration round of the deep neural network training, and also calculates a dynamic threshold value for each sample based on the probability-based confidence score output by the main branch module. Their iterative processes are represented as:
[0031]
[0032] wherein, is the iteration round index, and are two back-off ratios that control the degree of delay and the stability of the threshold value, and satisfy .
[0033] The system, in the iterative process, the back-off ratio of the dynamic threshold value based on LT-Diff The calculation method is
[0034]
[0035] This is called the "Double Layback-Ratio Trick" (DLRT). Considering incorporating it into the threshold iteration, a more stringent sample selection standard can be set. The double layback-ratio trick accelerates the convergence of the threshold value.
[0036] The system, the double dynamic threshold sample selection module divides the samples into a clean set and a difficult set The division criterion is given by:
[0037] .
[0038] The system, after dividing the samples in the training data set according to the sample division criterion, calculates the loss according to different loss functions to update the parameters of the deep neural network. Both loss functions involve an LT-Diff penalty term. The loss of the samples in the clean set is calculated using the loss function, and the loss of the samples in the difficult set is calculated using the loss function. The two loss functions are represented as:
[0039]
[0040] wherein, , and represent the size of the set , the set and the entire data set, serves as a scalar for the overall learning rate, and two LT-Diff based penalty terms and coefficients, similar to the original terms in cross-entropy loss and generalized cross-entropy loss.
[0041] The system, the constructed deep neural network model, and the overall loss function are expressed as:
[0042]
[0043] wherein (1) is a standard cross-entropy loss; (2) is a coefficient for balancing the two loss functions of the "clean" set and the "difficult" set; (3) and are weighted ramp functions with respect to the round index , which are expressed as follows:
[0044]
[0045] wherein and are two positive integers, and , respectively decrease and increase the overall loss function in the form of a Gaussian curve, and the preheating phase is introduced in the background of the ramp function, and is the index threshold of the preheating round.
[0046] Advantages
[0047] The application effectively simulates the progressive judgment mode of clinicians on adjacent period lesions by learning progressive feature tendency through a multi-branch auxiliary module, and excavates potential class condition factors. The double dynamic threshold sample selection mechanism combines a dynamic threshold based on probability and potential tendency difference to accurately distinguish clean samples and difficult samples, and reduce the interference of noise labels on model training. Experimental results show that the application has good robustness and diagnostic accuracy under different noise rates, and is significantly better than existing methods, providing an efficient and reliable solution for precise diagnosis of pneumoconiosis. BRIEF DESCRIPTION OF DRAWINGS
[0049] Figure 1 is a schematic diagram of progressive feature tendency in the progression of pneumoconiosis.
[0050] Figure 2 is a basic architecture diagram of the method provided by the application.
[0051] Figure 3 is a schematic diagram of progressive feature tendency (based on the first period).
[0052] Figure 4Illustration of the situation that auxiliary branches may encounter.
[0053] Figure 5 Illustration of the edge feature tendency (footing: dust-free lung).
[0054] Figure 6 Illustration of the accidental sample imbalance in some small batches.
[0055] Figure 7 Illustration of the progressive feature tendency evaluation of auxiliary branches.
[0056] Figure 8 Histogram of the threshold distribution based on LT-Diff changes with rounds.
[0057] Figure 9 Confusion matrix diagram of FTLTD-Net under different PPLN noise rates. DETAILED DESCRIPTION
[0059] The present application will be described in detail below in conjunction with specific embodiments.
[0060] A label noise-oriented classification system for chest X-ray images of pneumoconiosis includes three modules:
[0061] Main branch module: similar to a general deep neural network for classification tasks, the input is a chest X-ray image of pneumoconiosis, which is used to calculate and output a probability-based confidence score;
[0062] Multi-branch auxiliary module: this module contains auxiliary branches equal to the number of pneumoconiosis stages, and works in parallel with the main branch. The head of the backbone network serves as a feature extractor with shared parameters, and the tail diverges into multiple branches according to the number of stages. The module is divided into two stages. 1) Potential adjacent label mining stage: use auxiliary branches to train high-precision binary classifiers, input samples into two branches corresponding to the given label's generalized potential adjacent label set, and mine progressive feature tendency. For example, for a sample with a label of stage one, input into the branches corresponding to dust-free lung and stage two, learn the sample's feature tendency in "mild" (dust-free lung direction) and "severe" (stage two direction). Introduce branch consistent focal loss (BCFL) to balance the small batch sample imbalance problem, ensure the consistency of each branch learning, and make each auxiliary branch learn balanced feature tendency. 2) Progressive feature tendency evaluation stage: input the sample into the branch corresponding to its given label category, and each branch outputs a prediction vector containing two components, representing the sample's tendency in the "mild" and "severe" directions. Sharpen the regularization processing of the prediction confidence score to obtain the final confidence score, which is used for subsequent calculation of the potential tendency difference.
[0063] Double dynamic threshold sample selection module: according to the confidence score output by the multi-branch auxiliary module, the potential tendency difference (LT-Diff) of the sample is calculated, that is, the absolute difference of the predicted feature tendency of the sample in two directions. The dynamic threshold based on probability and the dynamic threshold based on LT-Diff are obtained through round-by-round iteration, and the double back-off ratio technique (DLRT) is used in the iteration process to speed up the convergence speed of the threshold based on LT-Diff. According to the double dynamic threshold, the samples are divided into clean set and difficult set: if the probability of the sample is greater than the threshold based on probability and the LT-Diff is less than the threshold based on LT-Diff, it is divided into clean set; if only one of the conditions is met, it is divided into difficult set. Different loss functions are used to process samples in different sets, and standard cross-entropy loss is used for clean set samples, and generalized cross-entropy loss and penalty term are used for difficult set samples to enhance the robustness of the model to noisy samples.
[0064] Based on the above label noise-oriented pneumoconiosis chest X-ray image classification system, a label noise-oriented pneumoconiosis chest X-ray image classification method is provided, which includes the following steps:
[0065] A1. Data preprocessing
[0066] The collected pneumoconiosis chest X-ray image (CXR) is center cropped to remove irrelevant information at the edges of the image and adjusted to a regular square of 512x512 pixels for easy model input and processing. Gray scale normalization is performed to normalize the pixel value to a 32-bit floating point number, with a mean of 0.5281 and a standard deviation of 0.2294, to improve the model's adaptability to different images. Random rotation (range [-15°, 15°]) and horizontal flipping are used for data augmentation to increase the diversity of training data and avoid model overfitting. When synthesizing label noise, based on the partial dependence on label noise simulation instance dependence factor, noise labels are generated for each instance according to truncated Gaussian distribution, and then asymmetric label noise mode is applied to simulate class condition factors, with a given label being confused with a generalized potential adjacent label with a certain probability, to obtain synthetic noise labels. The data set is divided into training set and test set, and the test set samples are strictly screened by more than three experienced radiologists to ensure the accuracy of the labels, which are used for model performance evaluation.
[0067] A2. Input training samples to multi-branch auxiliary module
[0068] The module explores the class-conditioned factors of PPLN by learning and predicting the progressive feature tendency, and its main structure contains multiple auxiliary branches equal in number to the number of classes, which work in parallel with the main branch. The auxiliary branches diverge from the tail of the backbone network (such as ResNet), and the head keeps sharing parameters. The tail diverges into 4 auxiliary branches according to the number of pneumoconiosis stages (such as no pneumoconiosis, stage 1, stage 2, and stage 3, a total of 4 stages), and each branch corresponds to a stage.
[0069] The module contains two stages in total, each of which starts with feature splitting, selects matching samples from a small batch of sample sets for potential neighboring label mining (PNLM) and progressive feature tendency determination (PFTD), respectively. Figure 2 As can be seen in b), the two stages maintain the same structure in the auxiliary branches, and each pair of corresponding branches shares the same parameters.
[0070] The specific steps of step A2 are as follows:
[0071] A21. Module input. The head of the backbone network (such as the first 3 layers of ResNet) is used as a feature extractor with shared parameters in the main branch and the auxiliary branch. For a given sample , the process is represented as
[0072]
[0073] The tail of the backbone network (such as the last layer of ResNet) diverges into multiple branches according to the number of classes , and each branch is denoted as , where . It should be noted that the feature extractors in the main branch and the auxiliary branch share parameters, and for simplicity, the main branch is indexed as , and accordingly the feature representation of sample in the i-th branch is
[0074]
[0075] where , is the feature dimension.
[0076] A22. Feature decomposition. The two-stage module learns and predicts the progressive feature bias by different processes, i.e., PNLM adopts samples of neighboring classes, while PFTD adopts samples of the corresponding class. Formally, for a given sample , all possible features matching to any stage in different branches are represented as follows
[0077]
[0078] where (1) and represent the features matching to the latent neighboring label mining and the progressive feature bias evaluation two stages, respectively; (2) denotes the features in the main branch. The probability vector corresponding to these features is represented as
[0079]
[0080] where and are the probabilities of the PNLM and PFTD stages, respectively, is the probability in the main branch.
[0081] A23. First stage: Latent neighboring label mining. In the first stage, the multi-branch auxiliary module utilizes auxiliary branches to alleviate the ambiguity of the latent inter-class boundary caused by the progressive pair label noise (PPLN), and according to the given label in the dataset, each sample is input into two branches corresponding to the classes contained in the generalized latent neighboring label set of the given label.
[0082] Branch input. The fundamental goal of latent neighboring label mining is to train high-precision binary classifiers (BCF). Each by adopting samples whose given label belongs to the generalized latent neighboring label set of the given label, to identify two latent neighboring labels in the set. For example, for a sample , whose GPNL set is , it will be input into branch 0 and branch 2, as shown in the example in Figure 2 . Conversely, samples with given labels of 0 and 2 are input into branch 1 to learn the implicit patterns of these two classes.
[0083] Progressive feature bias mining. In this stage, based on the hypothesis that the features of pneumoconiosis CXR will change with disease progression, each branch mines the progressive feature bias from the samples of its neighboring classes, for example, branch 1 mines the key discriminative features between class 0 and class 2. As Figure 3As shown, taking branch 1 as the starting point (corresponding to stage I pneumoconiosis), a relatively obvious feature tendency can be observed. That is, from the perspective of stage I pneumoconiosis, the samples labeled as non-pneumoconiosis or stage II input to this branch show a feature tendency of "mild" or "severe". Each branch acts as an "expert", only responsible for judging the progressive feature tendency of the input sample in two directions. After several iterations, all auxiliary branches can learn targeted "knowledge" highly related to the corresponding category and store it as parameters in the network structure.
[0084] Ideally, to acquire this specialized knowledge as accurately as possible, all Each classifier should be well-trained, using its output as a confidence score to measure the tendency of progressive features to distinguish samples from two adjacent classes, and the observed tendency of progressive features should strictly point to the two adjacent classes. However, under the PPLN setting, there are still several special cases that need to be discussed and handled, such as... Figure 4 As shown.
[0085] Scenario 1: Label noise within a given label. Due to the presence of PPLN in the dataset, the given label has a certain probability of being inaccurate. For example... Figure 4 As shown in “Scenario 1”, with branches For example, in principle, a branch is considered valid only if all samples in the input branch have the correct given labels. Only then can it be considered a well-trained binary classifier. In fact, even with label noise, the auxiliary branch still plays a role from the perspective of asymptotic feature bias. This is especially true when the given label is incorrect, for example, given a label... Some samples have potential neighboring labels that are and On the one hand, if the true labels of a portion of the samples are... Therefore, this portion of the samples generally has a label index less than [a certain value]. The samples have more similar latent semantic features compared to the branches. These samples, corresponding to different stages of pneumoconiosis, reflect symptoms that are "milder." On the other hand, if the true label of another portion of the samples is... This indicates that this portion of the samples is more likely to match the true label. The sample confusion, even with minimal interference, still indicates that they exhibit relatively "mild" feature characteristics. This well explains the underlying principle of the gradual feature tendency in the "mild" direction. Similarly, when the given label is... The time indicates a gradual tendency towards "serious" characteristics.
[0086] Scenario 2: Marginal category. If the sample tags belong Intuitively, it should have only one potential neighbor label, and these two categories are called marginal categories. The generalized potential neighbor label defined earlier is for this special case. For example, the GPNL set for category 0 is represented as... By treating category 0 itself as its own generalized potential neighbor label, Not only is the cardinality maintained at 2, but the structure of branch 0 is also made identical to that of other auxiliary branches. Furthermore, marginal branches are also significant in terms of latent feature tendencies. For example... Figure 4 As shown in “Scenario 2”, the sample It often has the slightest features among all training samples, while the sample This tendency tends to resemble samples with more severe characteristics (not the most severe, but at least moderate). Considering the presence of PPLN in Case 1, this tendency towards moderately severe characteristics also holds true. In Case 2, we call this tendency a marginal characteristic tendency, such as... Figure 5 As shown.
[0087] It has been demonstrated that, despite these two special cases, all auxiliary branches can independently learn progressive feature tendencies to enhance robustness to PPLN. To maintain consistency in the progressive feature tendencies learned by each branch of the model at this stage, a Branch-Consistent Focal Loss (BCFL) is introduced, defined as follows:
[0088]
[0089] In the formula, It is the first The focus loss of each auxiliary branch, In a small batch of sample sets The middle was entered into the first The number of samples in each branch, due to Represents tag space The index in the sample , yes The probability of belonging to class k in the main branch, and yes by The branch prediction belongs to the first... The probability of the class. Based on the original focal loss function, the two probability distributions are reconciled. In the above formula, the focal loss is defined as...
[0090]
[0091] in, and Represent two probability distributions, and It is a parameter of focus loss.
[0092] Furthermore, BCFL plays a crucial role in handling the occasional imbalance problem in mini-batch samples, which can lead to a biased learning tendency in an auxiliary branch. For example... Figure 6 As shown, the input branch In a given set of samples, the number of samples labeled as one category may far exceed the number of samples labeled as another. Therefore, BCFL enables each auxiliary branch to learn a balanced eigenvalue bias, thereby optimizing the next stage of the progressive eigenvalue bias evaluation process.
[0093] A24 Phase Two: Progressive Characteristic Tendency Assessment
[0094] In step A23, the main task of all auxiliary branches is to learn the progressive feature tendencies in both directions. In this stage, each branch continues to act as a binary classifier, evaluating the potential progressive feature tendencies of the input samples based on the "knowledge" acquired in the first stage.
[0095] Branch input and output. In the second stage, the auxiliary branches retain the same structure and parameters as the previous stage, except that the input samples for each branch are different. Each branch inputs samples of the form , that is, the samples are input into the branch corresponding to the category with their given label, so as to use a binary classifier to evaluate the potential tendency of these samples to be "mild" or "severe", such as Figure 7 As shown. (Refer to...) Figure 4 In this case, each branch at this stage will output a prediction vector containing two components.
[0096] Based on the asymptotic feature bias discussed above, each component of the vector represents a confidence score for the bias in the corresponding direction of the evaluation sample, intuitively reflecting the degree of PPLN. In fact, this score depicts the probability distribution of potentially adjacent true labels, and the bias direction should point towards the direction of the true label when not confused by PPLN.
[0097] Feature tendency confidence scores are regularized. For flexibility, the prediction confidence scores for evaluating the PFT are Sharpen regularized, as shown below.
[0098]
[0099] in, The temperature parameter determines the degree of sharpening or smoothing of the probability components. Ultimately, the confidence score obtained after the Sharpen operation is... This value was then used to calculate the latent propensity difference, which was then used as a threshold for dynamic sample selection based on the difference.
[0100] A3. Dual Dynamic Threshold Sample Selection Module
[0101] Sample selection aims to identify samples with potentially true labels to mitigate the memory effect of deep neural networks. To explore further evidence for measuring the strength of model memory and analyze the inherent instance dependence of PPLN, a Latent Tendency Difference (LT-Diff) is introduced based on asymptotic feature bias. To improve the sample selection strategy, this module introduces LT-Diff as a metric whose value is expected to decrease dynamically. Combined with two dynamic thresholds based on probability and difference, a dual-threshold sample selection method is formed.
[0102] A31. Calculate the latent propensity difference of the sample. (Sample) The LT-Diff is defined as
[0103]
[0104] in, Contains only two elements and , representing two generalized potential neighbor labels of the sample, Two corresponding components and These components represent and directly reflect the predicted feature tendencies in two directions, respectively. Given their equal importance, their absolute difference can serve as an indirect measure of asymptotic feature tendencies, thus reflecting the strength of the auxiliary branch's memory of the PPLN. Clearly, this memory of label noise should be minimized as much as possible; based on this fundamental principle, this module further introduces and calculates a dynamic threshold based on this difference.
[0105] A32. Dual Dynamic Thresholds. Both probability-based and difference-based dynamic thresholds are obtained through successive iterations, as follows:
[0106]
[0107] in, It is the iteration round index. and These are the backoff ratios for controlling both the degree of delay and the threshold stability, and they satisfy... The probability-based threshold follows the instance-specific dynamic threshold in the DISC method, as described in the first row of the equation above. Similarly, the dynamic threshold described in the second row of the equation is called the dynamic LT-Diff threshold. Although this dynamic LT-Diff threshold has a decreasing trend, its iterative form is consistent with that of the probability-based threshold. The above equation can be transformed to obtain:
[0108]
[0109] This indicates that, whether using probability or LT-Diff to subtract the threshold from the previous round, there should be a positive correlation between the round-by-round change in the threshold and the corresponding difference.
[0110] A33. Introducing the Double Layback-Ratio Trick (DLRT). Consider a very common mathematical principle: given two numbers, their sum is fixed, and when one number increases or decreases by 1 unit, their difference changes by 2 units. This property naturally exists when calculating the LT-Diff, so consider incorporating a probability-based threshold-related change factor. Double, satisfy Therefore, there is Incorporating it into the threshold iteration and setting a more stringent standard for sample selection can help further mitigate the memory effect of PPLN.
[0111] A34. A dual dynamic threshold sample selection strategy is used to partition the samples. This sample selection process evaluates all training samples in the dataset and divides them into two sets: a "clean set" and a "hard set." For each sample... The division criterion is based on the probability from the main branch. and potential propensity difference If a) is satisfied, its probability is greater than b) Its LT-Diff is less than Then the sample can be assigned to the "clean" set. In Chinese, its expression is as follows:
[0112]
[0113] This indicates that the selected set The samples not only contain key clean features that the main branch can learn, but also limit the memory strength of the auxiliary branches for the PPLN. If only one of the two conditions a) and b) above is true, then the sample can be classified into the "hard" set, as expressed below:
[0114]
[0115] This collection may contain, for example: Figure 1 The easily confused samples shown at the decision boundary demonstrate the inconsistency between the two evaluation criteria based on probability and difference, and require special handling to enhance the model's generalization ability.
[0116] A35. Introduce a loss function that includes an LT-Diff penalty term to update the parameters of the deep neural network. Subsequently, the set... and The samples in the dataset are processed in different ways according to different loss functions, and their expressions are as follows:
[0117]
[0118] in, , and Representing sets ,gather And the size of the entire dataset, Treated as a scalar value for the whole, and considered as the weighted learning rate. and Represents two penalty terms based on LT-Diff. and The coefficients are similar to the original terms in cross-entropy loss and generalized cross-entropy loss. Since LT-Diff should change inversely with the probability, It retains the unified form common to probability and further optimizes it by forcing the value of LT-Diff to decrease.
[0119] A4. Calculate the overall loss function
[0120] Based on the foregoing steps, the overall optimization objective of this invention is described by the following loss function:
[0121]
[0122] Among them: (1) It is the standard cross-entropy loss; (2) These are the coefficients used to balance the two loss functions for the "clean" set and the "hard" set; (3) and It's about round index. The weighted ramp function is expressed as follows:
[0123]
[0124] in, and They are two positive integers, and , By decreasing and increasing the overall loss function according to the form of a Gaussian curve, a preheating stage is introduced in the context of a ramp function. It is the index threshold for the number of warm-up rounds.
[0125] The advantages of the above setup are twofold. Firstly, the warm-up process forces all samples in the dataset to undergo several rounds of training using standard cross-entropy loss, making them susceptible to class-conditional label noise. It can be argued that in this multi-branch auxiliary model based on dual-threshold sample selection, the auxiliary branch acts as an "expert," often learning "knowledge" related to class-conditional factors in the PPLN. Furthermore, unlike confidence penalties and uncertainty assessment methods, the focus here should be on the "knowledge" within the auxiliary branch to more reasonably fit the potential distribution pattern of asymptotic pairwise label noise. Therefore, during the warm-up process… Keep it fixed and unchanged, while letting Increment until satisfied On the other hand, once the model has grasped the relatively simple patterns in PPLN that are highly correlated with class conditional factors, it should learn more complex "knowledge" in PPLN that is highly correlated with instance dependency factors through sample selection based on dual dynamic thresholds. The focus may shift to the loss functions for the "clean" and "hard" sets. Similarly, in subsequent rounds, It will decrease until the end of training, and Then it remains unchanged.
[0126] Experimental Analysis
[0127] Based on the assumptions and formal description of PPLN, it is assumed that the clinically collected pneumoconiosis CXR dataset should contain inherent label noise, referred to as the original PPLN, but its potential noise rate is still unclear. To further demonstrate the model's effectiveness, the experimental design also comprehensively considers the class-conditional factors and instance-dependent factors of PPLN. According to a given noise rate, i.e., the ratio of noisy labeled samples to the total number of samples, synthetic label noise is added to the dataset by mixing instance-dependent noise and instance-independent noise.
[0128] To delve deeper into the underlying reasons for the improved model performance of thresholding based on LT-Diff, histograms are used here to depict the changes in the threshold distribution based on LT-Diff with different PPLN noise rates over training iterations, as shown below. Figure 8As shown, the three rows correspond to PPLN noise rates of 10%, 20%, and 30%, respectively, with the vertical axis representing the proportion of samples within a specific threshold range for the corresponding instance. From the dynamic distribution changes of the histogram group, it can be observed that under all PPLN noise rates, the threshold distribution generally converges towards lower values, and after several rounds, it exhibits two centers. Models trained at higher PPLN noise rates tend to obtain a denser LT-Diff distribution at lower thresholds. From the perspective of LT-Diff, whether a sample is a "clean" sample locally depends on how its value changes between adjacent rounds. If the LT-Diff is temporarily less than the threshold, then it is considered a "clean" sample in that round. Globally, the number of times a sample is considered a "clean" sample reflects the decreasing trend of its threshold. Therefore, throughout the entire training round, samples closer to the low threshold center are more likely to be "clean" samples, while samples closer to the high threshold center are more likely to be "hard" samples. The remaining uncertainty stems from the combined selection strategy and the probability-based threshold.
[0129] This invention exhibits a slower performance degradation rate as the PPLN noise rate increases, such as... Figure 9 The confusion matrix shown illustrates the robustness of FTLTD-Net at various stages of pneumoconiosis. Stages without pneumoconiosis and stage III are less sensitive to lower PPLN noise rates, while the robustness of stages I and II decreases more significantly with increasing PPLN noise rates. This invention relates to an algorithm model that, in most cases, avoids classifying a sample as a non-adjacent category, thus preventing irrational classification. Results show that this invention maintains relatively stable robustness at all stages of pneumoconiosis with increasing noise rates. It significantly improves practical performance in handling the special label noise problem in the pneumoconiosis CXR dataset, contributing to further improving the accuracy and reliability of pneumoconiosis staging diagnosis in clinical practice.
[0130] It should be understood that those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
Claims
1. A classification system for chest X-ray images of pneumoconiosis oriented towards label noise, characterized in that, Considering both the class condition and instance dependency of PPLN, a deep neural network model is constructed to partition and classify samples by introducing progressive feature propensity analysis and calculating the potential propensity difference based on this propensity. This model is combined with a probability-based dynamic threshold and includes three modules: Main branch module: Similar to a typical deep neural network for classification tasks, it takes a chest X-ray image of pneumoconiosis as input and calculates and outputs a probability-based confidence score; Multi-branch auxiliary module: Contains auxiliary branches equal to the number of pneumoconiosis stages, working in parallel with the main branch. It operates through two stages: potential neighbor label mining and progressive feature propensity assessment. In the first stage, neighbor category images are input to each branch, and in the second stage, corresponding category images are input. The progressive feature propensity is learned and predicted, and a confidence score reflecting the sample's tendency towards "mild" or "severe" is output; Dual dynamic threshold sample selection module: Calculates the potential propensity difference based on the confidence score output by the multi-branch auxiliary module and combines it with a probability-based dynamic threshold to form a dual dynamic threshold sample selection mechanism, dividing the samples into a clean set and a hard set.
2. The system according to claim 1, characterized in that, The multi-branch auxiliary module consists of two stages, each starting with feature splitting from a mini-batch sample set. Matching samples were selected and used for Potential Neighboring Label Mining (PNLM) and Progressive Feature Tendency Determination (PFTD), respectively. Potential neighboring labels: for a given label... Its potential neighboring label (PNL) is represented as follows: The potential neighbor label set is defined as: The generalized latent neighbor label is defined as: These definitions describe the situation where adjacent stages of pneumoconiosis are easily confused and how this is reflected in the dataset. The tendency towards progressive features can be interpreted as a label... It may be related to it The degree of confusion of any GPN is quantified using a confidence score.
3. The system according to claim 2, characterized in that, The feature decomposition described herein involves a multi-branch auxiliary module divided into two stages, each employing different processes to learn and predict progressive feature tendencies. PNLM accepts samples from neighboring classes, while PFTD accepts samples from the corresponding class. Formally, for a given sample... The following represents all possible features in different branches that match any stage: ;in, and These represent features that match the two stages of potential neighbor label mining and progressive feature propensity evaluation, respectively. This represents the features in the main branch; the probability vectors corresponding to these features are represented as follows: ;in, and These are the probabilities of the PNLM and PFTD stages, respectively. It represents the probability in the main branch.
4. The system according to claim 1, characterized in that, In the PNLM stage of the multi-branch auxiliary module, a binary classifier is trained using auxiliary branches. Samples are input into branches corresponding to the set of generalized latent neighbor labels of a given label, learning progressive feature tendencies. Branch-Consistent Focal Loss (BCFL) is introduced to ensure the consistency of learning across branches, defined as: ;in, The progressive feature propensity assessment stage inputs the sample into the branch corresponding to the category with its given label, and outputs a prediction vector containing two components, which is then regularized by Sharpen to obtain the confidence score.
5. The system according to claim 1, characterized in that, In the dual dynamic threshold sample selection module, the absolute difference in the predicted feature tendencies of a sample in two generalized latent adjacent label directions is defined as the Latent Tendency Difference (LT-Diff). The LT-Diff is defined as: in, Contains only two elements and , representing two generalized potential neighbor labels of the sample, Two corresponding components and These represent and directly reflect the predicted characteristic tendencies in the two directions, respectively.
6. The system according to claim 5, characterized in that, Based on the aforementioned LT-Diff, a dynamic threshold is calculated for each sample according to the training iterations of the deep neural network. Simultaneously, a dynamic threshold is also calculated for each sample based on the probability-based confidence score output by the main branch module. Their iterative process is expressed as follows: in, It is the iteration round index. and These are the backoff ratios for controlling both the degree of delay and the threshold stability, and they satisfy... .
7. The system according to claim 6, characterized in that, The dual dynamic threshold sample selection module divides the samples into clean sets. and difficulties The classification criteria are given by the following formula: 。 8. The system according to claim 7, characterized in that, After dividing the training dataset into subsets according to the aforementioned sample partitioning criteria, losses are calculated using different loss functions to update the parameters of the deep neural network. Both loss functions incorporate an LT-Diff penalty term. Losses are calculated using the loss function for samples in the clean set and the loss function for samples in the hard set. The two loss functions are expressed as follows: in, , and Representing sets respectively ,gather And the size of the entire dataset, Treated as a scalar value for the whole, and considered as the weighted learning rate. and Represents two penalty terms based on LT-Diff. and The coefficients are similar to the original terms in cross-entropy loss and generalized cross-entropy loss.
9. The system according to claim 1, characterized in that, The overall loss function of the constructed deep neural network model is expressed as: Among them: (1) It is the standard cross-entropy loss; (2) These are the coefficients used to balance the two loss functions for the "clean" set and the "hard" set; (3) and It's about round index. The weighted ramp function is expressed as follows: in, and They are two positive integers, and , By decreasing and increasing the overall loss function according to the form of a Gaussian curve, a preheating stage is introduced in the context of a ramp function. It is the index threshold for the preheating round.