A multi-label classification method and device for medical images

By introducing a deep fusion method of image features and label topology in the medical image multi-label classification algorithm, the problem of neglected label correlation in the existing technology is solved, and the accuracy and robustness of multi-label classification are significantly improved.

CN119851048BActive Publication Date: 2025-06-10ANHUI UNIVERSITY OF TRADITIONAL CHINESE MEDICINE
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510329043.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-10
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

The existing medical imaging multi-label classification algorithm ignores the correlation between labels when dealing with multi-label classification, resulting in insufficient generalization capabilities of the model, especially in cross-center data heterogeneity and rare disease identification.

Method used

By designing a multi-label classification method, the image features are extracted using the trained convolutional neural network, and combined with the preset second feature, the semantic features and correlation information of each tag are characterized, and the prediction scores of each tag are calculated, thereby realizing the deep fusion of image features and tag topology.

Benefits of technology

This method not only considers the characteristics of the image itself, but also the correlation between the tags, significantly improving the accuracy and robustness of multi-label classification, especially when dealing with rare diseases and cross-center data heterogeneity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119851048B_ABST
    Figure CN119851048B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-label classification method and apparatus for images. The method includes: obtaining an image to be processed; extracting features of the image to be processed by using a trained first network to obtain first features; obtaining prediction scores of each label according to the first features and preset second features; and obtaining classification labels of the image to be processed according to the prediction scores of each label and a preset score threshold. This method extracts features of the image through a trained first network, and processes the features extracted by the first network according to the preset second features for characterizing semantic features of each label and their correlation information, so as to obtain prediction scores of all labels. When predicting, the prediction scores consider not only the features of the image itself, but also the correlation between each label, thereby realizing deep fusion of image features and label topology and making the multi-label classification result more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to a multi-label classification method and device for medical images. Background Art

[0002] With the breakthrough development of medical imaging technologies such as X-ray and computed tomography (CT), the application value of multi-modal imaging data in the entire clinical diagnosis and treatment cycle has become increasingly prominent. According to statistics, the global medical imaging data is expanding at a compound annual growth rate of 30%, which poses an urgent need for intelligent analysis of massive imaging data. Especially in the multi-label classification scenario, a single image may contain multiple pathological features (such as co-existing signs of pneumonia, lung nodules, and pneumothorax in chest X-rays), and there is an urgent need to establish a precise multi-label joint recognition system to achieve comprehensive diagnosis. Developing an efficient and reliable multi-label classification algorithm for medical images has become a core topic in the field of intelligent medicine.

[0003] Traditional image analysis methods rely on the manual observation and experience judgment of radiologists, which have two inherent defects: First, the efficiency of expert interpretation is limited by the manual processing speed (an average of 6 - 8 minutes per single case), making it difficult to cope with the exponentially growing diagnostic demands; Second, there are significant differences in the diagnostic coincidence rates of physicians with different qualifications (studies have shown that the detection differences of lung nodules between junior and senior physicians reach 15% - 20%), and the risk of misdiagnosis caused by subjective factors is prominent. Although the introduction of deep learning technology has partially solved the above problems, there are still key technical bottlenecks in the field of multi-label classification: (1) Existing algorithms mostly adopt an independent binary classification strategy, ignoring the strong correlation between pathological features in clinical practice, such as the correlation between pulmonary edema and abnormal cardiac function; (2) The long-tailed distribution of labels leads to a decrease in the recognition accuracy of the model for low-frequency diseases, such as the recognition rate of rare diseases is generally lower than 65%; (3) The cross-center data heterogeneity is significant, and the difference in Hounsfield Unit (HU) between different CT devices can reach 10% - 15%, resulting in insufficient model generalization ability.

[0004] The current mainstream solutions mainly alleviate the class imbalance problem by improving the loss function (such as weighted cross-entropy loss, Focal Loss, etc.). Typically, the Heuristic Stacking algorithm has increased the F1-score of rare diseases by nearly 10 percentage points through a dynamic sample weighting strategy; however, this does not solve the multi-classification problem. Summary of the Invention

[0005] In view of the above defects of the prior art, the present invention provides a multi-label classification method and device for medical images to solve the technical problem of ignoring the correlation between labels in the independent binary classification strategy.

[0006] To achieve the above and other related objectives, the present invention provides a multi-label classification method for medical images, including: obtaining an image to be processed; extracting features of the image to be processed by using a trained first network to obtain first features, wherein the first network is a convolutional neural network, and the first features are used to represent the visual features of the image to be processed; obtaining prediction scores of each label according to the first features and preset second features, wherein the second features are used to represent the semantic features of each label and their correlation information; and obtaining classification labels of the image to be processed according to the prediction scores of each label and a preset score threshold.

[0007] In an embodiment of the present invention, the first network includes a densely connected convolutional network integrating wavelet transform enhanced convolution.

[0008] In an embodiment of the present invention, the first network includes an initial vision layer, a wavelet transform dense block, a transition layer, and a global average pooling layer; extracting features of the image to be processed by using the trained first network to obtain first features includes: extracting low-level features of the image to be processed by using the initial vision layer; extracting multi-level features of the low-level features by using a plurality of the wavelet transform dense blocks; compressing and dimension-reducing the features output by the previous wavelet transform dense block by using the transition layer provided between two of the wavelet transform dense blocks; and compressing the spatial dimension of the features output by the last wavelet transform dense block to one by using the global average pooling layer to obtain the first features.

[0009] In an embodiment of the present invention, there are four wavelet transform dense blocks, and the four wavelet transform dense blocks respectively include 6, 12, 24, and 16 non-linear combination layers, and each non-linear combination layer includes batch normalization, an activation function, and wavelet transform enhanced convolution.

[0010] In an embodiment of the present invention, the second features are obtained through the following steps: obtaining a training set label file and a label word vocabulary according to a training set for training the first network; obtaining a correlation matrix according to the training set label file, and the correlation matrix is used to represent the correlation between labels; obtaining word vectors of all labels according to the label word vocabulary; and obtaining the second features according to the correlation matrix and the word vectors of all labels.

[0011] In one embodiment of the present invention, according to the training set label file, a correlation matrix is obtained, and the correlation matrix is used to characterize the correlation between labels, including: according to the training set label file, counting the number of times each pair of labels appear simultaneously to obtain a first matrix; preprocessing the first matrix to obtain a second matrix to reduce noise and redundant information; using a trained second network to process the second matrix to obtain a third matrix, where the second network dynamically models the relationship between label nodes through a self-attention mechanism; and optimizing the third matrix into a matrix suitable for use by a graph convolutional network to obtain the correlation matrix.

[0012] In one embodiment of the present invention, preprocessing the first matrix to obtain a second matrix includes: normalizing the first matrix to obtain a conditional probability matrix; filtering the conditional probability matrix according to a preset noise threshold to obtain a binary matrix; and reweighting the binary matrix according to preset hyperparameters to obtain the second matrix.

[0013] In one embodiment of the present invention, optimizing the third matrix into a matrix suitable for use by a graph convolutional network to obtain the correlation matrix includes: removing the first dimension of the third matrix to obtain a fourth matrix; adding an identity matrix to the fourth matrix to obtain a fifth matrix; calculating the degree matrix of the fifth matrix, taking the inverse matrix of the square root of the degree matrix, and then normalizing the fifth matrix using the inverse matrix to obtain the correlation matrix.

[0014] In one embodiment of the present invention, according to the correlation matrix and the word vectors of all labels, the second feature is obtained, including: using a trained third network to process the correlation matrix and the word vectors of all labels to obtain the second feature, where the third network is stacked by two layers of graph convolutional networks.

[0015] To achieve the above object and other related objects, the present invention further provides a multi-label classification device for medical images, including: a data acquisition unit for acquiring an image to be processed; a feature extraction unit for using a trained first network to extract the features of the image to be processed to obtain a first feature, where the first network is a convolutional neural network, and the first feature is used to characterize the visual features of the image to be processed; a score calculation unit for obtaining the predicted scores of each label according to the first feature and a preset second feature, where the second feature is used to characterize the semantic features and their correlation information of each label; and a label acquisition unit for obtaining the classification labels of the image to be processed according to the predicted scores of each label and a preset score threshold.

[0016] Advantages of the present invention: A multi-label classification method and device for medical images proposed by the present invention extracts features of an image through a trained first network, and processes the features extracted by the first network according to a second feature preset for characterizing semantic features of each label and their correlation information, so as to obtain prediction scores for all labels. When making predictions, the prediction scores consider not only the features of the image itself, but also the correlation between each label, thus realizing the deep fusion of image features and label topology, and making the multi-label classification result more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0018] Figure 1 Flowchart of the multi-label classification method provided in an embodiment of the present invention;

[0019] Figure 2 Structure diagram of the first network provided in an embodiment of the present invention;

[0020] Figure 3 Detailed flowchart of step S200 provided in an embodiment of the present invention;

[0021] Figure 4 Flowchart of the second feature acquisition step provided in an embodiment of the present invention;

[0022] Figure 5 Detailed flowchart of step S320 provided in an embodiment of the present invention;

[0023] Figure 6 Overall network structure diagram provided in an embodiment of the present invention;

[0024] Figure 7 Schematic diagram of the multi-label classification device provided in an embodiment of the present invention.

[0025] Explanation of reference numerals: 101, data acquisition unit; 102, feature extraction unit; 103, score calculation unit; 104, label acquisition unit. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] The following describes the implementation modes of the present invention through specific specific embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. In addition to the specific methods, devices, and materials used in the embodiments, according to the knowledge of those skilled in the art in the technical field and the description of the present invention, any methods, devices, and materials similar to or equivalent to those described in the embodiments of the present invention in the prior art can also be used to implement the present invention.

[0027] It should be understood that the terms used in the embodiments of the present invention are for the purpose of describing specific specific implementation schemes, rather than for limiting the protection scope of the present invention. Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those skilled in the technical field of the present invention.

[0028] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In some of these embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0029] The flowcharts and block diagrams in the accompanying drawings illustrate the architectures, functions, and operations that the methods and computer program products according to various embodiments disclosed in the present invention may implement. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0030] Please refer to Figure 1 , Figure 1 A multi-label classification method for medical images provided by an embodiment of the present invention, including steps S100 to S400.

[0031] Step S100: Obtain the image to be processed. The image to be processed can be, for example, a medical image taken by an X-ray or computed tomography (CT) device.

[0032] Step S200: Extract the features of the image to be processed using the trained first network to obtain first features. The first network is a convolutional neural network, and the first features are used to represent the visual features of the image to be processed.

[0033] In this step, it can be understood that the image to be processed needs to be preprocessed to adapt to the input size of the first network. A convolutional neural network (CNN) is a deep learning model specifically designed to process data with grid structures (such as images, videos, audio, etc.). Its core idea is to extract local features of the data through convolutional operations, reduce the feature dimension through pooling operations, and finally perform classification or regression through fully connected layers. In the present invention, the features of the image to be processed are extracted through a convolutional neural network, and the convolutional neural network can be, for example, network structures such as DenseNet, VGG, and ResNet.

[0034] In a specific embodiment of the present invention, the first network includes a dense connection convolutional network integrating wavelet transform enhanced convolution. The dense connection convolutional network (DenseNet) is an efficient convolutional neural network structure. Its core idea is to directly connect the output of each layer to the input of all subsequent layers through dense connections, enhancing feature reuse and gradient flow. This design reduces the number of parameters, alleviates the problem of gradient disappearance, and improves the model performance at the same time. The wavelet transform is a multi-scale signal analysis tool that can capture both the time-domain and frequency-domain information of the signal. Introducing the wavelet transform into the convolutional neural network to replace or supplement traditional convolutional operations to extract richer multi-scale features.

[0035] This embodiment incorporates the wavelet transform technology. Through wavelet decomposition, the first network can emphasize the low-frequency components in medical images and at the same time operate on multiple frequencies using compact convolutional kernels. This design enables the first network to better adapt to the noise and inhomogeneity in medical images, thereby significantly improving the robustness of classification. The introduction of the wavelet transform technology not only enhances the accuracy of feature extraction but also provides a more efficient and reliable solution for medical image analysis.

[0036] Please refer to Figure 2 , in a specific embodiment of the present invention, the first network includes an initial vision layer, a wavelet transform dense block, a transition layer, and a global average pooling layer. Step S200 includes steps S210 to S240.

[0037] Step S210: Extract the low-level features of the image to be processed using the initial vision layer. The initial vision layer is mainly used to extract low-level features, which, for example, includes convolutional layers and max pooling layers, etc.

[0038] Step S220: Extract multi-level features of low-level features using a number of wavelet transform dense blocks. The wavelet transform dense block is a core module designed in the present invention, mainly used for feature extraction of medical images. It extracts features through multiple non-linear combination layers (including batch normalization, ReLU activation function, and wavelet transform enhanced convolution), and concatenates the output of each layer with the input feature map in the channel dimension. This design can retain multi-level feature information and enhance the expression ability of the model.

[0039] Step S230: Use the transition layer set between two wavelet transform dense blocks to compress and reduce the dimension of the features output by the previous wavelet transform dense block. The transition layer is a transition module between wavelet transform dense blocks, mainly used for compression and dimensionality reduction of feature maps. By setting a transition layer between wavelet transform dense blocks, this alternating structure enables the first network to gradually extract more abstract and high-level features while maintaining efficient computation, and at the same time, the first network can better adapt to the complexity and diversity of medical image data, thereby improving the accuracy and generalization ability of classification.

[0040] Step S240: Use the global average pooling layer to compress the spatial dimension of the features output by the last wavelet transform dense block to one, obtaining the first feature. The global average pooling layer is mainly used to adjust the output feature size of the first network.

[0041] Please refer to Figure 3 , in a specific embodiment of the present invention, four wavelet transform dense blocks are provided. The four wavelet transform dense blocks respectively include 6, 12, 24, and 16 non-linear combination layers, and each non-linear combination layer includes batch normalization, activation function, and wavelet transform enhanced convolution. The first network will be described in detail below in combination with this structure.

[0042] (1) As mentioned before, the image to be processed needs to be preprocessed before being input into the first network. In this embodiment, for example, after adjusting the image to be processed to a size of 448×448, it then enters the initial vision layer.

[0043] (2) The initial vision layer first processes the input image through a common convolutional layer (convolution kernel size is 7×7, stride is 2) to obtain a feature map of 224×224×96, and then sequentially passes through batch normalization, ReLU activation function, and max pooling layer (convolution kernel size is 3×3, stride is 2) to obtain low-level features with a size of 112×112×96.

[0044] (3) The low-level features enter the first wavelet transform dense block. The first wavelet transform dense block processes the low-level features through 6 non-linear combination layers in sequence, and connects the outputs of each combination layer in a cross-channel manner. The output of the wavelet transform enhanced convolution and the input features are concatenated in the channel dimension, and finally a feature map with a size of 112×112×(96 + 6×48) is output. The wavelet transform dense blocks in the subsequent steps are repetitions of this structure.

[0045] (4) The transition layer includes a convolutional layer (with a kernel size of 1×1) and an average pooling layer (with a kernel size of 2×2). After the features in the previous step pass through the first transition layer, a feature map with a size of 56×56×(96 + 3×48) is output. All the transition layers in the first network adopt the same structure.

[0046] (5) After the features in the previous step pass through 12 non-linear combination layers in the second wavelet transform dense block in sequence, the output feature size is 56×56×(96 + 3×48 + 12×48).

[0047] (6) After the features in the previous step pass through the second transition layer, a feature map with a size of 28×28×(96 + 3×48 + 6×48) is output.

[0048] (7) After the features in the previous step pass through 24 non-linear combination layers in the third wavelet transform dense block in sequence, the output feature size is 28×28×(96 + 3×48 + 6×48 + 24×48).

[0049] (8) After the features in the previous step pass through the third transition layer, a feature map with a size of 14×14×(96 + 3×48 + 6×48 + 12×48) is output.

[0050] (9) After the features in the previous step pass through 16 non-linear combination layers in the fourth wavelet transform dense block in sequence, the output feature size is 14×14×(96 + 3×48 + 6×48 + 12×48 + 16×48).

[0051] (10) The feature map in the previous step finally passes through the global average pooling layer, compressing the spatial dimension of the feature map to 1, obtaining a one-dimensional first feature , where D represents the size of the first feature.

[0052] Step S300: Obtain the predicted scores of each label based on the first feature and a preset second feature, where the second feature is used to represent the semantic features of each label and their correlation information. The second feature mainly aims to introduce the semantic features of each label and their correlation information. It is the core output of the dynamic graph fusion mechanism. The second feature not only captures the dynamic correlation between labels but also realizes the joint optimization of label semantics and image visual features through fusion with the image feature matrix, thus significantly improving the accuracy and robustness of multi-label classification.

[0053] For a certain usage scenario, such as multi-label classification for chest medical images, the labels used have been determined, and the first network has also been trained. At this time, the semantic features of each label and their correlation information have been determined. Therefore, the second feature can be saved as a preset value. When predicting the image to be processed, the calculation can be directly performed according to the preset second feature.

[0054] Please refer to Figure 4 , in a specific embodiment of the present invention, the second feature is obtained through steps S310 to S340.

[0055] Step S310: Obtain the training set label file and the label word vocabulary based on the training set used to train the first network. Since the training set is required when training the first network, in this step, to ensure the correlation between the captured dynamic correlation between labels and the image visual features, the second feature is obtained step by step from the training set.

[0056] The present invention can adopt some existing data sets. The samples in the data set are composed of images and their corresponding labels. The data set is divided into a training set and a test set to complete the training and testing of the first network. In step S310, only the training set is used. The mapping relationship between the encoding of each image and its label is saved in the above-mentioned training set label file, and the label word vocabulary is a vocabulary containing all label words in the training set.

[0057] Step S320: Obtain the correlation matrix based on the training set label file, where the correlation matrix is used to represent the correlation between labels.

[0058] Please refer to Figure 5 , in a specific embodiment of the present invention, step S320 includes steps S321 to S324.

[0059] Step S321: According to the training set label file, count the number of times each pair of labels appear simultaneously to obtain the first matrix. In specific implementation, for example, nested loop statements can be used to traverse each line of data in the training set label file. The finally obtained first matrix can be denoted as M. , C represents the total number of labels. In this matrix, M ij represents the number of times the i-th label and the j-th label appear simultaneously among the multiple labels corresponding to the same image. Since each label does not count itself, the element values on the diagonal of the first matrix M are all 0.

[0060] Step S322, preprocess the first matrix to obtain a second matrix to reduce noise and redundant information. The preprocessing can include, for example, normalization, denoising, weighting, etc.

[0061] In a specific implementation of the present invention, step S322 specifically includes the following operations.

[0062] (1) Normalize the first matrix to obtain a conditional probability matrix. In this step, the label correlation and dependence are modeled in the form of conditional probability to obtain the conditional probability matrix M'. The conditional probability matrix M' is obtained by normalizing the first matrix M, and its element M ij ' represents the conditional probability that label j appears when label i appears.

[0063] (2) Filter the conditional probability matrix according to a preset noise threshold to obtain a binary matrix. This step is to avoid problems such as noise in the above conditional probability matrix M'. Therefore, the threshold τ can be used to filter the noisy edges to obtain the binary matrix A (A ij is the element value of the i-th row and the j-th column of matrix A):

[0064] .

[0065] (3) Re-weight the binary matrix according to a preset hyperparameter to obtain a second matrix. This step is mainly to solve the over-smoothing problem brought by the binary correlation matrix. Therefore, a re-weighting strategy can be used to obtain the second matrix N (N ij is the element value of the i-th row and the j-th column of matrix N):

[0066] ,

[0067] where p is a hyperparameter.

[0068] Step S323, process the second matrix using the trained second network to obtain a third matrix. The second network dynamically models the relationship between label nodes through the self-attention mechanism. The second network can be, for example, a GraphTransformer or a GraphSAGE network. Here, the Graph Transformer network is taken as an example to elaborate on step S323 in detail.

[0069] To obtain more abundant correlation structure information among label nodes, the Graph Transformer network is used to transform the obtained second matrix N into a new graph structure. The second matrix N obtained in step S322 is used as the input of the second network, and different Qs are obtained through different linear layers. i , K i , V i :

[0070] ,

[0071] ,

[0072] And through Q i , K i , V i Calculate the attention matrix:

[0073] ,

[0074] For the h-head attention layer, a subgraph G containing label node information can be obtained:

[0075] ,

[0076] For all subgraphs G obtained by different attention layers, the third matrix P can be obtained by matrix multiplication:

[0077] ,

[0078] In the above formulas, are all learnable parameters, Dh represents the dimension of each attention head, which is a hyperparameter, is the transpose of K i . The shape of the finally obtained third matrix P can be denoted as (1, num_classes, num_classes).

[0079] Step S324: Optimize the third matrix into a matrix suitable for use in the graph convolutional network to obtain the correlation matrix. In a specific embodiment of the present invention, step S324 specifically includes the following steps.

[0080] (1) Remove the first dimension of the third matrix to obtain the fourth matrix, and the shape of the fourth matrix is (num_classes, num_classes). This step is to remove the redundant batch dimension, simplify the matrix structure, and facilitate subsequent processing.

[0081] (2) Add the identity matrix to the fourth matrix to obtain the fifth matrix. The identity matrix is a matrix with diagonal elements equal to 1. The purpose of this step is to ensure that each node is connected to itself. In a graph convolutional network, the features of a node not only depend on its neighboring nodes but also on its own features. Adding the identity matrix can ensure that a node retains its own information during the information propagation process. Without self-loops, a node may lose its original features after multiple graph convolution operations.

[0082] (3) Calculate the degree matrix of the fifth matrix, take the inverse matrix of the square root of the degree matrix, and then use the inverse matrix to normalize the fifth matrix to obtain the correlation matrix. In this step, first calculate the degree matrix of the fifth matrix, that is, for each node, calculate its degree (i.e., the sum of each row of the adjacency matrix), and then take the reciprocal of its square root to obtain the degree matrix D; then take the inverse matrix of the square root of the degree matrix, and this inverse matrix can be denoted as D -1 / 2 ; Finally, use the inverse matrix to normalize the fifth matrix to obtain the final correlation matrix P'.

[0083] Step S330: Obtain the word vectors of all labels according to the label word vocabulary. A label is generally composed of multiple words or characters. The word vector of each label can be directly composed of the sum of the word vectors of the words or characters included in the label. For example, the word vector of the "lungopacity" category is composed of the sum of the word vectors of "lung" and "opacity". The word vectors can be trained by methods such as Glove, or FastText, or GoogleNews.

[0084] Step S340: Obtain the second feature according to the correlation matrix and the word vectors of all labels.

[0085] In a specific embodiment of the present invention, use the trained third network to process the correlation matrix and the word vectors of all labels to obtain the second feature. Among them, the third network is stacked by two layers of graph convolutional networks. The graph convolutional network (GCN, Graph Convolutional Network) is a neural network specifically used to process graph-structured data, and updates node features by aggregating the information of neighboring nodes. In this embodiment, the main purpose of using two layers of graph convolutional networks is to gradually propagate and aggregate label correlation information through multiple graph convolution operations, so as to more effectively model the relationship between label nodes.

[0086] The operation of the graph convolutional network can be expressed as:

[0087] ;

[0088] Among them, , which is the correlation matrix P' (obtained in step S320) with self-loops added, is the degree matrix of is the node feature of the l th layer, is the learnable weight matrix of the l th layer, and σ is the activation function (such as ReLU).

[0089] For the first graph convolutional network, l = 0, and in the formula corresponds to the labeled word vector obtained in step S330. What the first graph convolutional network outputs is ; for the second graph convolutional network, l = 1, and its final output is , that is, the second feature W. The second feature , where n is the number of classification labels.

[0090] In the step of obtaining the predicted scores of each label according to the first feature ( ) and the preset second feature ( ), the predicted score can be directly set. According to the sizes of the first feature and the second feature, the obtained , where each element value corresponds to the predicted score of a label respectively.

[0091] Step S400, according to the predicted scores of each label and the preset score threshold, obtain the classification label of the image to be processed. After obtaining the predicted scores, they can be compared with the preset score threshold. For example, the labels with scores exceeding 0.8 are the labels of the image to be processed.

[0092] In the above embodiments, the reason for taking the second feature as a preset value is that after the training of each network model is completed, the second feature is fixed and does not need to be recalculated. When the usage scenario changes or the labels change, the second feature needs to be recalculated.

[0093] Although in the above description of the steps, the first network, the second network, and the third network are described separately, in actual applications, the above three models together constitute an end-to-end model framework, and its structure is as Figure 6 shown (limited by the image size, some network structures are shown schematically and do not represent the actual quantity. For example Figure 6Only two wavelet transform dense blocks are schematically shown in the first network in [the figure] for the multi-label classification task of images. They need to complete training together. The training process is briefly introduced below. (1) Forward propagation: The first network is used to extract the features of the image to obtain the first feature x; the second network is used to generate the correlation matrix P'; the third network uses the correlation matrix P' and the label word vectors to propagate the label correlation information to obtain the second feature W. Finally, the first feature x and the second feature W are fused to obtain the prediction scores for multi-label classification. (2) Calculate the loss: Calculate the loss function according to the true labels and prediction results of the multi-label classification task. (3) Backward propagation: Calculate the gradients of the loss function with respect to the learnable parameters of the first network, the second network, and the third network through the backpropagation algorithm. (4) Update the parameters: Use the optimization algorithm to update the learnable parameters of all networks to minimize the loss function.

[0094] It should be noted that the step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as they contain the same logical relationship, they are all within the protection scope of this application; adding insignificant modifications to the algorithm or process or introducing insignificant designs, but not changing the core design of the algorithm and process are all within the protection scope of this patent.

[0095] Please refer to Figure 7 , Figure 7 A multi-label classification device for medical images provided in an embodiment of the present invention includes a data acquisition unit 101, a feature extraction unit 102, a score calculation unit 103, and a label acquisition unit 104. Among them, the data acquisition unit 101 is used to acquire the image to be processed; the feature extraction unit 102 is used to extract the features of the image to be processed by using the trained first network to obtain the first feature. The first network is a convolutional neural network, and the first feature is used to represent the visual features of the image to be processed; the score calculation unit 103 is used to obtain the prediction scores of each label according to the first feature and the preset second feature. The second feature is used to represent the semantic features of each label and their correlation information; the label acquisition unit 104 is used to obtain the classification labels of the image to be processed according to the prediction scores of each label and the preset score threshold.

[0096] It should be noted that the multi-label classification device in this embodiment is a device corresponding to the above multi-label classification method, and the functional modules in the multi-label classification device respectively correspond to the corresponding steps in the multi-label classification method. The multi-label classification device in this embodiment can be implemented in cooperation with the multi-label classification method. That is, without conflict, the relevant technical details mentioned in the above multi-label classification method can also be applied to the multi-label classification device in this embodiment.

[0097] Generally speaking, the present invention creatively proposes: (1) constructing a dynamic label correlation matrix and establishing a dynamic representation of disease associations through a graph attention mechanism; (2) designing a wavelet transform dense block that fuses wavelet transform enhanced convolutional kernels, whose multi-scale feature fusion ability significantly improves the accuracy of fine-grained feature extraction; (3) constructing a dual-branch heterogeneous network architecture, where the first network (the first branch) realizes image feature extraction, the third network (the second branch) initializes the label semantic space through Glove word vectors, generates a learnable adjacency matrix through the second network, and finally realizes the deep fusion of image features and label topology through a feature space alignment strategy, thereby further optimizing the model performance.

[0098] The application scope of the present invention is extensive. It is mainly aimed at multi-label classification tasks in the field of medical imaging and is applicable to the analysis and diagnosis of various medical imaging data such as chest X-ray images and CT images. In addition, it can also be applied to other scenarios with multi-label classification. Through the technology of the present invention, more accurate multi-label classification results can be provided, which is of great significance for medical image analysis. In actual medical images, a single image often may have multiple disease manifestations simultaneously. The present invention can more accurately capture the correlations between different disease labels through a method based on the correlation matrix, thereby significantly improving the classification accuracy. In addition, the present invention integrates a novel convolutional neural network structure, enhancing the model's ability in feature extraction, which helps to extract richer and more abstract image features and further improves the classification performance. Therefore, by incorporating the label correlation matrix, the present invention can effectively improve the accuracy of multi-label image classification and provide strong support for the accurate diagnosis of medical images.

[0099] The above embodiments are only illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.

Claims

1. A multi-label classification method for medical images, characterized in that: include: Get the image to be processed; Extracting features of the image to be processed using a trained first network to obtain a first feature, wherein the first network includes a densely connected convolutional network fused with wavelet transform enhanced convolution, and the first feature is used to characterize the visual features of the image to be processed; Obtaining a prediction score for each tag according to the first feature and a preset second feature, wherein the second feature is used to characterize the semantic features of each tag and its correlation information; Obtaining a classification label of the image to be processed according to the predicted score of each label and a preset score threshold; The second feature is obtained by the following steps: According to the training set used to train the first network, obtain a training set label file and a label word vocabulary; According to the training set label file, a correlation matrix is ​​obtained, wherein the correlation matrix is ​​used to characterize the correlation between labels; According to the label word vocabulary, word vectors of all labels are obtained; Obtaining the second feature according to the correlation matrix and the word vectors of all labels; According to the training set label file, a correlation matrix is ​​obtained, and the correlation matrix is ​​used to characterize the correlation between labels, including: According to the training set label file, the number of times each pair of labels appear at the same time is counted to obtain a first matrix; Preprocessing the first matrix to obtain a second matrix to reduce noise and redundant information; Processing the second matrix using the trained second network to obtain a third matrix, wherein the second network dynamically models the relationship between label nodes through a self-attention mechanism; The third matrix is ​​optimized to be a matrix suitable for use in a graph convolutional network to obtain the correlation matrix.

2. The multi-label classification method for medical images according to claim 1, characterized in that: The first network includes an initial visual layer, a wavelet transform dense block, a transition layer, and a global average pooling layer; The features of the image to be processed are extracted using the trained first network to obtain the first features, including: Extracting low-level features of the image to be processed using the initial visual layer; Extracting multi-level features of the low-level features using a plurality of the wavelet transform dense blocks; Using a transition layer disposed between two wavelet transform dense blocks, the features output by the previous wavelet transform dense block are compressed and dimensionally reduced; The global average pooling layer is used to compress the spatial dimension of the feature output by the last wavelet transform dense block to one, so as to obtain the first feature.

3. The multi-label classification method for medical images according to claim 2, characterized in that: There are four wavelet transform dense blocks, and the four wavelet transform dense blocks include 6, 12, 24, and 16 nonlinear combination layers respectively, and each of the nonlinear combination layers includes batch normalization, an activation function, and a wavelet transform enhanced convolution.

4. The multi-label classification method for medical images according to claim 1, characterized in that: Preprocessing the first matrix to obtain a second matrix includes: Normalizing the first matrix to obtain a conditional probability matrix; Filtering the conditional probability matrix according to a preset noise threshold to obtain a binary matrix; The binary matrix is ​​reweighted according to a preset hyperparameter to obtain the second matrix.

5. The multi-label classification method for medical images according to claim 4, characterized in that: Optimizing the third matrix to a matrix suitable for use in a graph convolutional network to obtain the correlation matrix includes: Removing the first dimension of the third matrix to obtain a fourth matrix; Adding a unit matrix to the fourth matrix to obtain a fifth matrix; The degree matrix of the fifth matrix is ​​calculated, and the inverse matrix of the square root of the degree matrix is ​​taken, and then the fifth matrix is ​​normalized by using the inverse matrix to obtain the correlation matrix.

6. The multi-label classification method for medical images according to claim 1, characterized in that: The second feature is obtained according to the correlation matrix and the word vectors of all labels, including: The correlation matrix and word vectors of all labels are processed using a trained third network to obtain the second feature, wherein the third network is formed by stacking two layers of graph convolutional networks.

7. A multi-label classification device for medical images, characterized in that: include: A data acquisition unit, used for acquiring an image to be processed; A feature extraction unit, configured to extract features of the image to be processed using a trained first network to obtain a first feature, wherein the first network includes a densely connected convolutional network fused with wavelet transform enhanced convolution, and the first feature is used to characterize visual features of the image to be processed; a score calculation unit, configured to obtain a prediction score for each tag according to the first feature and a preset second feature, wherein the second feature is used to characterize a semantic feature of each tag and its correlation information; and A label acquisition unit, used to obtain the classification label of the image to be processed according to the predicted score of each label and a preset score threshold; The second feature is obtained by the following steps: According to the training set used to train the first network, obtain a training set label file and a label word vocabulary; According to the training set label file, a correlation matrix is ​​obtained, wherein the correlation matrix is ​​used to characterize the correlation between labels; According to the label word vocabulary, word vectors of all labels are obtained; Obtaining the second feature according to the correlation matrix and the word vectors of all labels; According to the training set label file, a correlation matrix is ​​obtained, and the correlation matrix is ​​used to characterize the correlation between labels, including: According to the training set label file, the number of times each pair of labels appear at the same time is counted to obtain a first matrix; Preprocessing the first matrix to obtain a second matrix to reduce noise and redundant information; Processing the second matrix using the trained second network to obtain a third matrix, wherein the second network dynamically models the relationship between label nodes through a self-attention mechanism; The third matrix is ​​optimized to be a matrix suitable for use in a graph convolutional network to obtain the correlation matrix.

Citation Information

Patent Citations

  • Multi-label remote sensing image classification method with adaptive semantic information

    CN114821298A