Characteristic acquisition method of CXR image
By chunking high-resolution CXR images and combining graph Transformer and graph pooling technology, the problem of insufficient memory consumption and global context capture capabilities in high-resolution image processing is solved, and efficient and accurate image analysis is achieved.
Patent Information
- Application Number
- CN202510366741.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art consumes too much memory when processing high-resolution CXR images and has limited ability to capture global contexts. The Transformer model training process is complex and time-consuming, making it difficult to meet the dual requirements of real-time and accuracy.
The high-resolution CXR images are segmented into small pieces of subdivided images, and the low-level features are extracted using the convolutional token embedding component, and the high- and low-resolution image features are fused through graph Transformer, combined with graph pooling technology, local and global features are captured to reduce memory requirements and computational complexity.
It improves classification accuracy, multi-scale analysis capabilities, retains key details, enhances the robustness and generalization capabilities of the model, and is suitable for high-resolution image analysis.
Smart Images

Figure CN120298348A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical images, and particularly to a method for obtaining features of CXR images. Background Art
[0002] Chest X-ray (CXR) is the world's most popular medical imaging technology for diagnosing cardiopulmonary diseases. Timely CXR analysis is crucial for clinical care, but delays often occur due to overburdened healthcare center systems and a shortage of radiologists in rural areas. Automated CXR analysis can improve the efficiency of radiologists in detecting abnormalities and extend to underserved areas.
[0003] In recent years, deep learning has become a very successful medical image analysis tool, and many studies have been able to use deep learning for automatic chest radiograph diagnosis. Among these techniques, convolutional neural networks (CNNs) have proven to be particularly powerful. By using different convolutional kernels and network architectures, such as VGGNet, GoogLeNet, ResNet, and DenseNet, they can capture abstract representations of target images from different perspectives. CNNs can automatically extract features from images and perform efficient classification and detection, greatly improving the accuracy and efficiency of lung nodule detection. However, a key limitation of CNNs is their local bias, which hinders their ability to learn long-range dependencies in visual data. When dealing with high-resolution images, due to the limitation of their receptive fields, CNNs often have difficulty capturing long-range dependencies in images, which is particularly important in lung imaging analysis because information such as the location, size, and shape of lung nodules is crucial for diagnosis.
[0004] Although the Transformer architecture has become a promising alternative, it can handle long-range dependencies in sequence data well through the self-attention mechanism. For example, Keying et al. introduced RMTNet, a novel deep learning network that combines ResNet-50 with a Transformer architecture to capture local and long-range features in COVID-19 diagnosis. The model achieved high accuracy and efficiency on the adjusted 256×256×3 x-ray and CT image datasets, outperforming other models in terms of classification accuracy and detection speed; MedViT combines the locality of CNNs with the global connectivity of Transformers, proposing a robust and efficient CNN-Transformer hybrid model. To address the quadratic complexity of self-attention and improve cross-subspace information processing, it uses convolutional operations to design an efficient attention mechanism. However, Transformers rely on large-scale labeled datasets for pre-training. In the field of medical image classification, obtaining large-scale labeled samples is costly and challenging. Additionally, although publicly available chest x-ray datasets contain high-resolution images, due to GPU memory and training time limitations, most of the models discussed are trained on downscaled images. This makes it difficult for Transformers to meet the dual requirements of real-time performance and accuracy when processing high-resolution lung CT images.
[0005] Therefore, CNNs and Transformers, as well as methods combining them such as MedViT, have their respective advantages and disadvantages in automatic CXR analysis. For the MedViT method in the prior art, however, when dealing with high-resolution images or large images, it consumes excessive memory, and in some cases, its ability to capture global context is limited. The Transformer models therein usually contain a large number of parameters, which makes the model training process complex and time-consuming. Summary of the Invention
[0006] To overcome the problem in the above prior art that when dealing with high-resolution images or large images, it consumes excessive memory, and in some cases, its ability to capture global context is limited. The Transformer models therein usually contain a large number of parameters, which makes the model training process complex and time-consuming, the main object of the present invention is to provide a method for obtaining features of CXR images.
[0007] To achieve the above object, the present invention adopts the following technical solution. A method for obtaining features of chest x-ray images includes the following steps:
[0008] The obtained high-resolution chest x-ray image is divided twice to obtain small divided images;
[0009] Extract the low-level features of the small-block segmented image using a convolutional token embedding component, and use the low-level features as the initial input tokens. Adjust the feature dimension of the initial input tokens while introducing local inductive bias to obtain tokens with consistent dimensions.
[0010] Use the tokens with consistent dimensions as the nodes of the graph. Based on the eight-direction adjacency relationship of the tokens with consistent dimensions, establish an initial adjacency matrix, and optimize the initial adjacency matrix by obtaining weights to obtain an optimized adjacency matrix; the high-resolution chest X-ray image is processed through pooling to obtain a pooled feature.
[0011] Obtain the convolutional features of the low-resolution chest X-ray image; fuse the pooled features of the high-resolution chest X-ray image and the convolutional features of the low-resolution chest X-ray image through a graph Transformer based on cross-attention to obtain high-level features.
[0012] The high-level features are processed through graph pooling to obtain optimized high-level semantic features for classification processing.
[0013] The obtaining of the small-block segmented image includes:
[0014] Divide the obtained high-resolution CXR image into n different image patches, and each image patch is further divided into n individual small blocks to obtain a small-block segmented image.
[0015] The convolutional token embedding component includes two convolutional layers.
[0016] The extraction of the low-level features of the small-block segmented image using the convolutional token embedding component includes the following steps:
[0017] Adopt a 2D group convolution operation with a kernel size of 5×5 to share the convolution kernel across different feature channel groups, capture local context information, and obtain low-level features; the low-level features include edge features and texture features.
[0018] The feature dimension adjustment while introducing local inductive bias includes adjusting the dimension of the low-level features through linear projection, mapping the pixel values of each small block into a token representation with a fixed dimension. At the same time, in the convolutional layer, through the local receptive field characteristics of the convolution operation, pay attention to the spatial relationship of adjacent pixels, enhance the ability to model local structures, perform batch normalization and non-linear activation on the output token representation, and optimize the stability of the feature distribution to obtain low-level features with a fixed dimension.
[0019] Regarding the tokens with consistent dimensions as the nodes of the graph and establishing an initial adjacency matrix based on the eight-direction adjacency relationship of the tokens with consistent dimensions includes:
[0020] Regard each token with consistent dimensions as a node of a graph;
[0021] For each node, determine its adjacent nodes in eight directions;
[0022] Combine each node with its corresponding adjacent nodes to construct an N x N initial adjacency matrix, where N is the total number of nodes in the graph. Initially, all elements in the matrix are set to 0.
[0023] The obtaining of weights to optimize the initial adjacency matrix to obtain an optimized adjacency matrix includes the following steps:
[0024] Obtain weights based on the feature representation of the nodes and the spatial position information of the nodes;
[0025] Based on the obtained weights, create a weight matrix with the same size as the adjacency matrix. Initially, all elements are 0;
[0026] Traverse each node. For each node, check its adjacent nodes in eight directions. If there are adjacent nodes, mark the corresponding position in the adjacency matrix as 1 to obtain an updated weight value, and update the updated weight to the corresponding position in the adjacency matrix to obtain an optimized adjacency matrix.
[0027] The fusion of the pooled features of the high-resolution chest X-ray image and the convolutional features of the low-resolution chest X-ray image through a cross-attention-based graph Transformer to obtain high-level features includes the following steps:
[0028] Project the pooled features into query vectors;
[0029] Project the convolutional low-level features of the low-resolution image into key vectors and value vectors respectively;
[0030] Through the dot product of the query vector and the key vector, obtain the similarity score between the pooled features and the low-level features, introduce a learnable scaling factor, and apply the Softmax function to the similarity score to generate an attention weight matrix;
[0031] Use the attention weights to perform weighted summation on the value vectors to generate context-aware features as fusion features, add the fusion features and the original pooled features through residual connection, and output through layer normalization to obtain aggregated features;
[0032] Input the aggregated features into a feed-forward network for high-order interaction extraction to obtain a non-linear transformation result, and output the final high-level features through residual connection and normalization.
[0033] The processing of the high-level features through graph pooling includes the following steps:
[0034] Based on the embedding vectors of each node in the high-level features, a vector of weights is used to generate an importance score through a fully connected layer;
[0035] Combined with the optimized adjacency matrix, the importance scores of adjacent nodes are weighted and averaged to obtain the enhanced importance scores;
[0036] Sorted by the importance scores, the top k nodes are retained, and the retained nodes are hierarchically clustered to obtain similar nodes of multiple clusters, and the similar nodes of the clusters are merged into supernode features;
[0037] The supernode features are aggregated using feature pooling to obtain a new node set, and based on the new node set, the adjacency matrix is reconstructed to obtain the pooled features;
[0038] The pooled features are normalized and non-linearly activated through an activation function to obtain the optimized high-level semantic features.
[0039] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention can improve the classification accuracy, has the function of multi-scale analysis, can process and analyze images of multiple scales, capture features from local and global perspectives, and ensure the retention of key details. This method is particularly suitable for high-resolution images, effectively avoiding the loss of key details that may be caused by low-resolution images, thereby improving the classification accuracy. By fusing high-resolution and low-resolution images, and using graph pooling technology, the semantic parsing ability and processing efficiency of the model are further enhanced. This method not only retains the details in high-resolution images, but also utilizes the global information of low-resolution images, improving the robustness and accuracy of the model. Moreover, in the scenario of limited data, by fusing image information of multiple scales and adopting graph pooling technology, the risk of overfitting is effectively reduced, and the generalization ability of the model is improved. Even in the case of limited labeled data, it can provide accurate representations of the size, shape, and location of nodules for doctors, improving the diagnosis and treatment efficiency. Brief Description of the Drawings
[0040] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application.
[0041] Figure 1 is a schematic structural diagram of the present invention;
[0042] Figure 2 is a schematic flowchart of the present invention. Detailed Embodiments
[0043] In existing CXR analysis, CNNs and Transformers, as well as the methods combined with them such as MedViT, each have their own advantages and disadvantages in automatic CXR analysis. CNNs perform well in processing local features and classification tasks, but are limited by their receptive fields and local processing mechanisms; while Transformers can capture long-range dependencies, but face challenges such as high data requirements, high computational resource requirements, and training data limitations. Therefore, it is necessary to develop a more efficient and accurate automatic CXR analysis system.
[0044] The method closest to the present invention is the MedViT method, but this method consumes too much memory when processing high-resolution images or large images. And in some cases, its ability to capture global context is limited.
[0045] The Transformer models usually contain a large number of parameters, which makes the model training process complex and time-consuming. In addition, in order to prevent overfitting, a series of optimization strategies need to be adopted, which further increases the difficulty of training.
[0046] Referring to Figure 1 and Figure 2 The method of the present invention takes both high resolution and low resolution into account. The high-resolution CXR images are segmented into multiple squares, and each square is further segmented into smaller squares. Applying the graph Transformer, combined with graph convolutional projection, to process these two-layer image patches, reducing the number of input tokens through weight sharing in the hierarchical feature extraction process while maintaining the high-resolution characteristics of the images. The design of the initial adjacency matrix and the gradual expansion of the attention mechanism enable the model to focus on local features in the early stage while capturing global dependencies at higher levels. This local-to-global transformation helps to extract more reliable representations. By processing the images at different scales, the model can fuse features at different levels, thereby more comprehensively understanding the image content. Using two-dimensional group convolutional operations, by sharing convolutional kernels across different groups of feature channels, the memory requirements are effectively reduced, making the model more efficient in processing large-scale data. The introduction of techniques such as graph-pooling helps to reduce the number of tokens in the model, reduce the computational complexity, and at the same time promote more effective semantic information aggregation.
[0047] The present invention solves the problem of the limitations of local bias and remote dependence modeling. Traditional convolutional neural networks (CNNs) have made significant progress in the field of medical image analysis, but a key limitation of CNNs is their local bias, that is, they mainly focus on local features of images and are difficult to learn remote dependencies in images. This is particularly obvious when processing high-resolution chest X-ray images, because the significant differences in chest texture, shape, and size require the model to be able to capture more extensive information.
[0048] And the challenging issues of high-resolution image processing. Although high-resolution chest X-ray images provide rich details, most existing models are trained on downscaled images due to GPU memory and training time limitations. This approach may lead to the loss of critical details, thus affecting the classification accuracy. In addition, the processing of high-resolution images requires more computing resources, making it difficult to directly apply traditional methods to the clinical environment.
[0049] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0050] Embodiment:
[0051] As Figure 1 shown, first, the high-resolution CXR image is divided into n different blocks, and each block is further subdivided into n individual small patches. The graph Transformer extracts the small patches from each block as input tokens. Subsequently, the graph Transformer based on cross-attention uses all the pooled outputs from the high-resolution image and the convolutional features from the low-resolution image as inputs to derive high-level semantic features. The architecture of each module is as follows:
[0052] 1. Token Embedding: The convolutional token embedding component of the method of the present invention is used to capture the local context of the tokens present in high-resolution (784×784 or 1024×1024) and low-resolution (224×224) images, including the range from low-level edges to higher-order semantic primitives. This ability is achieved through a two-layer convolutional network. More specifically, for a given 2D image patch, the training function f(·) is used to meticulously convert each individual patch into a token. The function f(·) utilizes a 2D group convolutional operation with a kernel size of 5, effectively sharing the convolutional kernel across different groups of feature channels, thereby reducing memory requirements. The convolutional token embedding layer not only facilitates the adjustment of the token feature dimension but also introduces a local inductive bias at the initial stage of the model. This approach enhances the model's proficiency in accurately capturing the token context, especially in the earlier layers, ultimately increasing the overall performance.
[0053] 2. Graph Transformer: The token mapping input to each residual block is interpreted as graph data and processed by the graph convolutional layer. Essentially, the relationship between tokens in the mapping is captured by learning an adjacency matrix. In the bottom residual block of the graph Transformer, the initial adjacency matrix is established by recording the eight-directional adjacency relationship between tokens in the mapping. However, since the initial adjacency matrix only considers the preliminary form of the spatial relationship, it is necessary to consider the importance weights for a more detailed representation.
[0054] 3. Graph Transformer Based on Cross-Attention: The cross-attention mechanism plays a crucial role in capturing interactions between different sequences in the Transformer architecture. Specifically, the cross-attention layer computes attention scores by comparing queries with keys. These scores determine how much attention should be given to each position in the input sequence. The resulting attention weights are then used to compute a weighted sum of values, generating an attention-weighted representation that captures relevant information in the input. In the method of the present invention, the graph Transformer with cross-attention uses the pooled output from the high-resolution image and the convolutional features from the low-resolution image as inputs to derive high-level semantic features.
[0055] 4. Graph Pooling: Introducing graph pooling between consecutive blocks of the graph Transformer can reduce the number of nodes from n to t, thus significantly reducing the computational burden. Specifically, graph pooling reduces the scale of the graph by merging similar nodes or selecting representative nodes, while trying to preserve the important structural information of the graph. From the perspective of computational efficiency, compared with models based on Vision Transformer (ViT), using graph Transformer can significantly reduce the computational complexity. As shown in the formula:
[0056] where n represents the number of tokens or nodes, and d represents the size of the hidden layer. When n is much smaller than d, the method (GT) of the present invention has a significant reduction in computational complexity compared to ViT. Additionally, further reducing the number of tokens through graph pooling can more effectively reduce the computational cost.
[0057] It should be noted that in the present invention, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.
[0058] The above embodiments are merely illustrative examples of the present invention and do not constitute a limitation on the protection scope of the present invention. Any design identical or similar to the present invention falls within the protection scope of the present invention.
Claims
1. A method for obtaining features of a chest X-ray image, characterized in that, Including the following steps: The obtained high-resolution chest X-ray image is divided twice to obtain small-piece subdivision images; The convolutional token embedding component is used to extract the low-level features of the small-piece subdivision images, and the low-level features are used as the initial input tokens. The feature dimension of the initial input tokens is adjusted and a local inductive bias is introduced to obtain tokens with consistent dimensions; The tokens with consistent dimensions are used as the nodes of the graph. Based on the eight-direction adjacency relationship of the tokens with consistent dimensions, an initial adjacency matrix is established, and weights are obtained to optimize the initial adjacency matrix to obtain an optimized adjacency matrix; the high-resolution chest X-ray image is subjected to pooling processing to obtain pooling features; The convolutional features of the obtained low-resolution chest X-ray image; the pooling features of the high-resolution chest X-ray image and the convolutional features of the low-resolution chest X-ray image are fused through a graph Transformer based on cross-attention to obtain high-level features; The high-level features are subjected to graph pooling processing to obtain optimized high-level semantic features for classification processing.
2. The method for obtaining the features of a chest X-ray image according to claim 1, characterized in that, The obtaining of the small-piece subdivision images includes: The obtained high-resolution chest X-ray image is divided into n different image blocks, and each image block is subdivided into n individual small pieces to obtain small-piece subdivision images.
3. The method for obtaining the characteristics of a chest X-ray image according to claim 1, wherein, The convolutional token embedding component includes two convolutional layers; The using of the convolutional token embedding component to extract the low-level features of the small-piece subdivision images includes the following steps: A 2D group convolution operation with a kernel size of 5×5 is adopted to share the convolution kernel across different feature channel groups to capture local context information and obtain low-level features; the low-level features include edge features and texture features; The feature dimension adjustment and the introduction of a local inductive bias include adjusting the dimension of the low-level features through linear projection, mapping the pixel values of each small piece into a token representation with a fixed dimension, and at the same time, in the convolutional layer, through the local receptive field characteristic of the convolution operation, batch normalization and non-linear activation are performed on the output token representation to obtain low-level features with a fixed dimension.
4. The method for obtaining the features of a chest X-ray image according to claim 1, characterized in that, The using of the tokens with consistent dimensions as the nodes of the graph and establishing an initial adjacency matrix based on the eight-direction adjacency relationship of the tokens with consistent dimensions includes: Regarding each token with consistent dimensions as a node of the graph; Determining the adjacent nodes in its eight directions for each node; Combining each node and the corresponding adjacent nodes to construct an N x N initial adjacency matrix, where N is the total number of nodes in the graph. Initially, all elements in the matrix are set to 0.
5. The method for obtaining the characteristics of a chest X-ray image according to claim 4, wherein, The obtaining of weights to optimize the initial adjacency matrix to obtain an optimized adjacency matrix includes the following steps: Based on the feature representation of the nodes and the spatial position information of the nodes, weights are obtained; Based on the obtained weights, a weight matrix with the same size as the adjacency matrix is created, and all elements are 0 initially; Traverse each node. For each node, check its adjacent nodes in the eight directions. If there are adjacent nodes, mark the corresponding position in the adjacency matrix as 1 to obtain the updated weight value, and update the updated weight to the corresponding position in the adjacency matrix to obtain the optimized adjacency matrix.
6. The method for obtaining features of a chest X-ray image according to claim 1, wherein, Fusing the pooled features of the high-resolution chest X-ray image and the convolutional features of the low-resolution chest X-ray image through a graph Transformer based on cross-attention to obtain advanced features includes the following steps: Project the pooled features into query vectors; Project the convolutional low-level features of the low-resolution image into key vectors and value vectors respectively; Obtain the similarity scores between the pooled features and the low-level features through the dot product of the query vectors and the key vectors, introduce a learnable scaling factor, apply the Softmax function to the similarity scores to generate an attention weight matrix; Use the attention weights to perform weighted summation on the value vectors to generate context-aware features as fused features, add the fused features and the original pooled features through residual connection, and output through layer normalization to obtain aggregated features; Input the aggregated features into a feed-forward network for high-order interaction extraction to obtain a non-linear transformation result, and output the final advanced features through residual connection and normalization of the non-linear transformation result.
7. The method for obtaining the features of a chest X-ray image according to claim 1, wherein, The advanced features are processed by graph pooling, including the following steps: Based on the embedding vectors of each node in the advanced features, use a vector of weights to generate importance scores through a fully connected layer; Combine the optimized adjacency matrix, and perform weighted averaging on the importance scores of adjacent nodes to obtain enhanced importance scores; Sort by the importance scores, retain the top k nodes, perform hierarchical clustering on the retained nodes to obtain similar nodes in multiple clusters, and merge the similar nodes in the clusters into super-node features; Aggregate the super-node features using feature pooling to obtain a new node set, and reconstruct the adjacency matrix according to the new node set to obtain the pooled features; Perform normalization processing on the pooled features and perform non-linear activation through an activation function to obtain optimized advanced semantic features.