Computed tomography image processing method and application

Through the multi-scale feature extraction and multimodal alignment module combined with Swin UNETR and Electra model, the problem of insufficient accuracy and interpretation of liver cancer classification is solved, high-precision liver cancer classification is achieved and intuitive medical explanation is provided, and the robustness and clinical application value of the model are enhanced.

CN120298323APending Publication Date: 2025-07-11GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510341066.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art has problems in the classification of liver cancer classification with insufficient classification accuracy and insufficient interpretation of image and language description associations. Especially when dealing with liver lesions of diverse textures, shapes and sizes, the limitations of traditional CNN structures lead to degradation of classification performance, and existing methods fail to make full use of the rich association of vision-language pretrained models.

Method used

The multi-scale feature extraction module, multi-modal alignment module and classifier are used, combined with Swin UNETR and Electra models, and the multi-modal alignment module is used to achieve the matching of text and image embeddings, and the visual phrase matching module is used to generate medical interpretations, improving classification accuracy and providing intuitive medical interpretations.

Benefits of technology

The accuracy and stability of liver cancer classification are significantly improved, and the expression ability of visual embedding is enhanced through the multimodal alignment module, ensuring the consistency between image and text embedding, and automatically screening out relevant medical phrases and their confidence scores to provide a detailed explanation for clinical diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298323A_ABST
    Figure CN120298323A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and discloses a computed tomography image processing method and application, and the method comprises the steps: obtaining a computed tomography image and a corresponding image report; performing first preprocessing on the computed tomography image to obtain a preprocessed computed tomography image, and performing second preprocessing on the image report to obtain text embedding; dividing the preprocessed computed tomography image and text embedding into a training set, a verification set and a test set in proportion; training the classification model by using the training set, adjusting parameters of the trained classification model by using the verification set, and testing the classification model after parameter adjustment by using the test set to obtain a trained classification model; and inputting computed tomography data to be classified into the trained classification model to obtain a classification result and a corresponding medical description phrase. The liver cancer classification accuracy can be improved, and visual medical interpretation is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly relates to a method for processing computer tomography images and an application thereof. Background Art

[0002] At present, the goal of liver cancer classification is to construct an algorithm that can effectively distinguish primary liver cancer from metastatic liver cancer, so as to achieve early identification of liver cancer in medical examinations and improve the prognosis of patients. Among many medical images, computer tomography (CT) images are widely used in the diagnosis of liver tumors due to their high resolution and intuitive imaging advantages. In recent years, with the rapid development of deep learning technology, especially the application of convolutional neural networks (CNNs), significant progress has been made in liver cancer classification technology. For example, a framework based on a hierarchical convolutional neural network has been proposed for automatically detecting and classifying liver lesions in multi-phase CT; or a deep learning method has been used to perform five-class classification on liver masses in dynamic contrast-enhanced CT images, thus achieving a preliminary distinction of different types of liver cancer lesions.

[0003] However, the traditional CNN structure has certain limitations in improving encoding capabilities. To overcome this shortcoming, researchers have introduced hybrid structures in recent years to enhance the network feature expression capabilities and improve the classification and segmentation of tumors in the liver and other organs. Mainstream technologies such as U-Net and UNet++ use a symmetrical encoder-decoder structure and retain image detail information through jump connections. However, due to the inherent locality of convolution operations, they are limited in modeling long-range dependencies in images, especially when dealing with liver lesions with significant differences in texture, shape and size between patients, which can easily lead to a decrease in classification performance. To this end, new methods such as TransUNet and Swin UNETR introduce the global modeling capabilities of Transformer into the U-type network, and effectively improve the network's recognition ability for complex lesions by integrating long-range dependencies and multi-scale feature extraction. Although the above technologies have made great breakthroughs in encoding capabilities, they are still insufficient in the interpretability of image classification. Existing technologies mainly focus on the automatic extraction of image features and classification accuracy, but there is no sufficient and effective solution for how to mine the intrinsic relationship between images and language descriptions, and how to use this relationship to provide medical explanations for the final classification results. Some literature has tried to apply the vision-language pre-training model CILP to medical multimodal tasks, or to learn the association between visual features and word phrases through a CNN-based architecture, but these methods fail to fully utilize the rich associations between different receptive field features and text descriptions, resulting in insufficient explanation in clinical auxiliary diagnosis. Therefore, it is urgent to provide a new technical solution that can not only significantly improve the accuracy of liver cancer classification, but also generate explanations with clinical reference significance in combination with expert medical text information, thereby providing more comprehensive support for early clinical diagnosis and treatment decisions. Summary of the invention

[0004] The purpose of the present invention is to overcome the problems existing in the prior art and provide a method for processing computer tomography images. The present invention can improve the accuracy of liver cancer classification and provide an intuitive medical explanation.

[0005] To achieve the above object, in a first aspect, the present invention provides a classification model for computer tomography image processing. The classification model includes a multi-scale feature extraction module, a multi-modal alignment module, an aggregator, and a classifier. The multi-scale feature extraction module is configured to output an image embedding according to the input computer tomography image. The multi-modal alignment module is configured to output a matched image embedding, a plurality of medical phrases, and a confidence score corresponding to each medical phrase according to the image embedding and the text embedding. The aggregator is configured to generate a feature representation according to the matched image embedding, and aggregate the feature representation, the plurality of medical phrases, and the confidence score corresponding to each medical phrase to generate an aggregated feature. The classifier is configured to generate a classification result and a corresponding medical description phrase according to the aggregated feature.

[0006] Further, the multi-scale feature extraction module is configured to output an image embedding according to the input computer tomography image, specifically including: Input the computer tomography image into a Swin UNETR encoder to output a multi-scale feature map, denoted as ; Input the multi-scale feature map into a linear normalization layer, an adaptive pooling layer, and a three-dimensional convolutional layer in sequence to output multi-scale features, specifically as follows: ; Perform a splicing process on the multi-scale features, and input the spliced features into a fully connected layer to output an image embedding.

[0007] In a second aspect, the present invention further provides a method for computer tomography image processing. The processing method is based on the above-mentioned classification model for computer tomography image processing. The method includes the following steps: Obtain a computer tomography image and a corresponding image report; Perform a first preprocessing on the computer tomography image to obtain a preprocessed computer tomography image, and perform a second preprocessing on the image report to obtain a text embedding; Divide the preprocessed computer tomography image and the text embedding into a training set, a validation set, and a test set according to a ratio; Use the training set to train the classification model, use the validation set to adjust the parameters of the trained classification model, use the test set to test the classification model with adjusted parameters to obtain a trained classification model; Input the computer tomography data to be classified into the trained classification model to obtain a classification result and a corresponding medical description phrase.

[0008] Further, perform a first preprocessing on the computed tomography image to obtain the preprocessed computed tomography image. The first preprocessing specifically includes desensitization processing, naming standardization, HU value truncation, image cropping and slicing control, perspective unification, and / or window parameter adjustment.

[0009] Further, perform a second preprocessing on the image report to obtain text embeddings, specifically including: selecting multiple medical phrases from the image report and encoding the multiple medical phrases to obtain text embeddings.

[0010] Further, use a pre-trained Electra model to encode the multiple medical phrases to obtain text embeddings.

[0011] Further, the loss function for training the classification model includes a classification loss function, a contrastive loss function, and a matching loss function, specifically as follows:

[0012] Among them, is the contrastive loss function, is the matching loss, is the weight of the contrastive loss, is the weight of the matching loss, is the weight of the classification loss.

[0013] Further, the contrastive loss function is determined by the following formula:

[0014] Among them, is the boundary hyperparameter and , is the text embedding, is the image embedding, is the number of medical phrases, is the length of the embedding vector, is the th text embedding of the computed tomography image.

[0015] Further, the matching loss function is determined by the following formula:

[0016] Among them,, is the phrase existence label, is the Sigmoid function, is the matching confidence score.

[0017] In a third aspect, the present invention also provides an application of a method for processing computed tomography images, and the application is based on the method for processing computed tomography images and is used for liver cancer classification.

[0018] Compared with the prior art, the beneficial effects of the embodiments of the present invention are as follows: By using a fixed pre-trained Electra model to generate medical text embeddings and using SwinUNETR to extract multi-scale visual features of CT images, the present invention realizes the comprehensive extraction of two-modal information of text and image, ensures the full expression of their respective features, and significantly enhances the expression ability of visual embeddings, enabling the model to have higher robustness in processing the diversity of liver lesion textures, shapes, and sizes; also, through a multi-modal contrast learning module, it ensures the consistency of the processed visual embeddings and corresponding text embeddings in the semantic space, effectively promotes the complementarity of information between different modalities, and improves the overall classification accuracy and stability; furthermore, through a visual phrase matching module, it evaluates the matching situation between visual embeddings and text phrases, automatically screens out the most representative medical phrases and their confidence scores related to liver or liver cancer lesions, provides an intuitive medical interpretation for the classification results, and is convenient for clinicians to understand and apply. Description of the Drawings

[0019] Figure 1 is a flowchart of a method for processing computed tomography images according to Embodiment 1 of the present invention; Figure 2 is a block diagram of a computed tomography image processing system according to Embodiment 2 of the present invention; Figure 3 is a block diagram of a classification model according to Embodiment 1 of the present invention; Figure 4 is a schematic diagram of each module of the classification model according to Embodiment 1 of the present invention. Detailed Embodiments

[0020] The following combines the drawings and embodiments to further describe in detail the specific embodiments of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0021] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.

[0022] In the description of the present invention, it should be noted that unless otherwise clearly specified and defined, the terms "installed", "connected", "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0023] In addition, in the description of the present invention, unless otherwise stated, the meaning of "a plurality of" is two or more.

[0024] Embodiment 1 As Figure 1 shown, a computer tomography image processing method according to a preferred embodiment of the present invention includes the following steps: Step S1: Obtain a computer tomography image and a corresponding image report; Step S2: Perform a first preprocessing on the computer tomography image to obtain the preprocessed computer tomography image, and perform a second preprocessing on the image report to obtain a text embedding; Step S3: Divide the preprocessed computer tomography image and the text embedding into a training set, a validation set, and a test set according to a ratio; Step S4: Use the training set to train a classification model, use the validation set to adjust the parameters of the trained classification model, and use the test set to test the classification model with adjusted parameters to obtain a trained classification model; Step S5: Input the computer tomography data to be classified into the trained classification model to obtain a classification result and a corresponding medical description phrase.

[0025] Next, the technical solution of the present invention will be described in detail as follows: S1: Obtain a computer tomography image and a corresponding image report.

[0026] S2: Perform the first preprocessing on the computed tomography (CT) images to obtain the preprocessed CT images, and perform the second preprocessing on the image reports to obtain text embeddings. The first preprocessing includes: desensitization, naming standardization, HU value truncation, image cropping and slicing control, perspective unification, and window parameter adjustment. Specifically, first perform desensitization on the CT images to protect patient privacy and remove personal sensitive information (such as name, ID number, etc.) from the CT images, only retaining medical-related metadata. Then, label the desensitized CT images as primary liver cancer and metastatic liver cancer, and number the labeled CT images according to rules. This step (naming standardization) is to unify the data identifiers for subsequent processing. Then, truncate the Hounsfield unit (HU) values of the numbered CT images to the range [-200, 200]. The gray value (Hounsfield Unit, HU) range of CT images is extremely large (-1000 to +3000), but in practical applications, only specific tissues (such as the liver) need to be concerned. Truncating the HU values to [-200, 200] covers 99% of the liver area and removes the interference of irrelevant tissues (such as bones, air). Next, perform cropping and slicing control processing on the truncated CT images. The cropping standard is 256 256, and limit each CT image to contain 60 slices. The purpose is to unify the input size to meet the input requirements of the classification model and reduce the computational burden. Then, unify the perspective to the patient perspective with the back facing the observer (observing from top to bottom) to avoid misjudgment by the model due to direction confusion. Finally, adjust the window level to 40 - 60 and the window width to 300 - 400 to enhance the contrast between the liver and tumors, optimize the image display effect, highlight the lesion features, and obtain the preprocessed CT images.

[0027] Furthermore, perform the second preprocessing on the image reports to obtain text embeddings. Specifically, from the image reports, use the term frequency-inverse document frequency (TF-IDF) method to select the 30 most frequently occurring medical phrases, and then use the pre-trained Electra model to encode these 30 medical phrases to obtain text embeddings (one-hot encoding). The one-hot encoding is used to represent the presence of the selected phrases in the report (0 means absent, 1 means present).

[0028] S3: Divide the preprocessed CT images and text embeddings into a training set, a validation set, and a test set according to a certain proportion. In this embodiment, divide them into a training set, a validation set, and a test set according to the ratio of 8:1:1.

[0029] S4: Train a classification model using the training set, adjust the parameters of the trained classification model using the validation set, test the classification model with adjusted parameters using the test set to obtain a trained classification model. In this embodiment, as Figures 3-4 shown, the classification model includes a multi-scale feature extraction module, a multi-modal alignment module, an aggregator, and a classifier. The multi-scale feature extraction module is used to output an image embedding according to the input computed tomography image; the multi-modal alignment module is used to output a matched image embedding, multiple medical phrases, and a confidence score corresponding to each medical phrase according to the image embedding and the text embedding; the aggregator is used to generate a feature representation according to the matched image embedding, and aggregate the feature representation, the multiple medical phrases, and the confidence score corresponding to each medical phrase to generate an aggregated feature; the classifier is used to generate a classification result and a corresponding medical description phrase according to the aggregated feature.

[0030] Specifically, the multi-scale feature extraction module is used to output an image embedding according to the input computed tomography image, and specifically includes: Input the computed tomography image into the Swin UNETR encoder to output a multi-scale feature map, denoted as ; Input the multi-scale feature map into a linear normalization layer, an adaptive pooling layer, and a three-dimensional convolutional layer in sequence to output multi-scale features, specifically as follows: ; Perform a splicing process on the multi-scale features, and input the spliced features into a fully connected layer to output an image embedding.

[0031] The multi-modal alignment module includes: 1. A multi-modal contrastive learning sub-module: This sub-module uses a contrastive loss function, takes the text description corresponding to each CT image scan as a positive sample, and forms negative samples with other irrelevant texts. By calculating the similarity between the embeddings, it effectively distinguishes cross-modal positive and negative samples, so that the text and the image features are more closely aligned in the shared representation space. Specifically, in the multi-modal contrastive learning module, the contrastive loss function is defined as:

[0032] Among them, is a boundary hyperparameter and , is the text embedding, is the image embedding, is the number of medical phrases, is the embedding vector length, is the th text embedding of the Refers to the th embedded representation, refers to the th tensor in the embedded representation.

[0033] 2. Visual Phrase Matching Sub-module: Based on the aligned visual embeddings, each of the medical phrases is matched and evaluated through binary classification, automatically screening out the top three most representative medical phrases related to liver or liver cancer lesions, and calculating their confidence scores to provide the intuitive medical explanation for the final decision of the model. The loss function adopted by this sub-module during training is as follows: Matching loss function is determined by the following formula:

[0034] where,, is the phrase existence label, is the Sigmoid function, is the matching confidence score.

[0035] 3. Overall Alignment Mechanism: During the multi-modal alignment process, both forward and backward information flow mechanisms are introduced to ensure the two-way propagation of text and image embeddings in the shared space, further enhancing the cross-modal alignment effect.

[0036] The aggregator and classifier are used to integrate the aligned multi-modal features and perform the final liver cancer classification on the CT images; specifically, the aggregator: inputs the visually embedded representation processed by multi-modal alignment into the fully connected layer, generates the final feature representation through linear mapping; at the same time, fuses the medical phrases and their confidence scores output by the visual phrase matching module with this feature representation to form enhanced aggregated features. The classifier: uses the aggregated feature representation for liver cancer category prediction, and adopts the binary cross-entropy loss function to supervise the training of the classification results of primary and metastatic liver cancers, thus achieving a high-precision classification task.

[0037] During the aggregation process, through iterative optimization strategies such as gradient descent, the synergistic effects between sub-modules are continuously adjusted to ensure that the overall model continuously improves the cross-modal feature fusion and semantic alignment effects during training. Specifically, the total loss function is the classification loss , the contrast loss and the matching loss weighted sum of:

[0038] Finally, use the Adam gradient descent algorithm to update the network model parameters, thereby obtaining a trained code processing model.

[0039] S5: Input the CT image to be classified into the trained code processing model to obtain the liver cancer type and the corresponding medical description phrases.

[0040] In this embodiment, by using a fixed pre-trained Electra model to generate medical text embeddings and using SwinUNETR to extract multi-scale visual features of CT images, comprehensive extraction of two-modal information of text and image is realized, ensuring the full expression of their respective features, and significantly enhancing the expression ability of visual embeddings, making the model more robust in dealing with the diversity of liver lesion textures, shapes and sizes; also through the multi-modal contrast learning module, the consistency of the processed visual embeddings and the corresponding text embeddings in the semantic space is ensured, effectively promoting the complementarity of information between different modalities and improving the accuracy and stability of overall classification; also through the visual phrase matching module, the matching situation between visual embeddings and text phrases is evaluated, and the most representative medical phrases related to liver or liver cancer lesions and their confidence scores are automatically selected, providing an intuitive medical explanation for the classification results and facilitating clinical doctors to understand and apply.

[0041] Embodiment 2 As Figure 2 shown, an embodiment of the present invention further provides a computer tomography image processing system, including: an acquisition module: used to acquire computer tomography images and corresponding image reports; a preprocessing module: used to perform a first preprocessing on the computer tomography image to obtain the preprocessed computer tomography image, and perform a second preprocessing on the image report to obtain text embeddings; a partitioning module: partitioning the preprocessed computer tomography image and the text embeddings into a training set, a validation set and a test set according to a ratio; a training module: training the classification model using the training set, adjusting the parameters of the trained classification model using the validation set, testing the classification model with adjusted parameters using the test set to obtain a trained classification model; a classification module: inputting the computer tomography data to be classified into the trained classification model to obtain a classification result and the corresponding medical description phrases.

[0042] The system proposed in this embodiment is based on a computer tomography image processing method proposed in Embodiment 1. It can be understood that the optional items proposed in Embodiment 1 also apply to this embodiment and will not be elaborated here.

[0043] Embodiment 3 An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of a computer tomography image processing method are implemented.

[0044] In summary, the embodiments of the present invention provide a computer tomography image processing method, system, and storage medium. By using a fixed pre-trained Electra model to generate medical text embeddings and the Swin UNETR to extract multi-scale visual features of CT images, it realizes the comprehensive extraction of two-modal information of text and image, ensures the full expression of their respective features, significantly enhances the expression ability of visual embeddings, and enables the model to have higher robustness in dealing with the diversity of liver lesion textures, shapes, and sizes. It also ensures the consistency of the processed visual embeddings and the corresponding text embeddings in the semantic space through a multi-modal contrast learning module, effectively promotes the complementarity of information between different modalities, and improves the accuracy and stability of overall classification. It also evaluates the matching situation between visual embeddings and text phrases through a visual phrase matching module, automatically screens out the most representative medical phrases and their confidence scores related to liver or liver cancer lesions, provides an intuitive medical explanation for the classification results, and is convenient for clinicians to understand and apply.

[0045] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and replacements can be made, and these improvements and replacements should also be regarded as the protection scope of the present invention.

Claims

1. A classification model for computer tomography image processing, characterized in that, The classification model includes a multi-scale feature extraction module, a multi-modal alignment module, an aggregator, and a classifier. The multi-scale feature extraction module is used to output an image embedding according to the input computed tomography (CT) image; the multi-modal alignment module is used to output a matched image embedding, multiple medical phrases, and a confidence score corresponding to each medical phrase according to the image embedding and the text embedding; The aggregator is used to generate a feature representation according to the matched image embedding, and aggregate the feature representation, the multiple medical phrases, and the confidence score corresponding to each medical phrase to generate an aggregated feature; The classifier is used to generate a classification result and a corresponding medical description phrase according to the aggregated feature.

2. The classification model for computer tomography image processing according to claim 1, characterized in that, The multi-scale feature extraction module is used to output an image embedding according to the input computed tomography (CT) image, and specifically includes: Input the computed tomography image into the Swin UNETR encoder to output a multi-scale feature map, denoted as ; Sequentially inputting the multi-scale feature map into a linear normalization layer, an adaptive pooling layer, and a three-dimensional convolutional layer to output multi-scale features, specifically as follows: ; Performing a splicing process on the multi-scale features, and inputting the spliced features into a fully connected layer to output an image embedding.

3. A computer tomography image processing method, the processing method being based on a classification model for computer tomography image processing according to any one of claims 1 to 2, characterized in that, The method includes the following steps: Obtaining a computed tomography (CT) image and a corresponding image report; Performing a first preprocessing on the computed tomography (CT) image to obtain a preprocessed computed tomography (CT) image, and performing a second preprocessing on the image report to obtain a text embedding; Dividing the preprocessed computed tomography (CT) image and the text embedding into a training set, a validation set, and a test set according to a proportion; Training the classification model using the training set, adjusting the parameters of the trained classification model using the validation set, and testing the classification model with adjusted parameters using the test set to obtain a trained classification model; Inputting the computed tomography (CT) data to be classified into the trained classification model to obtain a classification result and a corresponding medical description phrase.

4. A computer tomography image processing method according to claim 3, characterized in that, Performing a first preprocessing on the computed tomography (CT) image to obtain a preprocessed computed tomography (CT) image. The first preprocessing specifically includes desensitization processing, naming standardization, HU value truncation, image cropping and slicing control, view angle unification, and / or window parameter adjustment.

5. A computer tomography image processing method according to claim 3, wherein, And performing a second preprocessing on the image report to obtain a text embedding, specifically including: selecting multiple medical phrases from the image report, and encoding the multiple medical phrases to obtain a text embedding.

6. A computer tomography image processing method according to claim 5, characterized in that, Encoding the multiple medical phrases using a pre-trained Electra model to obtain a text embedding.

7. A computer tomography image processing method according to claim 3, characterized in that, The loss function for training the classification model includes a classification loss function, a contrastive loss function, and a matching loss function, specifically as follows: Among them, is the contrast loss function, is the matching loss, is the weight of the contrast loss, is the weight of the matching loss, is the weight of the classification loss.

8. A method for processing computed tomography images according to claim 7, characterized in that, The contrastive loss function is determined by the following formula: wherein, is a boundary hyperparameter and , is a text embedding, is an image embedding, is the number of medical phrases, is the embedding vector length, is the text embedding of the th computed tomography image.

9. A method for processing computed tomography images according to claim 7, characterized in that, Matching loss function Determined by the following formula: Among them, is the phrase existence label, is the Sigmoid function, is the matching confidence score.

10. Application of a computer tomography image processing method, the application is based on the computer tomography image processing method according to any one of claims 3 to 9, characterized in that, For liver cancer classification.