A liver cancer detection system based on image recognition, an electronic device and a readable storage medium

CN118799298BActive Publication Date: 2026-10-09FOSHAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411019862.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-29
Publication Date
2026-10-09
Estimated Expiration
2044-07-29

AI Technical Summary

Technical Problem

但这些研究方法存在一定的局限性,首先上述实验设计都需要医生手动标注肿瘤位置,无法自动标注肿瘤感兴趣区域(Region ofInterest,ROI),使用手工方法定位肝癌目标费时费力;其次,网络用于鉴别肿瘤的特征有限

Benefits of technology

[0035] The image recognition-based liver cancer detection system of this invention receives the path to the liver cancer ultrasound data folder uploaded by the user through a liver cancer ultrasound data acquisition module, and reads the liver cancer ultrasound data according to the path. A liver cancer lesion localization module extracts B-mode ultrasound images from the liver cancer ultrasound data and uses a trained YOLOX-CBAM model to locate liver cancer lesions in the B-mode ultrasound images, obtaining a tumor ROI image. The tumor ROI image includes the ROI image of the B-mode ultrasound image and the ROI image of the CEUS image. A multimodal ultrasound fusion model classification module extracts brightness data from the tumor ROI image and uses the brightness data to fit a brightness curve, thus obtaining the ROI from the B-mode ultrasound image. The system selects the Region of Interest (ROI) image from a B-mode ultrasound image with a clear tumor boundary. A CEUS video is generated based on the ROI image of the CEUS image. The brightness curve, the ROI image of the B-mode ultrasound image with a clear tumor boundary, and the CEUS video are used as multimodal ultrasound data. A trained multimodal ultrasound fusion model is employed to classify the multimodal ultrasound data, resulting in a liver cancer classification prediction. The visualization analysis module uses a clustering and gradient-weighted class activation mapping visualization model to generate a liver cancer classification heatmap from the features learned from the multimodal ultrasound data. The prediction result output module outputs and displays the liver cancer classification prediction result, the liver cancer classification probability map, the liver cancer classification heatmap, and the brightness curve. This invention provides a convenient and easy-to-use image recognition-based liver cancer detection system. This system can achieve fully automatic localization and classification of liver cancer, providing doctors with a reference and helping them make more accurate diagnoses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118799298B_ABST
    Figure CN118799298B_ABST
Patent Text Reader

Abstract

The application provides a liver cancer detection system based on image recognition, an electronic device and a readable storage medium, and comprises a liver cancer ultrasound data acquisition module, a liver cancer lesion positioning module, a multi-modal ultrasound fusion model classification module, a visual analysis module and a prediction result output module.The liver cancer ultrasound data acquisition module receives a liver cancer ultrasound data folder path to read liver cancer ultrasound data.The liver cancer lesion positioning module extracts a B-ultrasound image from the liver cancer ultrasound data and obtains a tumor ROI image by positioning using a YOLOX-CBAM model.The multi-modal ultrasound fusion model classification module extracts brightness data from the tumor ROI image to fit a brightness curve, selects an ROI image of a B-ultrasound image with a clear tumor boundary, generates a CEUS video according to the tumor ROI image, obtains multi-modal ultrasound data, classifies the multi-modal ultrasound data using a multi-modal ultrasound fusion model to obtain a liver cancer classification prediction result, and generates a liver cancer classification heat map using a Grad-CAM visualization model.The prediction result output module outputs the classification prediction result.The application provides a liver cancer detection system based on image recognition, which is convenient and easy to use, provides a reference for doctors, and helps doctors make more accurate diagnoses.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical image processing technology, and in particular to an image recognition-based liver cancer detection system, an electronic device, and a readable storage medium. Background Technology

[0002] Primary liver cancer is the sixth most frequently diagnosed cancer worldwide and the third leading cause of cancer death. Primary liver cancer includes hepatocellular carcinoma (HCC) (accounting for 85%-90% of cases), intrahepatic cholangiocarcinoma (ICC), and other mixed hepatocellular-cholangiocarcinoma (CHC). HCC generally has a better prognosis, typically treated with chemotherapy, while non-HCC types such as ICC have a poorer prognosis, requiring surgical interventions such as liver resection for a possible cure. Because HCC and non-HCC differ in malignancy and require different treatments, accurate liver cancer screening is clinically significant in reducing misdiagnosis and overtreatment of HCC, and misdiagnosis and delayed treatment of non-HCC.

[0003] Contrast-Enhanced Ultrasound (CEUS) is the most commonly used imaging method for the clinical diagnosis of liver cancer. When performing CEUS examinations, clinicians use the Liverimaging Reporting and Data System (LI-RADS) to classify patients into standardized categories (LR-1 to LR-5). Cases in the LR-M category exhibit atypical imaging features, constituting atypical liver cancer, and are difficult to diagnose. Furthermore, this diagnostic method suffers from high subjectivity and is labor-intensive.

[0004] Currently, many researchers have used 3D-CNN to train CEUS image data to extract features that can be used to identify tumors. This network can not only learn the spatial features of tumor morphology and structure in CEUS images, but also monitor the temporal changes in the relevance of each frame. For example, Fengxin Pan et al. used 3D-CNN to design a CEUS-based computer-aided liver lesion identification system to distinguish between focal nodular hyperplasia (FNH) and HCC, achieving an experimental accuracy of 93.1%. Furthermore, the work of Tian Jie et al. demonstrated that 3D-CNN-based radiomics methods can more effectively utilize CEUS videos to accurately predict the response of hepatocellular carcinoma patients to transarterial chemoembolization (TACE). This method uses artificial intelligence to establish R-DLCEUS, R-BMode, and R-TIC models to quantitatively analyze CEUS data. The results showed that the R-DLCEUS model based on 3D-CNN improved the AUC by at least 12% compared to the other two models, and its heatmap visualization results provided some assistance to physicians in predicting TACE responses. More advanced studies have used multimodal learning to guide and assist networks in learning important features for tumor classification. For example, Qiuping Ma et al. combined CEUS, B-mode Ultrasound (BUS), and clinical data to build a model (2D-CNN, attention module, bidirectional long short-term memory module) for predicting recurrence of hepatocellular carcinoma after thermal ablation. Experiments showed that the combined model achieved an AUC of 84%, demonstrating good robustness and the ability to stratify and predict high-risk late-stage recurrence. Chen Chen et al. combined CEUS and US data to build a domain knowledge-driven deep learning model (backbone 3D-CNN, domain knowledge-guided temporal attention module, domain knowledge-guided channel attention module) for classifying breast tumors. Experiments showed that this model achieved a maximum sensitivity of 97.2% and a maximum accuracy of 86.3%, outperforming other comparable models.

[0005] Therefore, the approach of combining deep learning with radiomics in liver cancer diagnosis has been proven to uncover information that can assist doctors in diagnosis, making it easier and faster to diagnose and personalize treatment plans. However, these research methods have certain limitations. First, the experimental designs mentioned above require doctors to manually label the tumor location, and cannot automatically label the region of interest (ROI). Using manual methods to locate liver cancer targets is time-consuming and laborious. Second, the network has limited features for identifying tumors. It cannot fully utilize the multimodal information of radiomics to assist in the diagnosis of liver cancer. In addition, the single prediction result does not output effective features for doctors to refer to in diagnosis, and it remains at the research stage, making it difficult to implement a user-friendly computer-aided diagnostic system. Summary of the Invention

[0006] In view of the above problems, the present invention is proposed to provide an image recognition-based liver cancer detection system, an electronic device, and a readable storage medium that overcome or at least partially solve the above problems.

[0007] This invention provides a liver cancer detection system based on image recognition, comprising:

[0008] The liver cancer ultrasound data acquisition module is used to receive the liver cancer ultrasound data folder path uploaded by the user and read the liver cancer ultrasound data according to the liver cancer ultrasound data folder path.

[0009] The liver cancer lesion localization module is used to extract ultrasound images from the liver cancer ultrasound data and use a trained YOLOX-CBAM model to locate liver cancer lesions in the ultrasound images to obtain tumor ROI images; the tumor ROI images include the ROI images of the ultrasound images and the ROI images of the CEUS images.

[0010] The multimodal ultrasound fusion model classification module is used to extract brightness data from the tumor ROI image and fit a brightness curve using the brightness data, select the ROI image of the ultrasound image with clear tumor boundaries from the ROI image of the B-ultrasound image, generate a CEUS video based on the ROI image of the CEUS image, use the brightness curve, the ROI image of the ultrasound image with clear tumor boundaries and the CEUS video as multimodal ultrasound data, and use a trained multimodal ultrasound fusion model to classify the multimodal ultrasound data to obtain the liver cancer classification prediction result;

[0011] The visualization analysis module is used to generate liver cancer classification heatmaps from features learned from multimodal ultrasound data using a clustering and gradient-weighted class activation mapping visualization model.

[0012] The prediction result output module is used to output and display the liver cancer classification prediction results, liver cancer classification heatmap, and brightness curve.

[0013] Optionally, the liver cancer lesion localization module is further configured to convert the video data in the liver cancer ultrasound data into frame images, place the frame images and the image data in the liver cancer ultrasound data in chronological order and remove the CEUS image from the original ultrasound image to obtain a B-ultrasound image, and use a trained YOLOX-CBAM model to locate the liver cancer lesion in the B-ultrasound image to obtain a tumor ROI image.

[0014] Optionally, the multimodal ultrasound fusion model classification module is further used to extract brightness data of the tumor region and liver parenchyma region from the ROI image of the CEUS image, and to fit the corresponding brightness curves using the brightness data of the tumor region and liver parenchyma region respectively; the brightness data of the region is the average value of the region pixels in the image.

[0015] Optionally, the multimodal ultrasound fusion model classification module is also used to randomly select a ROI image with a clear tumor boundary from the ROI images of each patient's ultrasound images as the ultrasound image data of the multimodal ultrasound data.

[0016] Optionally, the multimodal ultrasound fusion model classification module is further configured to convert the ROI image of the CEUS image into a grayscale image, mark the moment when the tumor echo is first observed as the starting point of the CEUS video and take the image at the moment when the tumor echo is first observed as the starting frame of the CEUS video, and continuously extract image frames according to a preset first time interval until a preset first frame number threshold is reached, and continue to continuously extract image frames in subsequent frames according to a preset second time interval until a preset second frame number threshold is reached, and generate a CEUS video using the extracted images.

[0017] Optionally, the YOLOX-CBAM model includes a YOLOX module and a CBAM module, wherein the CBAM module includes a channel attention module (CAM) and a spatial attention module (SAM).

[0018] The channel attention module is used to apply max pooling and average pooling to the ultrasound image in the channel dimension to extract features. After sigmoid activation, the features are multiplied with the input feature map to obtain an enhanced CAM feature map.

[0019] The spatial attention module is used to perform max pooling and average pooling on each spatial location of the enhanced CAM feature map, concatenate the results and generate a SAM weight matrix through sigmoid activation, and multiply the SAM weight matrix and the enhanced CAM feature map to obtain a feature map with dual channel and spatial attention weights.

[0020] The YOLOX module uses the channel and spatial dual attention-weighted feature maps to locate liver cancer lesions and obtain tumor ROI images.

[0021] Optionally, the multimodal ultrasound fusion model includes a backbone composed of a CEUS model and a BUS model, a dual-channel feature fusion module and a time-intensity feature fusion module, and a fully connected layer.

[0022] The backbone composed of the CEUS model and the BUS model is used to extract CEUS modal features from the multimodal ultrasound data using the CEUS model and to extract BUS modal features from the multimodal ultrasound data using the BUS model.

[0023] The dual-channel feature fusion module is used to splice the extracted CEUS modal features and BUS modal features in the channel dimension and use the SE module to perform adaptive weighting to obtain the CEUS-BUS feature map;

[0024] The time intensity feature fusion module is used to divide the brightness data into time intervals to obtain the brightness change range, the brightness change range is a vector representing the degree of attention of the interval, and the CEUS-BUS feature map is compressed in the time dimension to obtain a CEUS-BUS one-dimensional vector. The brightness change range and the CEUS-BUS one-dimensional vector are multiplied to obtain the feature weight coefficient, and the feature weight coefficient is multiplied with the CEUS-BUS feature map to obtain a CEUS-BUS feature map with time attention.

[0025] The fully connected layer outputs liver cancer classification prediction results based on the CEUS-BUS feature map with time attention.

[0026] Optionally, the image recognition-based liver cancer detection system further includes an information management module, which comprises a result query submodule, a patient information management submodule, an operation history management submodule, an image information management submodule, and a user information management submodule.

[0027] The result query submodule is used to respond to user-inputted patient information keyword queries and display the corresponding liver cancer analysis results for the patient.

[0028] The patient information management submodule is used to display all patient information and perform add, delete, and modify operations on patient information;

[0029] The operation history management submodule is used to record and query the operation history of all users;

[0030] The image information management submodule is used to query the original image path, number of original images, storage path, and user notes for patients during diagnosis for all patients.

[0031] The user information management submodule is used to display all user information and supports users to edit and cancel their personal accounts, as well as support administrators to add, delete, and modify all users.

[0032] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the electronic device executes the computer program, it loads the liver cancer detection system based on image recognition as described in any one of the embodiments of the present invention.

[0033] The present invention also provides a readable storage medium having a computer program stored thereon, which, when executed by a processor, loads the image recognition-based liver cancer detection system as described in any one of the embodiments of the present invention.

[0034] This invention has the following advantages:

[0035] The image recognition-based liver cancer detection system of this invention receives the path to the liver cancer ultrasound data folder uploaded by the user through a liver cancer ultrasound data acquisition module, and reads the liver cancer ultrasound data according to the path. A liver cancer lesion localization module extracts B-mode ultrasound images from the liver cancer ultrasound data and uses a trained YOLOX-CBAM model to locate liver cancer lesions in the B-mode ultrasound images, obtaining a tumor ROI image. The tumor ROI image includes the ROI image of the B-mode ultrasound image and the ROI image of the CEUS image. A multimodal ultrasound fusion model classification module extracts brightness data from the tumor ROI image and uses the brightness data to fit a brightness curve, thus obtaining the ROI from the B-mode ultrasound image. The system selects the Region of Interest (ROI) image from a B-mode ultrasound image with a clear tumor boundary. A CEUS video is generated based on the ROI image of the CEUS image. The brightness curve, the ROI image of the B-mode ultrasound image with a clear tumor boundary, and the CEUS video are used as multimodal ultrasound data. A trained multimodal ultrasound fusion model is employed to classify the multimodal ultrasound data, resulting in a liver cancer classification prediction. The visualization analysis module uses a clustering and gradient-weighted class activation mapping visualization model to generate a liver cancer classification heatmap from the features learned from the multimodal ultrasound data. The prediction result output module outputs and displays the liver cancer classification prediction result, the liver cancer classification probability map, the liver cancer classification heatmap, and the brightness curve. This invention provides a convenient and easy-to-use image recognition-based liver cancer detection system. This system can achieve fully automatic localization and classification of liver cancer, providing doctors with a reference and helping them make more accurate diagnoses. Attached Figure Description

[0036] Figure 1 This is a structural block diagram of a liver cancer detection system based on image recognition provided in an embodiment of the present invention;

[0037] Figure 2 This is a schematic diagram of the login interface of the liver cancer detection system based on image recognition provided in an embodiment of the present invention;

[0038] Figure 3 This is a schematic diagram of the homepage of the liver cancer detection system based on image recognition provided in an embodiment of the present invention;

[0039] Figure 4 This is a schematic diagram of the data preprocessing interface of the liver cancer detection system based on image recognition provided in an embodiment of the present invention;

[0040] Figure 5 This is a schematic diagram of the liver cancer classification function interface of the liver cancer detection system based on image recognition provided in an embodiment of the present invention;

[0041] Figure 6 This is a schematic diagram of the result query interface of the liver cancer detection system based on image recognition provided in an embodiment of the present invention;

[0042] Figure 7 This is a schematic diagram of the patient information management interface of the liver cancer detection system based on image recognition provided in an embodiment of the present invention;

[0043] Figure 8 This is a schematic diagram of the operation history management interface of the liver cancer detection system based on image recognition provided in an embodiment of the present invention;

[0044] Figure 9 This is a schematic diagram of the image information management interface of the liver cancer detection system based on image recognition provided in an embodiment of the present invention;

[0045] Figure 10 This is a schematic diagram of the user information management interface of the liver cancer detection system based on image recognition provided in an embodiment of the present invention. Detailed Implementation

[0046] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0047] Reference Figure 1 The diagram illustrates a structural block diagram of an image recognition-based liver cancer detection system provided in an embodiment of the present invention, which may specifically include the following modules:

[0048] The liver cancer ultrasound data acquisition module is used to receive the liver cancer ultrasound data folder path uploaded by the user and read the liver cancer ultrasound data according to the liver cancer ultrasound data folder path.

[0049] The liver cancer lesion localization module is used to extract ultrasound images from the liver cancer ultrasound data and use a trained YOLOX-CBAM model to locate liver cancer lesions in the ultrasound images to obtain tumor ROI images; the tumor ROI images include the ROI images of the ultrasound images and the ROI images of the CEUS images.

[0050] The multimodal ultrasound fusion model classification module is used to extract brightness data from the tumor ROI image and fit a brightness curve using the brightness data, select the ROI image of the ultrasound image with clear tumor boundaries from the ROI image of the B-ultrasound image, generate a CEUS video based on the ROI image of the CEUS image, use the brightness curve, the ROI image of the ultrasound image with clear tumor boundaries and the CEUS video as multimodal ultrasound data, and use a trained multimodal ultrasound fusion model to classify the multimodal ultrasound data to obtain the liver cancer classification prediction result;

[0051] The visualization analysis module is used to generate liver cancer classification heatmaps from features learned from multimodal ultrasound data using a clustering and gradient-weighted class activation mapping visualization model.

[0052] The prediction result output module is used to output and display the liver cancer classification prediction results, liver cancer classification heatmap, and brightness curve.

[0053] This invention relates to an image recognition-based liver cancer detection system that integrates image processing, machine learning, and visualization analysis technologies to achieve fully automated localization and classification of liver cancer. Specifically, the system first verifies the legitimacy of the user's identity through a user login module. After successful login, the user is taken to the system homepage. Unregistered users can complete account registration through simple steps. Figure 2 This is a schematic diagram of the login interface of the liver cancer detection system based on image recognition of the present invention. Figure 3 This is a schematic diagram of the homepage of the image recognition-based liver cancer detection system of the present invention.

[0054] After entering the system homepage, users can upload the folder path containing the patient's liver cancer ultrasound data through the liver cancer ultrasound data acquisition module. Figure 4 This is a schematic diagram of the data preprocessing interface of the liver cancer detection system based on image recognition of the present invention. The system then reads these data and automatically performs liver cancer localization and classification. During this process, the liver cancer lesion localization module extracts ultrasound images from the liver cancer ultrasound data and uses a pre-trained YOLOX-CBAM model to efficiently and accurately localize liver cancer lesions in the ultrasound images, obtaining a region of interest (ROI) image. The ROI image can include the ROI image of the ultrasound image and the ROI image of the CEUS image.

[0055] To further enhance diagnostic reliability, the multimodal ultrasound fusion model classification module extracts brightness data from tumor ROI images and fits a brightness curve accordingly. This curve reflects the characteristics of tumor region changes over time. Next, the multimodal ultrasound fusion model classification module intelligently selects the ultrasound ROI image with the clearest tumor boundary as the ultrasound image data for multimodal ultrasound data, and generates a dynamic CEUS video based on the CEUS image. The brightness curve, the ROI image with a clear tumor boundary, and the CEUS video together constitute multimodal ultrasound data. This means that multimodal ultrasound data integrates multi-dimensional information, using three-dimensional CEUS video, two-dimensional ultrasound image, and one-dimensional brightness curve to provide effective features, thereby improving classification accuracy. This multimodal ultrasound data is input into the trained multimodal ultrasound fusion model for in-depth analysis, ultimately outputting the classification result for liver cancer. Figure 5 This is a schematic diagram of the liver cancer classification function interface of the liver cancer detection system based on image recognition of the present invention.

[0056] To facilitate doctors' intuitive understanding of the model's decision-making basis, the system is equipped with a visualization analysis module. This module uses clustering technology and a gradient-weighted class activation mapping (Grad-CAM) visualization model to transform features learned from multimodal ultrasound data into easily understandable liver cancer classification heatmaps, intuitively demonstrating the model's focus on the images. In the heatmap, the closer a region's color is to red, the more attention the model pays to that region, and the more likely that region is to contain features of classified liver cancer.

[0057] Finally, through the prediction result output module, the system clearly displays the liver cancer classification prediction results, probability map, liver cancer classification heatmap, and brightness curve to the user. This comprehensive output provides doctors with accurate diagnostic references, helping them make more precise diagnoses.

[0058] In one embodiment of the present invention, the liver cancer lesion localization module is further configured to convert the video data in the liver cancer ultrasound data into frame images, place the frame images and the image data in the liver cancer ultrasound data in chronological order and remove the CEUS image from the original ultrasound image to obtain a B-ultrasound image, and use a trained YOLOX-CBAM model to locate the liver cancer lesion in the B-ultrasound image to obtain a tumor ROI image.

[0059] In this embodiment of the invention, the liver cancer lesion localization module can process liver cancer ultrasound data and automatically annotate the tumor ROI region using a trained YOLOX-CBAM model. Specifically, the liver cancer lesion localization module first converts the liver cancer ultrasound data from video data into frame images, eight frames per second, corresponding to eight images per second. These frame images are then placed together with subsequent frame images (the original ultrasound image data in the liver cancer ultrasound data) in chronological order, and redundant regions, i.e., CEUS images, are removed, leaving only the ultrasound image. Then, the trained YOLOX-CBAM model is used to localize the liver cancer lesion on the remaining ultrasound image, thus obtaining the tumor ROI image.

[0060] The YOLOX-CBAM model was trained using a localization dataset with labeled ROIs. Specifically, an ultrasound examination was performed using an ultrasound diagnostic system, and 80 seconds of continuous DCM format video (8 frames per second) was recorded without changing any machine settings. After 80 seconds, a five-minute intermittent scan of the lesion was performed and recorded as images to characterize flushing features. For each patient, both ultrasound and CEUS examination results were stored, collecting a total of 131 HCC data points and 30 non-HCC data points. The collected video data was then converted into images (8 frames per second, corresponding to 8 images per second). These images were then placed chronologically along with subsequent frames, and redundant areas (i.e., CEUS images) were removed, leaving only the ultrasound images. After filtering, all patients' ultrasound images totaled 4830 images, which served as the localization dataset because ultrasound images contain more tumor texture information. Finally, the localization dataset was divided into training, validation, and test sets in a 6:2:2 ratio. The improved YOLOX-CBAM model network structure was trained using this dataset. The model takes a patient's ultrasound image as input and outputs a tumor ROI image (containing the tumor ROI image from the ultrasound image and the ROI image from the CEUS image), and the trained YOLOX-CBAM model is tested using a validation set and a test set.

[0061] In one embodiment of the present invention, the multimodal ultrasound fusion model classification module is further used to extract brightness data of the tumor region and the liver parenchyma region from the ROI image of the CEUS image, and to fit the corresponding brightness curves using the brightness data of the tumor region and the liver parenchyma region respectively; the brightness data of the region is the average value of the region pixels in the image.

[0062] In this embodiment of the invention, a CEUS (Continuous Ultrasound in the Tumor Region) image sequence can be extracted from the tumor ROI image. Brightness data of the tumor region and the liver parenchyma region are then extracted from the CEUS ROI image sequence, with the brightness data representing the average value of pixels in that region within the image. Two brightness curves are then generated using the brightness data of the tumor region and the liver parenchyma region, respectively, serving as one-dimensional data of the multimodal ultrasound data. The two brightness curves represent the changes in brightness of the tumor region and the liver parenchyma region over time, respectively.

[0063] In one embodiment of the present invention, the multimodal ultrasound fusion model classification module is further used to randomly select an ROI image with a clear tumor boundary from the ROI images of each patient's ultrasound images as the ultrasound image data of the multimodal ultrasound data, that is, the two-dimensional data of the multimodal ultrasound data.

[0064] In this embodiment of the invention, the ROI image sequence of the patient's B-ultrasound can be extracted from the tumor ROI image. For each patient, a ROI image with a clear tumor boundary is randomly selected from the B-ultrasound image ROI sequence as the BUS image data of the multimodal ultrasound data.

[0065] In one embodiment of the present invention, the multimodal ultrasound fusion model classification module is further configured to convert the ROI image of the CEUS image into a grayscale image, mark the moment when the tumor echo is first observed as the starting point of the CEUS video and take the image at the moment when the tumor echo is first observed as the starting frame of the CEUS video, and continuously extract image frames according to a preset first time interval until a preset first frame number threshold is reached, and continue to continuously extract image frames in subsequent frames according to a preset second time interval until a preset second frame number threshold is reached, and generate a CEUS video using the extracted images.

[0066] In this embodiment of the invention, the ROI sequence of the CEUS image in the tumor ROI image is converted into a grayscale image. The moment when the tumor echo is first observed is marked as the starting point of the CEUS video, and the image at this moment is used as the starting frame of the CEUS video. Then, image frames are continuously extracted according to a preset first time interval until a preset first frame number threshold is reached. In subsequent frames, image frames are continuously extracted according to a preset second time interval until a preset second frame number threshold is reached. Finally, these extracted images are used to generate a CEUS video as the three-dimensional data of multimodal ultrasound data.

[0067] In one embodiment of the present invention, the YOLOX-CBAM model includes a YOLOX module and a CBAM module, wherein the CBAM module includes a channel attention module (CAM) and a spatial attention module (SAM).

[0068] The channel attention module is used to apply max pooling and average pooling to the ultrasound image in the channel dimension to extract features. After sigmoid activation, the features are multiplied with the input feature map to obtain an enhanced CAM feature map.

[0069] The spatial attention module is used to perform max pooling and average pooling on each spatial location of the enhanced CAM feature map, concatenate the results and generate a SAM weight matrix through sigmoid activation, and multiply the SAM weight matrix and the enhanced CAM feature map to obtain a feature map with dual channel and spatial attention weights.

[0070] The YOLOX module uses the channel and spatial dual attention-weighted feature maps to locate liver cancer lesions and obtain tumor ROI images.

[0071] The improved YOLOX-CBAM model network addresses the issue that while using the multi-scale feature fusion method of the YOLOX model alone can effectively alleviate tumor heterogeneity and accurately locate small tumors, ultrasound images suffer from numerous noise spots and low contrast, resulting in a large number of redundant features in non-tumor regions in the feature map. To address this, this invention incorporates a CBAM (Convolutional Block Attention Module) into YOLOX to enhance the model's perception of lesion features. CBAM combines the Channel Attention Module (CAM) and the Spatial Attention Module (SAM). The CAM module extracts features along the channel dimension using max pooling and average pooling, and after sigmoid activation, multiplies them with the input feature map to obtain an enhanced CAM feature map. The SAM module performs max and average pooling on each spatial location of the CAM feature map, concatenates the results, and generates a SAM weight matrix through sigmoid activation. Finally, the SAM weight matrix is ​​multiplied with the CAM feature map to obtain the output of the CBAM module, i.e., a feature map weighted by both channel and spatial attention, thus enhancing the model's focus on tumor features.

[0072] The YOLOX-CBAM model trained based on the YOLOX-CBAM model network designed in this invention can locate liver cancer lesions in ultrasound images and obtain tumor ROI images. Specifically, the channel attention module of the YOLOX-CBAM model applies max pooling and average pooling to the ultrasound image in the channel dimension to extract features. After sigmoid activation, these features are multiplied with the input feature map to obtain an enhanced CAM feature map. The spatial attention module performs max pooling and average pooling on each spatial location of the enhanced CAM feature map, concatenates the results, and generates a SAM weight matrix through sigmoid activation. The SAM weight matrix is ​​then multiplied with the enhanced CAM feature map to obtain a feature map with dual channel and spatial attention weights. The YOLOX module uses the dual channel and spatial attention weighted feature map to locate liver cancer lesions and obtain tumor ROI images.

[0073] In one embodiment of the present invention, the multimodal ultrasound fusion model includes a backbone composed of a CEUS model and a BUS model, a dual-channel feature fusion module and a time-intensity feature fusion module, and a fully connected layer.

[0074] The backbone composed of the CEUS model and the BUS model is used to extract CEUS modal features from the multimodal ultrasound data using the CEUS model and to extract BUS modal features from the multimodal ultrasound data using the BUS model.

[0075] The dual-channel feature fusion module is used to splice the extracted CEUS modal features and BUS modal features in the channel dimension and use the SE module to perform adaptive weighting to obtain the CEUS-BUS feature map;

[0076] The time intensity feature fusion module is used to divide the brightness data into time intervals to obtain the brightness change range, the brightness change range is a vector representing the degree of attention of the interval, and the CEUS-BUS feature map is compressed in the time dimension to obtain a CEUS-BUS one-dimensional vector. The brightness change range and the CEUS-BUS one-dimensional vector are multiplied to obtain the feature weight coefficient, and the feature weight coefficient is multiplied with the CEUS-BUS feature map to obtain a CEUS-BUS feature map with time attention.

[0077] The fully connected layer outputs liver cancer classification prediction results based on the CEUS-BUS feature map with time attention.

[0078] This invention constructs a multimodal ultrasound fusion model within the PyTorch deep learning framework. The model consists of three parts: a backbone composed of a CEUS model and a BUS model; a dual-channel feature fusion module (DCFFM); and a temporal intensity feature fusion module (TIFFM). The specific design scheme is as follows:

[0079] (1) Using C3D and VGG13 models with fully connected layers and some convolutional layers removed, CEUS and BUS models were designed to extract ultrasound data features. Combining the characteristics of simultaneously observing CEUS and BUS images, a DCFFM adaptive fusion of the two modal features was designed. First, the features extracted by the CEUS and BUS models were concatenated along the channel dimension. Then, an SE module was used for adaptive weighting. Finally, a fully connected layer was used to output the classification results. This fusion strategy aims to fully integrate CEUS and BUS modal information to improve the performance of HCC diagnosis.

[0080] (2) In CEUS videos, each frame is considered equally important. However, in actual clinical practice, doctors tend to focus on keyframes with dynamic changes. Therefore, TIFFM is designed to make the model pay more attention to keyframes. First, the brightness data is divided into time intervals, and a vector representing the degree of attention for each interval is obtained from the range of brightness changes. Second, the CEUS feature map is compressed in the time dimension to form a one-dimensional vector. Finally, the two vectors are multiplied to obtain the feature weight coefficients. The weight coefficients are then multiplied with the feature map to obtain a feature map with temporal attention.

[0081] Specifically, after inputting multimodal ultrasound data into the multimodal ultrasound fusion model, the CEUS model in the multimodal ultrasound fusion model extracts CEUS modal features from the multimodal ultrasound data, and the BUS model extracts BUS modal features from the multimodal ultrasound data. The dual-channel feature fusion module concatenates the extracted CEUS modal features and BUS modal features in the channel dimension and uses the SE module for adaptive weighting to obtain the CEUS-BUS feature map. The temporal intensity feature fusion module divides the brightness data into brightness variation ranges at equal time intervals, and the brightness variation range is a vector representing the degree of attention of the interval. The CEUS-BUS feature map is compressed in the time dimension to obtain a one-dimensional CEUS-BUS vector. The brightness variation range is multiplied by the one-dimensional CEUS-BUS vector to obtain the feature weight coefficients. The feature weight coefficients are multiplied by the CEUS-BUS feature map to obtain a CEUS-BUS feature map with temporal attention. Finally, the fully connected layer outputs the liver cancer classification prediction result based on the CEUS-BUS feature map with temporal attention.

[0082] The multimodal ultrasound fusion model is trained using a pre-classified and labeled multimodal ultrasound dataset. Specifically, the localization dataset with labeled ROIs is processed to obtain the pre-classified and labeled multimodal ultrasound dataset.

[0083] The 3D CEUS video from the multimodal ultrasound dataset is obtained through the following steps:

[0084] 1) Convert the ROI sequence of CEUS images in the labeled ROI localization dataset into grayscale images;

[0085] 2) Mark the moment when the tumor echo is first observed as the starting point of the CEUS video, and the image at this moment is used as the starting frame of the CEUS video;

[0086] 3) After the initial frame, one frame is captured every 2 seconds, for a total of 25 frames. In subsequent frames, one frame is captured every 30 seconds, for a total of 5 frames. In total, each patient's CEUS video contains 30 frames.

[0087] Two-dimensional BUS images from a multimodal ultrasound dataset are obtained through the following steps:

[0088] For each patient, a clearly defined ROI image was randomly selected from the ultrasound image ROI sequence in the labeled ROI localization dataset and used as the BUS image data of the multimodal ultrasound data.

[0089] The one-dimensional brightness curve of the multimodal ultrasound dataset is obtained through the following steps:

[0090] 1) Extract the ROI sequence from each patient's CEUS image in the labeled ROI localization dataset;

[0091] 2) Extract brightness data of the tumor region and liver parenchyma region from the ROI sequence of the CEUS image; (the brightness data of the region is the average value of the pixels in that region in the image);

[0092] 3) Using the brightness data of the tumor area and the liver parenchyma area, two brightness curves are generated respectively; the two brightness curves represent the brightness changes of the tumor area and the liver parenchyma area over time.

[0093] After obtaining the one-dimensional, two-dimensional, and three-dimensional data of the multimodal ultrasound dataset, data augmentation was performed during the generation of the multimodal ultrasound dataset to compensate for the data imbalance caused by the scarcity of non-HCC data. Specifically, for CEUS videos, the images one second before and after the current frame of non-HCC cases were resampled to form new CEUS videos, thus tripling the amount of non-HCC video data. Subsequently, each frame of all case videos was horizontally flipped, doubling the HCC data and increasing the non-HCC data sixfold. This frame-by-frame resampling method effectively augments the data, achieving good results even with small datasets and data imbalance.

[0094] After obtaining the multimodal ultrasound dataset, the model can be trained using it, specifically by dividing the dataset into a training set:test set ratio of 8:2. The training set data is then input into the constructed multimodal ultrasound fusion model network, with an epoch set of 100 and an initial learning rate of 1e-4. To reduce the risk of overfitting, dropout layers with a dropout rate of 0.2 are set in the fully connected layers, and a cosine defiring learning rate adjustment strategy is used to optimize the training process.

[0095] In one embodiment of the present invention, the image recognition-based liver cancer detection system further includes an information management module, which comprises a result query submodule, a patient information management submodule, an operation history management submodule, an image information management submodule, and a user information management submodule.

[0096] The result query submodule is used to respond to user-inputted patient information keyword queries and display the corresponding liver cancer analysis results for the patient.

[0097] The patient information management submodule is used to display all patient information and perform add, delete, and modify operations on patient information;

[0098] The operation history management submodule is used to record and query the operation history of all users;

[0099] The image information management submodule is used to query the original image path, number of original images, storage path, and user notes for patients during diagnosis for all patients.

[0100] The user information management submodule is used to display all user information and supports users to edit and cancel their personal accounts, as well as support administrators to add, delete, and modify all users.

[0101] In this embodiment of the invention, the image recognition-based liver cancer detection system may further include an information management module. The information management module may further include a result query submodule, a patient information management submodule, an operation history management submodule, an image information management submodule, and a user information management submodule. Each submodule is specifically used for:

[0102] The results query submodule allows users to search for relevant patients and their predicted results by entering keywords through the "Results Query" interface. Figure 6 ;

[0103] Patient Information Management Submodule: The "Patient Information" interface displays all patient information and allows for adding, deleting, and modifying patient information. (Refer to...) Figure 7 ;

[0104] Operation History Management Submodule: The "Operation History" interface allows you to view the operation history of all system users. Figure 8 ;

[0105] Image Information Management Submodule: The "Image Information" interface displays the original image paths, number of original images, storage paths, and user-defined notes for each patient during diagnosis. Figure 9 ;

[0106] The User Information Management submodule displays all user information in the system on the "User Information" interface. Doctors can edit and cancel their personal accounts, while administrators can add, delete, and modify user information. (See also...) Figure 10 .

[0107] This invention develops a fully automated, image-recognition-based liver cancer detection system that is easy for clinicians to use. The interface is designed using PyQt5, and the database is designed using MySQL. It can manage and query information related to the operator and the patient. Inputting the patient's CEUS and BUS clinical data yields a multimodal model-predicted liver cancer type, the patient's image brightness curve, and the tumor ROI region and its characteristic heatmap labeled by the target detection network. Furthermore, an improved YOLOX-CBAM model is used to capture multi-scale tumor information, alleviate positive sample imbalance and tumor heterogeneity problems, improve model robustness, and automatically identify tumor regions in ultrasound images (B-mode ultrasound images), effectively overcoming the time-consuming and subjective drawbacks of manual localization, making it faster than existing technologies. A deep learning model is built based on the characteristics of ultrasound data, using a multimodal fusion method to simulate the diagnostic mode of radiologists, providing more features for liver cancer classification and improving classification accuracy. An attention mechanism is used to fuse information from various dimensions of the image, extracting useful information from each dimension to screen for atypical liver cancers. This model can fully utilize multimodal ultrasound data to classify liver cancer, which is superior to single-modal ultrasound methods. It uses features learned by the Grad-CAM visualization model to generate heat maps. At the same time, the heat maps and brightness curves are added to the system to provide more quantitative information to assist doctors in diagnosis.

[0108] Based on the same inventive concept, another embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to load the image recognition-based liver cancer detection system as described in any embodiment of the present invention.

[0109] Specifically, the electronic device includes a memory and a processor, which are connected via a bus. The memory stores a computer program that can run on the processor to load the image recognition-based liver cancer detection system according to any one of the first aspects of the present invention.

[0110] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0111] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0112] Based on the same inventive concept, another embodiment of the present invention provides a computer-readable storage medium having a computer program / instructions stored thereon, which, when executed by a processor, loads the image recognition-based liver cancer detection system as described in any of the first aspects of the present invention.

[0113] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0114] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, electronic devices, storage media, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0115] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

[0116] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0117] The foregoing has provided a detailed description of the image recognition-based liver cancer detection system, electronic device, and readable storage medium provided by this invention. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are merely for the purpose of helping to understand the method and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application. The above embodiments are merely preferred embodiments given to fully illustrate this invention, and the scope of protection of this invention is not limited thereto. Equivalent substitutions or modifications made by those skilled in the art based on this invention are all within the scope of protection of this invention.

Claims

1. A liver cancer detection system based on image recognition, characterized in that, include: The liver cancer ultrasound data acquisition module is used to receive the liver cancer ultrasound data folder path uploaded by the user and read the liver cancer ultrasound data according to the liver cancer ultrasound data folder path. The liver cancer lesion localization module is used to extract ultrasound images from the liver cancer ultrasound data and use a trained YOLOX-CBAM model to locate liver cancer lesions in the ultrasound images to obtain tumor ROI images; the tumor ROI images include the ROI images of the ultrasound images and the ROI images of the CEUS images. The multimodal ultrasound fusion model classification module is used to extract brightness data from the tumor ROI image and fit a brightness curve using the brightness data, select the ROI image of the ultrasound image with clear tumor boundaries from the ROI image of the B-ultrasound image, generate a CEUS video based on the ROI image of the CEUS image, use the brightness curve, the ROI image of the ultrasound image with clear tumor boundaries and the CEUS video as multimodal ultrasound data, and use a trained multimodal ultrasound fusion model to classify the multimodal ultrasound data to obtain the liver cancer classification prediction result; The visualization analysis module is used to generate liver cancer classification heatmaps from features learned from multimodal ultrasound data using a clustering and gradient-weighted class activation mapping visualization model. The prediction result output module is used to output and display the liver cancer classification prediction results, liver cancer classification heatmap, and brightness curve; The YOLOX-CBAM model includes a YOLOX module and a CBAM module. The CBAM module includes a channel attention module (CAM) and a spatial attention module (SAM). The channel attention module is used to apply max pooling and average pooling to the ultrasound image in the channel dimension to extract features. After sigmoid activation, the features are multiplied with the input feature map to obtain an enhanced CAM feature map. The spatial attention module is used to perform max pooling and average pooling on each spatial location of the enhanced CAM feature map, concatenate the results and generate a SAM weight matrix through sigmoid activation, and multiply the SAM weight matrix and the enhanced CAM feature map to obtain a feature map with dual channel and spatial attention weights. The YOLOX module uses the channel and spatial dual attention-weighted feature maps to locate liver cancer lesions, obtaining tumor ROI images.

2. The liver cancer detection system based on image recognition according to claim 1, characterized in that, The liver cancer lesion localization module is also used to convert the video data in the liver cancer ultrasound data into frame images, place the frame images and the image data in the liver cancer ultrasound data in chronological order and remove the CEUS image from the original ultrasound image to obtain a B-ultrasound image, and use a trained YOLOX-CBAM model to locate the liver cancer lesion in the B-ultrasound image to obtain a tumor ROI image.

3. The liver cancer detection system based on image recognition according to claim 1, characterized in that, The multimodal ultrasound fusion model classification module is also used to extract brightness data of the tumor region and liver parenchyma region from the ROI image of the CEUS image, and to fit the corresponding brightness curves using the brightness data of the tumor region and liver parenchyma region respectively; the brightness data of the region is the average value of the region pixels in the image.

4. The liver cancer detection system based on image recognition according to claim 1, characterized in that, The multimodal ultrasound fusion model classification module is also used to randomly select a ROI image with a clear tumor boundary from the ROI images of each patient's ultrasound images, as the ultrasound image data of the multimodal ultrasound data.

5. The liver cancer detection system based on image recognition according to claim 1, characterized in that, The multimodal ultrasound fusion model classification module is also used to convert the ROI image of the CEUS image into a grayscale image, mark the moment when the tumor echo is first observed as the starting point of the CEUS video and take the image at the moment when the tumor echo is first observed as the starting frame of the CEUS video, and continuously extract image frames according to a preset first time interval until a preset first frame number threshold is reached, and continue to continuously extract image frames in subsequent frames according to a preset second time interval until a preset second frame number threshold is reached, and use the extracted images to generate the CEUS video.

6. The liver cancer detection system based on image recognition according to claim 1, characterized in that, The multimodal ultrasound fusion model includes a backbone composed of a CEUS model and a BUS model, a dual-channel feature fusion module and a time-intensity feature fusion module, and a fully connected layer. The backbone composed of the CEUS model and the BUS model is used to extract CEUS modal features from the multimodal ultrasound data using the CEUS model and to extract BUS modal features from the multimodal ultrasound data using the BUS model. The dual-channel feature fusion module is used to splice the extracted CEUS modal features and BUS modal features in the channel dimension and use the SE module to perform adaptive weighting to obtain the CEUS-BUS feature map; The time intensity feature fusion module is used to divide the brightness data into time intervals to obtain the brightness change range, the brightness change range is a vector representing the degree of attention of the interval, and the CEUS-BUS feature map is compressed in the time dimension to obtain a CEUS-BUS one-dimensional vector. The brightness change range and the CEUS-BUS one-dimensional vector are multiplied to obtain the feature weight coefficient, and the feature weight coefficient is multiplied with the CEUS-BUS feature map to obtain a CEUS-BUS feature map with time attention. The fully connected layer outputs liver cancer classification prediction results based on the CEUS-BUS feature map with time attention.

7. The liver cancer detection system based on image recognition according to claim 1, characterized in that, The image recognition-based liver cancer detection system also includes an information management module, which comprises a result query submodule, a patient information management submodule, an operation history management submodule, an image information management submodule, and a user information management submodule. The result query submodule is used to respond to user-inputted patient information keyword queries and display the corresponding liver cancer analysis results for the patient. The patient information management submodule is used to display all patient information and perform add, delete, and modify operations on patient information; The operation history management submodule is used to record and query the operation history of all users; The image information management submodule is used to query the original image path, number of original images, storage path, and user notes for patients during diagnosis for all patients. The user information management submodule is used to display all user information and supports users to edit and cancel their personal accounts, as well as support administrators to add, delete, and modify all users.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the electronic device executes the computer program, it loads the liver cancer detection system based on image recognition as described in any one of claims 1 to 7.

9. A readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it loads the image recognition-based liver cancer detection system as described in any one of claims 1 to 7.