Image processing method and device and electronic equipment
By using a lesion segmentation model based on the Bottleneck structure and V-Net network, combined with multi-kernel inverted residuals and inverted residual attention modules, the accuracy issues of automatic segmentation of lung cancer lesion regions and prediction of ALK rearrangement status were resolved, realizing a non-invasive and rapid diagnostic method and improving the reliability and repeatability of diagnostic results.
Patent Information
- Application Number
- CN202511338651.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2025-12-19
AI Technical Summary
Existing technologies have low accuracy and efficiency in automatically segmenting lung cancer lesions, and cannot accurately predict ALK rearrangement status, resulting in poor reproducibility of diagnostic results and making it difficult to widely apply in clinical practice.
A lesion segmentation model based on the Bottleneck structure and V-Net network is adopted, combined with multi-kernel inverted residual and inverted residual attention modules, to automatically segment lesion regions from CT images, and to predict ALK rearrangement status using a rearrangement status prediction model.
It achieves efficient and accurate automatic segmentation of lesion areas and non-invasive and rapid prediction of ALK rearrangement status, improving the reliability and repeatability of diagnostic results and reducing reliance on human experience.
Smart Images

Figure CN121169880A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of deep learning, in particular to an image processing method and device and electronic equipment. BACKGROUND
[0002] Radiomics technology can effectively predict a variety of molecular variations of lung cancer, including Epidermal Growth Factor Receptor (EGFR) mutation and Programmed Death Ligand 1 (PD-L1) expression level. If the Anaplastic Lymphoma Kinase (ALK) rearrangement state can be predicted from the Computed Tomography (CT) image, the non-invasive screening of ALK targeted therapy patients will be realized, which has great clinical value. However, the existing researches have limitations in performance and model application, and clinical landing faces multiple challenges: on the one hand, it is time-consuming and laborious to manually outline the lesion area, the tumor boundary is irregular, and the contrast with the surrounding structure is low, which leads to large differences between measurers and poor repeatability; on the other hand, due to the problems of overlapping image features, tumor heterogeneity and the like, it is difficult to timely and effectively predict the ALK rearrangement state.
[0003] Therefore, how to improve the automatic segmentation accuracy and efficiency of the lesion area and accurately predict the ALK rearrangement state has become a technical problem to be solved. SUMMARY
[0004] Therefore, the embodiments of the present application provide an image processing method and device and electronic equipment, which can improve the automatic segmentation accuracy and efficiency of the lesion area and the accuracy of ALK rearrangement state prediction.
[0005] In a first aspect, the embodiments of the present application provide an image processing method applied to a server, the method comprising: acquiring a first to-be-processed image, wherein the first to-be-processed image comprises a lesion area; inputting the first to-be-processed image into a lesion segmentation model to acquire a second to-be-processed image comprising a lesion segmentation result, wherein the lesion segmentation model is constructed based on a Bottleneck structure and a V-Net network, and the lesion segmentation result is a segmentation result of the lesion segmentation model for the lesion area; inputting the second to-be-processed image into a rearrangement state prediction model to acquire a rearrangement prediction result for the lesion segmentation result, wherein the rearrangement state prediction model is used to predict the prediction result of the Anaplastic Lymphoma Kinase rearrangement.
[0006] In a second aspect, an embodiment of the present application provides an image processing apparatus, comprising: an acquisition module configured to acquire a first to-be-processed image, wherein the first to-be-processed image comprises a lesion region; a segmentation module configured to input the first to-be-processed image into a lesion segmentation model to acquire a second to-be-processed image comprising a lesion segmentation result, wherein the lesion segmentation model is constructed based on a Bottleneck structure and a V-Net network, and the lesion segmentation result is a segmentation result for the lesion region output by the lesion segmentation model; and a rearrangement prediction module configured to input the second to-be-processed image into a rearrangement state prediction model to acquire a rearrangement prediction result for the lesion segmentation result, wherein the rearrangement state prediction model is configured to predict a rearrangement prediction result of an anaplastic lymphoma kinase.
[0007] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor; and a memory configured to store processor-executable instructions, wherein the processor is configured to execute the image processing method of the first aspect.
[0008] In a fourth aspect, an embodiment of the present application provides a computer program product comprising a computer program, which, when executed by a processor, implements the image processing method of the first aspect.
[0009] The image processing method and apparatus, and the electronic device provided in the embodiments of the present application can realize automatic segmentation and rearrangement state prediction of the entire lesion region (for example, a tumor), can predict the possibility of the ALK gene state noninvasively, quickly and at low cost through a conventional CT image, improve the efficiency and accuracy of lesion delineation, enhance the repeatability and reliability of the diagnosis result, reduce the dependence on manual experience, and can also provide timely decision-making reference for doctors. BRIEF DESCRIPTION OF DRAWINGS
[0010] The accompanying drawings are included to provide a further understanding of the present disclosure and constitute a part of the specification, illustrate embodiments of the present disclosure and are used to explain the present disclosure together with the embodiments of the present disclosure, but do not constitute a limitation on the present disclosure. The above and other features and advantages will become more apparent from the detailed description of the specific embodiments described below, taken in conjunction with the accompanying drawings, in which: Figure 1 FIG. 1 is a schematic diagram of an implementation environment provided by an embodiment of the present application.
[0011] Figure 2is a flowchart of an image processing method provided by an example embodiment of the present application.
[0012] Figure 3 is a flowchart of an image processing method provided by another example embodiment of the present application.
[0013] Figure 4 is a flowchart of a training method of a lesion segmentation model provided by an example embodiment of the present application.
[0014] Figure 5 is a flowchart of a training method of a rearrangement state prediction model provided by an example embodiment of the present application.
[0015] Figure 6 is a structural diagram of an image processing apparatus provided by an example embodiment of the present application.
[0016] Figure 7 is a block diagram of an electronic device for image processing provided by an example embodiment of the present application. DETAILED DESCRIPTION
[0017] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0018] Lung cancer is still the leading cause of cancer-related deaths worldwide, of which non-small cell lung cancer (NSCLC) accounts for about 85% of all cases. With the discovery of targetable driver gene mutations, the treatment model of NSCLC has undergone a revolutionary change, and targeted therapy has significantly improved the quality of life and prognosis of patients compared with traditional chemotherapy. Anaplastic lymphoma kinase (ALK) gene rearrangement as an important predictor of the efficacy of ALK tyrosine kinase inhibitors (TKIs), the development of its precise non-invasive detection technology has important clinical significance.
[0019] ALK rearrangement occurs in about 3-7% of NSCLC patients, and these patients have unique clinical characteristics and are sensitive to TKI treatment. Therefore, early and accurate detection of ALK rearrangement status is crucial for developing individualized treatment plans. Traditional detection methods such as fluorescence in situ hybridization (FISH) and immunohistochemistry (IHC) require invasive tissue biopsy and have problems such as sampling error and inter-observer heterogeneity.
[0020] The emergence of radiomics technology provides a new way for non-invasive prediction of tumor molecular characteristics. CT scanning, as a routine clinical examination method, can provide rich data for radiomics analysis. Recent studies have shown that CT radiomics can effectively predict various molecular variations of lung cancer, including epidermal growth factor receptor (EGFR) mutations and programmed death ligand 1 (PD-L1) expression levels. If the ALK rearrangement status can be predicted from CT images, it will achieve non-invasive screening of ALK-targeted therapy patients, which has great clinical value.
[0021] However, existing studies have limitations in performance and application models, such as manual delineation of lesion regions, which is not only time-consuming and laborious, but also has large variability and poor repeatability due to irregular tumor boundaries, low contrast with surrounding structures, and strong dependence on experimenter experience, making it difficult to be widely promoted in clinical practice.
[0022] To solve the above problems, the image processing method provided by the embodiments of the present application is provided, and various non-limiting embodiments of the present application will be specifically introduced below with reference to the accompanying drawings.
[0023] Figure 1 is a schematic diagram of an implementation environment provided by the embodiments of the present application. The implementation environment includes a CT scanner 130, a server 120, and a computer device 110. The computer device 110 can obtain a first to-be-processed image (i.e., a medical image) from the CT scanner 130 for X-ray scanning of human tissue, and the computer device 110 can also be connected to the server 120 through a communication network. Optionally, the communication network is a wired network or a wireless network.
[0024] The computer device 110 can be a general-purpose computer or a computer device composed of a dedicated integrated circuit, and the like, and the embodiments of the present application do not limit the same. For example, the computer device 110 can be a mobile terminal device such as a tablet computer, or can also be a personal computer (PC), such as a laptop computer and a desktop computer, and the like. Those skilled in the art can know that the number of the computer device 110 can be one or more, and the types thereof can be the same or different. For example, the computer device 110 can be one, or the computer device 110 can be tens or hundreds, or more. The number and the type of the computer device 110 are not limited in the embodiments of the present application.
[0025] In some optional embodiments, the computer device 110 obtains medical sample images from the CT scanner 130, and the medical sample images can include a plurality of two-dimensional medical sample images. The computer device 110 trains the neural network through the medical sample images to obtain a network model (for example, a lesion segmentation model) for segmenting a lesion of a set of medical sample images.
[0026] The server 120 is a server, or is composed of a plurality of servers, or is a virtualization platform, or is a cloud computing service center.
[0027] Figure 2 FIG. 1 is a flowchart of an image processing method provided by an example embodiment of the present application. Figure 2 The method of FIG. 1 is performed by a computing device, for example, a server. As shown in FIG. 1, the image processing method includes the following contents. Figure 2
[0028] S210: Obtain a first to-be-processed image.
[0029] In an embodiment, the first to-be-processed image includes a lesion region.
[0030] Specifically, the first to-be-processed image can be a medical image including a lesion region, for example, a medical image or a medical image of a target organ region. The first to-be-processed image can be a CT image obtained by a CT scanner, and the first to-be-processed image can be a two-dimensional image or a three-dimensional image, and the embodiments of the present application do not limit the first to-be-processed image.
[0031] S220: Input the first to-be-processed image into a lesion segmentation model to obtain a second to-be-processed image including a lesion segmentation result.
[0032] In an embodiment, the lesion segmentation model is constructed based on a Bottleneck structure and a V-Net network, and the lesion segmentation result is a segmentation result for the lesion region output by the lesion segmentation model. That is, the lesion segmentation model is a pre-trained model constructed based on a Bottleneck structure and a V-Net network.
[0033] Specifically, the Bottleneck structure can also be referred to as a bottleneck structure, and the V-Net network can be a three-dimensional medical image segmentation network, which is not limited in the embodiments of the present application. The lesion segmentation model can be a lightweight segmentation model, referred to as UltraLight-Trans-VBNet. The lesion segmentation model can combine the Transformer network and the VBNet network to enhance the automatic identification of the lesion region (i.e., the NSCLC region) of the CT image and improve the recognition ability of the model.
[0034] For example, the lesion segmentation model can be a super-light multi-core attention network improved based on VBNet. It should be noted that VBNet is an advanced model combining V-Net and Bottleneck network structure. The Bottleneck structure reduces the parameter amount of the convolution kernel while ensuring the segmentation effect of the model, so that the model loading and inference speed is faster.
[0035] In an embodiment, the first to-be-processed image is input into the lesion segmentation model to obtain a high-resolution second to-be-processed image including the lesion segmentation result. The second to-be-processed image can be understood as the first to-be-processed image labeled with the lesion segmentation result. It should be noted that the specific description of this step is described in detail in the following embodiments. To avoid repetition, it will not be repeated here.
[0036] S230: inputting the second to-be-processed image into the rearrangement state prediction model to obtain a rearrangement prediction result for the lesion segmentation result.
[0037] In an embodiment, the rearrangement state prediction model is used to predict the prediction result of anaplastic lymphoma kinase rearrangement. The rearrangement state prediction model is a pre-trained model used to predict anaplastic lymphoma kinase rearrangement. It should be understood that the rearrangement state prediction model is used to predict the prediction result of ALK rearrangement positive.
[0038] Specifically, the second to-be-processed image including the lesion segmentation result is input into the rearrangement state prediction model, which can automatically identify and output the rearrangement prediction result for the lesion segmentation result.
[0039] Therefore, the embodiment of the present application can realize automatic segmentation and rearrangement state prediction of the entire lesion area (for example, tumor), can non-invasively, quickly and low-costly infer the possibility of ALK gene state through a conventional CT image, improves the efficiency and accuracy of lesion delineation, enhances the repeatability and reliability of the diagnosis result, reduces the dependence on artificial experience, and can also provide timely decision-making reference for doctors.
[0040] Figure 3 is a flowchart of an image processing method provided by another exemplary embodiment of the present application. Figure 3 The embodiment is Figure 2 Examples of the embodiment, which mainly describe the differences, are as follows. Figure 3 As shown in FIG. 2, step S220 further includes the following contents.
[0041] In an embodiment, the lesion segmentation model includes a multi-core reverse residual module and a multi-core reverse residual attention module.
[0042] S310: Preprocessing the first to-be-processed image to obtain a third to-be-processed image after preprocessing.
[0043] In an embodiment, the preprocessing includes resampling, normalization, contrast adjustment, and cropping and padding.
[0044] Specifically, the third to-be-processed image can be understood as the first to-be-processed image after the preprocessing operations of resampling, normalization, contrast adjustment, and cropping and padding.
[0045] Resampling can adjust the spatial resolution of the first to-be-processed image to ensure that the first to-be-processed image has the same pixel spacing. The median of the voxel spacing in the data set, that is, 1.0x1.0x1.0mm, is used as the standard spacing for resampling. For example, the first to-be-processed image and its metadata are read, and then the medical image processing library (for example, SimpleITK or PyTorch's torchio) is used to read the first to-be-processed image and its voxel spacing. Furthermore, an interpolation method (for example, linear interpolation or nearest neighbor interpolation) can be used to adjust the voxel spacing of the first to-be-processed image to the target spacing.
[0046] The standardization can reduce the influence of the value range on the model training through Z-score standardization. The specific operation may be, for example, calculating the mean and standard deviation, i.e., calculating the mean and standard deviation of the pixel values for the entire data set or each image respectively. For example, subtract the mean from each pixel value and divide by the standard deviation.
[0047] The contrast adjustment can enhance the contrast of the image through histogram equalization, gamma correction and the like, so that the target region (i.e., the lesion region) is more prominent. The specific operation may be, for example, histogram equalization, i.e., performing histogram equalization on the first to-be-processed image to make the histogram distribution of the first to-be-processed image more uniform. The specific operation may also be gamma correction, i.e., adjusting the brightness and contrast of the first to-be-processed image through gamma correction.
[0048] The cropping and padding can crop the first to-be-processed image to a fixed size and pad it to meet the requirements of the model input size (e.g., 128x128x128 voxels). The specific operation may be, for example, cropping, in which if the size of the first to-be-processed image is larger than the target size, the region of the target size is cropped from the center of the first to-be-processed image. The specific operation may also be padding, in which if the size of the first to-be-processed image is smaller than the target size, zeros are padded around the first to-be-processed image until the target size is reached.
[0049] S320: performing feature extraction on the third to-be-processed image through the multi-kernel inverse residual module to obtain a third to-be-processed image after feature extraction.
[0050] Specifically, step S320 is an operation performed in the encoder stage, and the multi-kernel inverse residual module is used to perform multi-scale feature extraction on the third to-be-processed image, and maximum pooling is used for down-sampling to obtain the third to-be-processed image after feature extraction.
[0051] It should be noted that the encoder stage of the embodiment of the present application extracts multi-scale features through the MKIR module and performs down-sampling through maximum pooling.
[0052] S330: performing image processing on the third to-be-processed image after feature extraction through the multi-kernel inverse residual attention module to obtain a second to-be-processed image.
[0053] Specifically, step S330 is an operation performed in the decoder stage, and the convolutional multi-focus attention mechanism and the multi-kernel inverse residual module are used to perform feature refinement on the third to-be-processed image after feature extraction, and up-sampling is used to restore the spatial resolution to obtain a lesion segmentation result, and then the second to-be-processed image including the lesion segmentation result is obtained.
[0054] It should be noted that the decoder stage of the embodiment of the present application refines the feature map through the MKIRA module and restores the spatial resolution through upsampling.
[0055] It should also be noted that the encoder and core decoder stages of the lesion segmentation model of the embodiment of the present application respectively introduce multi-kernel inverted residual (MKIR) and multi-kernel inverted residual attention (MKIRA) to improve feature encoding and refinement capabilities.
[0056] Therefore, by introducing the multi-kernel inverted residual module and the multi-kernel inverted residual attention module and combining image preprocessing operations, the embodiment of the present application can effectively improve the precision and robustness of lesion segmentation, while enhancing the model's ability to capture lesion features and improving the quality and reliability of the segmentation results.
[0057] In an embodiment of the present application, the multi-kernel inverted residual module is used to extract features from the third to-be-processed image to obtain a third to-be-processed image after feature extraction, including: performing multi-scale feature extraction on the third to-be-processed image through the multi-kernel inverted residual module, and performing down-sampling through maximum pooling to obtain the third to-be-processed image after feature extraction.
[0058] It should be noted that the embodiment of the present application is in the encoder stage (Encoder), and the main task of the encoder stage is to extract features from the input third to-be-processed image and gradually reduce the spatial resolution to obtain higher-level semantic information. Each encoder stage uses an MKIR module to generate a feature map. It should be noted that the addition of the MKIR module captures the context relationship through multiple kernels, further enriching the feature map.
[0059] Specifically, the purpose of the multi-kernel inverted residual module (i.e., the MKIR module) is to capture features of different scales through multi-kernel convolution to better understand granular details and context information. The multi-kernel inverted residual module uses a 1x1 convolution layer to reduce the number of channels of the input feature map (i.e., the third to-be-processed image) through a bottleneck structure to reduce the amount of calculation and reduce the computational burden for subsequent convolution operations. Then, a batch normalization (BN) operation is performed to normalize each channel of the third to-be-processed image, stabilize the training process, speed up convergence, and output the normalized third to-be-processed image.
[0060] Further, a rectified linear unit (RELU) is performed, and a ReLU activation function is applied to the normalized third to-be-processed image, so as to introduce nonlinearity to enable the network to learn complex features, and output an activated third to-be-processed image. Then, a multi-kernel depthwise convolution operation is performed, and a plurality of convolution kernels (for example, 3x3, 5x5, etc.) of different sizes are used to perform a depthwise convolution operation on the activated third to-be-processed image, so as to capture features of different scales, enhance the understanding of details and context, and output a multi-scale feature map for the third to-be-processed image.
[0061] Further, the original channel number of the third to-be-processed image is restored again using a 1x1 convolution layer, to prepare for subsequent operations, and output a third to-be-processed image with restored channel numbers. Then, a batch normalization (BN) operation is performed again, and the third to-be-processed image with restored channel numbers is normalized to stabilize the training process and accelerate convergence, thereby outputting a normalized third to-be-processed image.
[0062] Finally, a max pooling operation is performed, and a max pooling operation is performed on the normalized third to-be-processed image to reduce the spatial resolution. Through downsampling, the amount of calculation is reduced while the key information is preserved, and then a third to-be-processed image after feature extraction (i.e., a feature map after downsampling) is output.
[0063] Therefore, it can be seen that the embodiments of the present application can effectively capture multi-level features of an image, enhance the recognition ability of a model for a complex lesion region, and thus improve the accuracy and robustness of lesion segmentation, by performing multi-scale feature extraction on the image through the multi-kernel inverted residual module and combining max pooling for downsampling.
[0064] In an embodiment of the present application, the third to-be-processed image after feature extraction is processed by the multi-kernel inverted residual attention module to obtain a second to-be-processed image, including: performing feature refinement on the third to-be-processed image after feature extraction through a convolutional multi-focus attention mechanism and a multi-kernel inverted residual module, and restoring the spatial resolution through upsampling to obtain a second to-be-processed image including a lesion segmentation result.
[0065] It should be noted that the embodiments of the present application are in the decoder stage (Decoder), and the main task of the decoder stage is to gradually restore the spatial resolution of the feature map and generate a final segmentation result. Each decoder stage uses an MKIRA module to refine the feature map. The MKIRA module enhances the ability of the network to focus on key channels and spatial regions, thereby ensuring that the most prominent features are enhanced Specifically, the multi-kernel inverse residual attention module (i.e., MKIRA module) combines a convolutional multi-focal attention (CMFA) mechanism and the MKIR module to enhance the contextual information and local structure of the feature map. The third processed image after feature extraction output by the encoder is input into the MKIRA module, and through the convolutional multi-focal attention mechanism, adaptive max-pooling and average-pooling, the third processed image after feature extraction is spatially compressed to generate an attention map, thereby enhancing relevant channels and suppressing irrelevant channels, improving the robustness to local structural changes, and further outputting the third processed image after attention enhancement.
[0066] Further, similar to the MKIR module in the encoder stage, the MKIR module in the decoder stage captures features of different scales through multi-kernel convolution, refines the feature map (i.e., the third processed image after attention enhancement), enhances the understanding of details and context, and outputs the third processed image after refinement. Then, a 1x1 convolutional layer is used to reduce the number of channels of the third processed image, reduce the computational load, and prepare for subsequent operations to output the third processed image with reduced number of channels.
[0067] Further, the third processed image with reduced number of channels is normalized to stabilize the training process, speed up convergence, and output the third processed image after normalization. Then, a RELU activation operation is performed to apply a ReLU activation function to the third processed image after normalization, introduce nonlinearity, enable the network to learn complex features, and output the third processed image after activation. Finally, an upsampling operation is performed on the third processed image after activation to restore the spatial resolution, gradually restore the spatial resolution of the feature map, and output the third processed image after upsampling. Finally, in the last stage of the decoder, a high-resolution lesion segmentation result is output, thereby obtaining the second processed image including the lesion segmentation result for medical image analysis and completing the task of segmenting the lesion region.
[0068] It should be noted that the lesion segmentation model can maintain high precision while minimizing computational overhead. The connection between the encoder and the decoder is through grouped attention, which guides the information flow by utilizing the gating signal from higher resolution features, thereby more lightly improving the segmentation performance, making it easy to deploy the segmentation network to the cloud or mobile applications. While ensuring the segmentation effect of the model, the parameter amount of the convolution kernel is reduced, making the model loading and inference faster.
[0069] Therefore, the embodiment of the application refines the features of the image after feature extraction by the convolution multi-focus attention mechanism and the multi-core reverse residual module, and restores the spatial resolution by upsampling, thereby effectively improving the accuracy and detail performance of the lesion segmentation result and enhancing the recognition and positioning ability of the model for the lesion area.
[0070] Figure 4 is a flowchart of a training method of a lesion segmentation model provided by an exemplary embodiment of the application. As shown in Figure 4 , the training method includes the following contents.
[0071] S410: Obtain sample image data.
[0072] In an embodiment, the sample image data includes image data obtained by a data augmentation method.
[0073] Specifically, the sample image data can include original image data and image data obtained by a data augmentation method, wherein the data augmentation method can be to generate more training samples and improve the robustness of the model by rotating, flipping, scaling, etc. the original image data. For example, a random rotation method is used to randomly rotate the image within a certain range. And / or, a random flipping method is used to randomly select horizontal flipping or vertical flipping. And / or, a random scaling method is used to randomly scale the image within a certain range. The data augmentation method can be flexibly set according to the actual situation. The application does not make specific limitations on the data included in the sample image data.
[0074] S420: Input the sample image data into the initial lesion segmentation model to obtain a predicted segmentation result.
[0075] Specifically, the sample image data is used to train the initial lesion segmentation model (for example, the initial UltraLight-Trans-VBNet) model, wherein the preset result (for example, the gold standard) referred to by the model is the NSCLC region labeled by a physician, and the back propagation algorithm and the stochastic gradient descent algorithm are used to obtain the predicted segmentation result according to the network forward propagation.
[0076] It should be noted that the initial lesion segmentation model can be understood as a model that has not been trained.
[0077] S430: Obtain a similarity coefficient between the predicted segmentation result and the preset result.
[0078] Specifically, the similarity coefficient can be a Dice (Dice Coefficient) coefficient. The Dice coefficient can be used as a loss function, and the smaller the loss function, the better the convergence of the trained initial lesion segmentation model. The score of the Dice coefficient is between 0 and 1, and the closer the score is to 1, the higher the overlap between the predicted segmentation result and the preset result of the standard annotation, and the stronger the similarity. Generally speaking, a Dice coefficient greater than 0.5 means that the predicted segmentation result has a high degree of overlap, thereby reflecting that the trained initial lesion segmentation model has good segmentation performance.
[0079] In an embodiment, the Dice coefficient between the predicted segmentation result output by the initial lesion segmentation model and the preset result (e.g., the gold standard) is calculated.
[0080] S440: In the case where the similarity coefficient meets the preset condition, the lesion segmentation model is obtained.
[0081] Specifically, the similarity coefficient (i.e., the Dice coefficient) can be used as a loss function, and when the loss value is low to a certain extent and cannot be further reduced with the number of iterations, the training of the initial lesion segmentation model is completed and stopped, and the final segmentation model, i.e., the lesion segmentation model, is obtained.
[0082] It should be noted that the trained lesion segmentation model can also be applied to the validation set for further parameter optimization. The best model parameters and hyperparameter configurations are selected by the Dice coefficient. The trained UltraLight-Trans-VBNet model is evaluated on the test set, and the lesion region is visualized. The Dice coefficient is used to measure the segmentation accuracy, and the overlap between the output result (i.e., the predicted segmentation result) of the model and the gold standard (i.e., the preset result) of the physician is calculated to evaluate the accuracy and consistency of the segmentation result.
[0083] Therefore, the embodiments of the present application enrich sample image data through data augmentation, evaluate the accuracy of the predicted segmentation result using the similarity coefficient, and obtain the lesion segmentation model when the similarity coefficient meets the preset condition, thereby improving the robustness and segmentation accuracy of the model and enhancing the accuracy and reliability of the lesion segmentation.
[0084] Figure 5 is a flowchart of a training method of a rearrangement state prediction model provided by an exemplary embodiment of the present application. As shown in Figure 5 , the training method includes the following contents.
[0085] S510: Obtain sample training data.
[0086] In an embodiment, before extracting the radiomics features, data preprocessing is needed on the sample lesion segmentation result (e.g., an image of a target organ region, i.e., an original sample image including the lesion segmentation result) to improve the quality of the data, reduce noise and inconsistency, and thus improve the accuracy and generalization ability of the model.
[0087] Specifically, the preprocessing can include image resampling, which adjusts the spatial resolution of the image to ensure that all images have the same pixel spacing. The median of the voxel spacing in the statistical dataset, i.e., 1.0x1.0x1.0 millimeter, is used as the standard spacing for resampling. The preprocessing can also include image standardization, which normalizes the image by window width and window level, and concatenates the maximum and minimum normalization to make the data from different devices or sources consistent and reduce the bias caused by device differences.
[0088] Next, radiomics features are extracted for the lesion region. The automatically extracted Region Of Interest (ROI) radiomics features that meet the IBSI standard have 2264, and the radiomics parameters (e.g., shape features) can be extracted from the sample lesion segmentation result (i.e., the original sample image) according to the ROI, and the radiomics parameters (e.g., texture features, gray scale statistical quantity features, and high-order features) are extracted from the original sample lesion segmentation result and the filtered image.
[0089] In an embodiment, the radiomics parameters are extracted from the sample lesion segmentation result; based on the radiomics parameters and the clinical information, sample training data is obtained, wherein the sample training data includes a sample training dataset, a sample validation dataset, and a sample test dataset.
[0090] Specifically, the radiomics parameters are extracted from the sample lesion segmentation result, and the radiomics parameters are combined with the clinical information to obtain the sample training data. The sample training dataset, the sample validation dataset, and the sample test dataset can be divided in a certain ratio, for example, randomly divided into the sample training dataset, the sample test dataset, and the sample validation dataset in a ratio of 8:1:1.
[0091] Therefore, by extracting the radiomics parameters from the sample lesion segmentation result and combining the clinical information to construct the sample training data, the embodiments of the present application can fully utilize the imaging and clinical information, improve the training effect and generalization ability of the model, and thus improve the accuracy and reliability of the diagnosis.
[0092] In an embodiment, all the extracted radiomics parameters and clinical information features of the sample training data set are normalized by z-score, so as to eliminate the influence of inconsistent dimensions of feature values. Then, the first K (K is 1 / 2 of the number of original features) significant features are selected by chi-square test, and finally the optimal features are selected by least absolute shrinkage and selection operator (LASSO). Then, step S520 is performed to construct the ALK rearrangement state prediction model by using the selected features and machine learning algorithms.
[0093] S520: training the initial rearrangement state prediction model by using multiple machine learning algorithms and sample training data to obtain multiple ALK rearrangement state prediction results.
[0094] In an embodiment, the multiple ALK rearrangement state prediction results are in the form of score values.
[0095] Specifically, the machine learning algorithm can be, for example, a random forest, a decision tree, a support vector machine, a logistic regression, etc. The initial rearrangement state prediction model is trained by using multiple machine learning algorithms and sample training data, and after the training of the initial rearrangement state prediction model is completed, each model is applied to the sample test data set and the sample validation data set. It should be noted that the initial rearrangement state prediction model can be understood as a model that has not been trained. And each machine learning algorithm is trained with the initial rearrangement state prediction model to obtain multiple ALK rearrangement state prediction results, wherein the machine learning algorithm and the ALK rearrangement state prediction result correspond one-to-one.
[0096] In an embodiment, the ALK rearrangement state prediction result can be a result of the comprehensive performance of the model. The ALK rearrangement state prediction result can be quantified by a receiver operating characteristic (ROC) curve, and the value range of the ROC curve and the area under the curve (AUC) is [0.5, 1], wherein the closer the AUC is to 1, the better the performance of the model. That is, the ALK rearrangement state prediction result is in the form of a score value, and the closer the ALK rearrangement state prediction result is to 1, the better the performance of the model corresponding to the ALK rearrangement state prediction result.
[0097] It should be noted that the performance of the model can also be evaluated in detail by sensitivity, specificity, accuracy, precision, F1 score, etc., which are not limited in the embodiments of the present application.
[0098] S530: taking the initial rearrangement state prediction model corresponding to the ALK rearrangement state prediction result with the highest score value as the rearrangement state prediction model.
[0099] Specifically, after training the initial rearrangement state prediction model by applying multiple machine learning algorithms, the model with the best performance (i.e., the highest score value) is finally selected as the rearrangement state prediction model, i.e., the ALK prediction model.
[0100] It should be noted that there are currently studies exploring the relationship between CT radiomics features and ALK rearrangement status. For example, a retrospective study on lung adenocarcinoma patients found that a specific combination of radiomics features can accurately predict ALK fusion gene expression. Another study integrated radiomics features and traditional clinical CT features to construct an ALK rearrangement prediction model with high prediction performance.
[0101] Therefore, the embodiments of the present application train the sample training data by machine learning algorithms, select the model corresponding to the ALK rearrangement state prediction result with the highest score value as the final rearrangement state prediction model, thereby ensuring the accuracy and reliability of the model and improving the performance of rearrangement state prediction and the clinical application value.
[0102] In an embodiment of the present application, the radiomics parameters include shape features, texture features, gray level statistics features, and high-order features.
[0103] Specifically, the radiomics parameters can include shape features, texture features, gray level statistics features, and high-order features, etc. The shape features can include 14 three-dimensional features reflecting the shape and size of the region, such as maximum diameter, volume, surface area, similarity to a sphere, etc. The gray level statistics features can be obtained from 18 features extracted from the original sample image, which quantitatively describe the distribution of voxel intensity in the image through commonly used and basic metrics, including mean, peak, maximum, minimum, etc. The texture features can be features reflecting the gray level distribution of pixels and their surrounding spatial neighborhoods, obtained from 72 features extracted from the original sample image, and calculated according to the Gray Level Cooccurence Matrix (GLCM), Gray Level Run Length Matrix (GLRLM), Gray Level Size Zone Matrix (GLSZM), Neighbouring Gray Tone Difference Matrix (NGTDM), and Gray Level Dependence Matrix (GLDM). The high-order features can be obtained from 432 gray level statistics features and 1728 texture features extracted from the filtered image, and the filtering methods include mean filtering, Gaussian filtering, logarithmic filtering, wavelet transform, etc. 24 kinds.
[0104] Therefore, by comprehensively applying various radiomics parameters, the embodiments of the present application can more comprehensively and accurately quantify the features of the lesion region, thereby improving the accuracy and repeatability of the diagnosis, reducing the dependence on human experience, and improving the efficiency and reliability of clinical application.
[0105] Figure 6 is a structural schematic diagram of an image processing device provided by an exemplary embodiment of the present application. As shown in Figure 6 the image processing device 600 includes an acquisition module 610, a segmentation module 620, a rearrangement prediction module 630, and a training module 640.
[0106] The acquisition module 610 is configured to acquire a first to-be-processed image, where the first to-be-processed image includes a lesion region; the segmentation module 620 is configured to input the first to-be-processed image into a lesion segmentation model to acquire a second to-be-processed image including a lesion segmentation result, where the lesion segmentation model is constructed based on a Bottleneck structure and a V-Net network, and the lesion segmentation result is a segmentation result for the lesion region output by the lesion segmentation model; and the rearrangement prediction module 630 is configured to input the second to-be-processed image into a rearrangement state prediction model to acquire a rearrangement prediction result for the lesion segmentation result, where the rearrangement state prediction model is used to predict a prediction result of anaplastic lymphoma kinase rearrangement.
[0107] The image processing apparatus provided in the embodiments of the present application can realize automatic segmentation of the entire lesion region (for example, a tumor) and prediction of a rearrangement state, can predict the possibility of an ALK gene state noninvasively, quickly and at low cost through a conventional CT image, can improve the efficiency and accuracy of lesion delineation, can enhance the repeatability and reliability of a diagnosis result, can reduce the dependence on manual experience, and can provide a doctor with timely decision-making reference.
[0108] According to an embodiment of the present application, the lesion segmentation model includes a multi-core reverse residual module and a multi-core reverse residual attention module, the segmentation module 620 is configured to pre-process the first to-be-processed image to obtain a third to-be-processed image after pre-processing, where the pre-processing includes resampling, standardization, contrast adjustment, and cropping and padding; the multi-core reverse residual module is configured to perform feature extraction on the third to-be-processed image to obtain a third to-be-processed image after feature extraction; and the multi-core reverse residual attention module is configured to perform image processing on the third to-be-processed image after feature extraction to obtain the second to-be-processed image.
[0109] According to an embodiment of the present application, the segmentation module 620 is configured to perform multi-scale feature extraction on the third to-be-processed image through the multi-core reverse residual module, and perform down-sampling through maximum pooling to obtain the third to-be-processed image after feature extraction.
[0110] According to an embodiment of the present application, the segmentation module 620 is configured to perform feature refinement on the third to-be-processed image after feature extraction through a convolution multi-focus attention mechanism and the multi-core reverse residual module, and restore the spatial resolution through up-sampling to obtain the second to-be-processed image including the lesion segmentation result.
[0111] According to an embodiment of the present application, the training module 640 is configured to obtain sample image data, wherein the sample image data comprises image data obtained through data enhancement; input the sample image data into the initial lesion segmentation model to perform model training, and obtain a predicted segmentation result; obtain a similarity coefficient between the predicted segmentation result and a preset result; and in a case where the similarity coefficient meets a preset condition, obtain the lesion segmentation model.
[0112] According to an embodiment of the present application, the training module 640 is configured to obtain sample training data; train the initial rearrangement state prediction model through a plurality of machine learning algorithms and the sample training data, and obtain a plurality of ALK rearrangement state prediction results, wherein the plurality of ALK rearrangement state prediction results are in the form of score values; and take the initial rearrangement state prediction model corresponding to the ALK rearrangement state prediction result with the highest score value as the rearrangement state prediction model.
[0113] According to an embodiment of the present application, the training module 640 is configured to extract radiomics parameters from sample lesion segmentation results; and obtain sample training data based on the radiomics parameters and clinical information, wherein the sample training data comprises a sample training data set, a sample verification data set and a sample test data set.
[0114] According to an embodiment of the present application, the radiomics parameters comprise shape features, texture features, gray scale statistical quantity features and high-order features.
[0115] It should be understood that the specific working processes and functions of the obtaining module 610, the segmentation module 620, the rearrangement prediction module 630 and the training module 640 in the above embodiments can refer to the descriptions of the image processing method provided by the embodiments, which will not be repeated here. Figures 2 to 5 The embodiments provide an image processing method.
[0116] Figure 7 is a block diagram of an electronic device 700 for image processing provided by an exemplary embodiment of the present application.
[0117] Referring to Figure 7 , the electronic device 700 comprises a processing assembly 710, which further comprises one or more processors, and a memory resource represented by a memory 720, for storing instructions executable by the processing assembly 710, such as an application program. The application program stored in the memory 720 can comprise one or more than one module each corresponding to a set of instructions. In addition, the processing assembly 710 is configured to execute the instructions to perform the image processing method described above.
[0118] The electronic device 700 can further include a power supply component configured to perform power management of the electronic device 700, a wired or wireless network interface configured to connect the electronic device 700 to a network, and an input / output (I / O) interface. The electronic device 700 can be operated based on an operating system stored in the memory 720, such as Windows Server TM , Mac OSX TM , Unix TM , Linux TM , FreeBSD TM or the like.
[0119] A non-transitory computer-readable storage medium, when instructions stored in the storage medium are executed by a processor of the electronic device 700 described above, enable the electronic device 700 to perform an image processing method, comprising: obtaining a first to-be-processed image, wherein the first to-be-processed image includes a lesion region; inputting the first to-be-processed image into a lesion segmentation model to obtain a second to-be-processed image including a lesion segmentation result, wherein the lesion segmentation model is constructed based on a Bottleneck structure and a V-Net network, and the lesion segmentation result is a segmentation result for the lesion region output by the lesion segmentation model; inputting the second to-be-processed image into a rearrangement state prediction model to obtain a rearrangement prediction result for the lesion segmentation result, wherein the rearrangement state prediction model is used to predict a prediction result of anaplastic lymphoma kinase rearrangement.
[0120] Embodiments of the present application also provide a computer program product, comprising a computer program for executing the steps of the method of turning on the turn signal lamp described in the method embodiments, which can be specifically referred to the method embodiments described above, and will not be repeated here. The computer program product can be specifically realized by hardware, software or a combination thereof. In one optional embodiment, the computer program product is specifically embodied as a computer storage medium, and in another optional embodiment, the computer program product is specifically embodied as a software product, such as a software development kit (Software Development Kit, SDK) and the like.
[0121] All the optional technical solutions described above can be combined to form optional embodiments of the present application, and will not be repeated here.
[0122] Those skilled in the art can appreciate that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solutions. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0123] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the system, device and unit described above can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.
[0124] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.
[0125] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0126] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit.
[0127] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the essential part or part of the technical solutions that make contributions to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various storage program code storage media.
[0128] It should be noted that in the description of the present application, the terms "first", "second", "third" and the like are used only for descriptive purposes and are not to be construed as indicating or implying relative importance. In addition, in the description of the present application, the meaning of "a plurality of" is two or more, unless otherwise specified.
[0129] The above only the preferred embodiment of the present application has, and does not limit the present application, any modification, equivalent replacement, etc. made within the spirit and principle of the present application, should be included in the protection scope of the present application.
Claims
1. An image processing method, characterized by, Applied to a server, comprising: obtaining a first to-be-processed image, wherein the first to-be-processed image comprises a lesion region; inputting the first to-be-processed image into a lesion segmentation model to obtain a second to-be-processed image comprising a lesion segmentation result, wherein the lesion segmentation model is constructed based on a bottleneck structure and a V-Net network, and the lesion segmentation result is a segmentation result for the lesion region output by the lesion segmentation model; inputting the second to-be-processed image into a rearrangement state prediction model to obtain a rearrangement prediction result for the lesion segmentation result, wherein the rearrangement state prediction model is used to predict a prediction result of anaplastic lymphoma kinase rearrangement.
2. The image processing method of claim 1, wherein, The lesion segmentation model comprises a multi-core reverse residual module and a multi-core reverse residual attention module, wherein the inputting the first to-be-processed image into the lesion segmentation model to obtain the second to-be-processed image comprising the lesion segmentation result comprises: preprocessing the first to-be-processed image to obtain a third to-be-processed image after preprocessing, wherein the preprocessing comprises resampling, standardization, contrast adjustment, and cropping and padding; performing feature extraction on the third to-be-processed image through the multi-core reverse residual module to obtain a third to-be-processed image after feature extraction; performing image processing on the third to-be-processed image after feature extraction through the multi-core reverse residual attention module to obtain the second to-be-processed image.
3. The image processing method of claim 2, wherein, The performing feature extraction on the third to-be-processed image through the multi-core reverse residual module to obtain the third to-be-processed image after feature extraction comprises: performing multi-scale feature extraction on the third to-be-processed image through the multi-core reverse residual module, and performing down-sampling through maximum pooling to obtain the third to-be-processed image after feature extraction.
4. The image processing method of claim 2, wherein, The performing image processing on the third to-be-processed image after feature extraction through the multi-core reverse residual attention module to obtain the second to-be-processed image comprises: performing feature refinement on the third to-be-processed image after feature extraction through a convolutional multi-focus attention mechanism and a multi-core reverse residual module, and restoring the spatial resolution through up-sampling to obtain the second to-be-processed image comprising the lesion segmentation result.
5. The image processing method of claim 1, wherein, The training process of the lesion segmentation model comprises: obtaining sample image data, wherein the sample image data comprises image data obtained through data enhancement; inputting the sample image data into an initial lesion segmentation model for model training to obtain a predicted segmentation result; obtaining a similarity coefficient between the predicted segmentation result and a preset result; in a case where the similarity coefficient meets a preset condition, obtaining the lesion segmentation model.
6. The image processing method of claim 1, wherein, The training process of the rearrangement state prediction model comprises: obtaining sample training data; training an initial rearrangement state prediction model through a plurality of machine learning algorithms and the sample training data to obtain a plurality of ALK rearrangement state prediction results, wherein the plurality of ALK rearrangement state prediction results are in the form of score values; taking an initial rearrangement state prediction model corresponding to an ALK rearrangement state prediction result with the highest score value as the rearrangement state prediction model.
7. The image processing method of claim 6, wherein, The sample training data is obtained, including: extracting radiomics parameters from the sample lesion segmentation result; obtaining the sample training data based on the radiomics parameters and the clinical information, wherein the sample training data includes a sample training data set, a sample validation data set and a sample test data set.
8. The image processing method of claim 7, wherein, The radiomics parameters include shape features, texture features, gray scale statistical quantity features and high-order features.
9. An image processing apparatus characterized by comprising: It comprises: an acquisition module configured to acquire a first to-be-processed image, wherein the first to-be-processed image includes a lesion region; a segmentation module configured to input the first to-be-processed image into a lesion segmentation model to obtain a second to-be-processed image including a lesion segmentation result, wherein the lesion segmentation model is constructed based on a Bottleneck structure and a V-Net network, and the lesion segmentation result is a segmentation result for the lesion region output by the lesion segmentation model; a rearrangement prediction module configured to input the second to-be-processed image into a rearrangement state prediction model to obtain a rearrangement prediction result for the lesion segmentation result, wherein the rearrangement state prediction model is used to predict a prediction result of anaplastic lymphoma kinase rearrangement.
10. An electronic device, comprising: It comprises: a processor; a memory for storing instructions executable by the processor, wherein the processor is configured to execute the image processing method of any one of claims 1 to 8.
11. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the image processing method of any one of claims 1 to 8.