Esophageal cancer PD-L1 expression state prediction method, device, equipment, medium and product
By training deep learning models based on multimodal data based on 18F-FDG PET/CT imaging and clinical pathological characteristics, the invasiveness and time-consuming problems of PD-L1 expression status detection in esophageal carcinoma are solved, and non-invasive, fast and efficient PD-L1 prediction is achieved.
Patent Information
- Application Number
- CN202510311857.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-29
AI Technical Summary
The existing PD-L1 expression status detection method for esophageal carcinoma requires invasive biopsy, which is long and local, and has low accuracy.
By acquiring multimodal data, including 18F-FDG PET/CT images and clinicopathological features, deep learning models are trained to predict the expression status of PD-L1 in the tumor region of interest, and avoid invasive manipulation.
Achieve non-invasive, fast and efficient PD-L1 prediction, providing systemic immune status assessment, reducing costs and improving detection efficiency.
Smart Images

Figure CN120388726A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing, and in particular, to a method, device, equipment, medium and product for predicting the PD-L1 expression status of esophageal cancer. Background Art
[0002] With the rapid development of the field of tumor immunotherapy, immune checkpoint inhibitors (ICIs) centered on programmed death receptor-1 (PD-1) and programmed death ligand-1 (PD-L1) have made remarkable progress in cancer treatment. How to detect the status of PD-L1 has become an urgent problem to be solved.
[0003] In the related art, the expression detection of PD-L1 in esophageal cancer generally uses immunohistochemistry (IHC) method. However, using the above method not only requires obtaining tumor specimens through invasive biopsy, but also needs to send them to the laboratory for testing, which often takes a long time. In addition, the biopsy results are often local and the accuracy is not high. Summary of the Invention
[0004] The present disclosure provides a method, device, equipment, medium and product for predicting the PD-L1 expression status of esophageal cancer, which is used to solve the technical problems that the existing methods for detecting the PD-L1 expression status of esophageal cancer require invasive biopsy, take a long time, and are local.
[0005] The first aspect of the present disclosure is to provide a data processing method, including:
[0006] Obtain a training data set, where the training data set includes a plurality of training data pairs, the training data pairs include the associated data of esophageal cancer patients and label information, the associated data includes the first tomographic image and the second tomographic image of the tumor region of interest and the clinicopathological feature vector, the first tomographic image and the second tomographic image are associated with the same image parameters, and the image parameters include voxel size, window level and width, image size, and numerical scale, and the label information is used to indicate the true expression status of programmed death ligand 1 in the tumor region of interest;
[0007] Based on the training data set, perform iterative training on a preset network model until the network model meets the preset convergence condition to obtain a target prediction model, and the target prediction model is used to predict the predicted expression status of programmed death ligand 1 in the tumor region of interest.
[0008] The second aspect of the present disclosure is to provide a data processing method, including:
[0009] Obtain the data to be recognized associated with the target patient, where the data to be recognized includes a first image, a second image, and target clinicopathological data. The first image and the second image are 18F-FDG PET / CT images associated with the target patient, and the target patient is a patient with esophageal cancer;
[0010] Perform a preprocessing operation on the data to be recognized to obtain target data, where the image parameters associated with the first image and the second image in the target data are consistent. The image parameters include voxel size, window level and width, image size, and numerical scale;
[0011] Input the target data into a preset target prediction model to obtain a prediction result output by the target prediction model. The prediction result is the predicted expression status of programmed death ligand 1 in the tumor region of interest of the target patient;
[0012] The target prediction model is trained based on the data processing method described in the first aspect.
[0013] The third aspect of the present disclosure is to provide a data processing device, including:
[0014] An acquisition module, configured to acquire a training data set, where the training data set includes a plurality of training data pairs. The training data pair includes the associated data of a patient with esophageal cancer and label information. The associated data includes a first tomographic image, a second tomographic image, and a clinicopathological feature vector of the tumor region of interest. The image parameters associated with the first tomographic image and the second tomographic image are consistent. The image parameters include voxel size, window level and width, image size, and numerical scale. The label information is used to indicate the true expression status of programmed death ligand 1 in the tumor region of interest;
[0015] A training module, configured to iteratively train a preset network model based on the training data set until the network model meets a preset convergence condition, to obtain a target prediction model. The target prediction model is used to predict the predicted expression status of programmed death ligand 1 in the tumor region of interest.
[0016] The fourth aspect of the present disclosure is to provide a data processing device, including:
[0017] A data acquisition module, configured to acquire the data to be recognized associated with the target patient, where the data to be recognized includes a first image, a second image, and target clinicopathological data. The first image and the second image are 18F-FDG PET / CT images associated with the target patient, and the target patient is a patient with esophageal cancer;
[0018] A preprocessing module for preprocessing the data to be recognized to obtain target data, where the image parameters associated with the first image and the second image in the target data are consistent, and the image parameters include voxel size, window level and width, image size, and numerical scale;
[0019] A prediction module for inputting the target data into a preset target prediction model to obtain a prediction result output by the target prediction model, where the prediction result is the predicted expression status of programmed death ligand 1 in the tumor region of interest of the target patient;
[0020] The target prediction model is trained based on the data processing device described in the third aspect.
[0021] The fifth aspect of the present disclosure is to provide an electronic device, including: a processor and a memory;
[0022] The memory stores computer execution instructions;
[0023] The processor executes the computer execution instructions stored in the memory, so that the processor executes the data processing method described in the first aspect or the second aspect.
[0024] The sixth aspect of the present disclosure is to provide a computer-readable storage medium, in which computer execution instructions are stored, and when the processor executes the computer execution instructions, the data processing method described in the first aspect or the second aspect is implemented.
[0025] The seventh aspect of the present disclosure is to provide a computer program product, including a computer program, and when the computer program is executed by the processor, the data processing method described in the first aspect or the second aspect is implemented.
[0026] The esophageal cancer PD-L1 expression status prediction method, device, equipment, medium and product provided by the present disclosure pre-obtain a training data set, which may include multiple training data pairs. Each training data pair includes the associated data of an esophageal cancer patient and label information, where the label information is used to indicate the true expression status of programmed death ligand 1 in the tumor region of interest of the esophageal cancer patient. Then, based on the training data set, a preset network model is trained. After the network model meets the preset convergence condition, a target prediction model can be obtained. Furthermore, after obtaining the to-be-identified data associated with the target patient, the target prediction model can accurately predict the predicted expression status of programmed death ligand 1 in the tumor region of interest of the target patient. In addition, since the associated data includes multi-modal data such as the first tomographic imaging, the second tomographic imaging, and the clinicopathological feature vector of the tumor region of interest, the trained target prediction model can combine multi-modal data such as the 18F-FDG PET / CT image and the clinicopathological data of the target patient to accurately predict the predicted expression status of programmed death ligand 1 in the tumor region of interest, and there is no need to perform invasive biopsy on the target patient, effectively avoiding the invasive risk. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can also be obtained according to these drawings.
[0028] Figure 1 Schematic diagram of the system architecture based on the present disclosure;
[0029] Figure 2 Schematic flowchart of the data processing method provided by the embodiment of the present disclosure;
[0030] Figure 3 Schematic flowchart of the data processing method provided by another embodiment of the present disclosure;
[0031] Figure 4 Schematic flowchart of the data processing method provided by another embodiment of the present disclosure;
[0032] Figure 5 Schematic flowchart of the data processing method provided by another embodiment of the present disclosure;
[0033] Figure 6 Schematic diagram of the structure of the network model provided by the embodiment of the present disclosure;
[0034] Figure 7 Schematic diagram of the ROC curve provided by the embodiment of the present disclosure;
[0035] Figure 8 Schematic flowchart of the data processing method provided by an embodiment of the present disclosure;
[0036] Figure 9 Schematic structural diagram of the data processing device provided by an embodiment of the present disclosure;
[0037] Figure 10 Schematic structural diagram of the data processing device provided by an embodiment of the present disclosure;
[0038] Figure 11 Schematic structural diagram of the electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some but not all of the embodiments of the present disclosure. All other embodiments obtained based on the embodiments in the present disclosure fall within the scope of protection of the present disclosure.
[0040] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0041] Glossary:
[0042] Programmed death ligand-1 (PD-L1 for short), a transmembrane protein belonging to the B7 family, is mainly expressed on the surfaces of tumor cells and certain immune cells.
[0043] Programmed death-1 (PD-1 for short) is mainly expressed on the surface of T cells. After binding to PD-L1, it inhibits the activity of T cells. Tumor cells inhibit the activation and proliferation of T cells by expressing PD-L1 and binding it to PD-1 on the surface of T cells.
[0044] 18F-FDG PET / CT images are imaging techniques that combine positron emission tomography (PET) and computed tomography (CT). Using 18F-FDG as a tracer, they provide both functional metabolic information and anatomical structure information.
[0045] The area under the curve (AUC) index, which refers to the area under the receiver operating characteristic (ROC) curve, comprehensively considers both false positives and true positives. Compared to individual sensitivity or specificity, it can evaluate model performance more comprehensively. The ROC curve shows the performance of a model by plotting the relationship between the false positive rate (FPR) and the true positive rate (TPR, i.e., sensitivity), reflecting the model's performance at different thresholds. The AUC value is the area under the ROC curve, indicating the model's ability to distinguish positive and negative samples. The value of AUC ranges from 0 to 1, and the closer it is to 1, the better the model performance.
[0046] In related technologies, immunohistochemistry (IHC) methods are mainly relied on for detecting PD-L1 expression in esophageal cancer. This not only requires obtaining tumor specimens through invasive biopsies but also has the following disadvantages: high cost because surgeries or biopsies involve equipment and labor costs; low efficiency as patients need to wait for laboratory test results, which takes a long time; invasive risks as biopsies themselves may cause complications, increasing the patient's risk; and the problem of locality as biopsy results may not fully reflect the overall immune status of the tumor.
[0047] In the process of solving the above problems, the inventors found through research that a deep learning model based on 18F-FDG PET / CT images and clinicopathological information can be developed to provide a non-invasive, fast, and efficient PD-L1 prediction method. This method can avoid invasive operations, reduce costs, improve detection efficiency, and provide a more comprehensive assessment of the systemic immune status, facilitating precise treatment decisions.
[0048] To solve the technical problems that the existing methods for detecting the PD-L1 expression status in esophageal cancer require invasive biopsies, are time-consuming, and have locality issues, the present disclosure provides a method, device, equipment, medium, and product for predicting the PD-L1 expression status in esophageal cancer.
[0049] It should be noted that the method, device, equipment, medium, and product for predicting the PD-L1 expression status in esophageal cancer provided by the present disclosure can be applied to any PD-L1 prediction scenario.
[0050] Figure 1 Schematic diagram of the system architecture on which the present disclosure is based. As Figure 1 shown, the system architecture on which the present disclosure is based at least includes: a server 11 and a data server 12. Among them, a data processing device is provided in the server 11. The data processing device can be written in languages such as C / C++, Java, Shell, or Python; the data server 12 can be a cloud server or a server cluster, and a large amount of training data is stored therein.
[0051] Based on the above system architecture, the server 11 can obtain a training data set from the data server 12, and perform iterative training operations on a preset network model based on the training data set. Therefore, the obtained target prediction model can have the ability to predict the PD-L1 expression status of the tumor region of interest of the target patient based on tomographic image data and clinical pathological data.
[0052] Figure 2 Schematic diagram of the flow of the data processing method provided by the embodiment of the present disclosure. As Figure 2 shown, the method includes:
[0053] Step 201, obtain a training data set, where the training data set includes a plurality of training data pairs, the training data pair includes the associated data of an esophageal cancer patient and label information, the associated data includes the first tomographic imaging, the second tomographic imaging, and the clinical pathological feature vector of the tumor region of interest, the image parameters associated with the first tomographic imaging and the second tomographic imaging are the same, the image parameters include voxel size, window level and width, image size, and numerical scale, and the label information is used to indicate the true expression status of programmed death ligand 1 of the tumor region of interest.
[0054] The execution subject of this embodiment is a data processing device. The data processing device can be coupled to the server.
[0055] In this embodiment, in order to be able to determine the PD-L1 expression status of the tumor region of interest of an esophageal cancer patient without invasive biopsy, a network model can be pre-trained to perform a PD-L1 expression status prediction operation by identifying image data through the network model.
[0056] In the related art, generally only data of one data modality is used for model training. For example, only single imaging data or clinicopathological data is used for model training. However, due to the relatively single training data, it often limits the model's comprehensive understanding and prediction of complex disease states. For example, in the prediction of the PD-L1 expression status of esophageal cancer, single imaging data may not be sufficient to comprehensively capture the biological characteristics and immune microenvironment of the tumor, and pure clinicopathological data cannot fully reflect the imaging characteristics and systemic immune status of the tumor.
[0057] Therefore, in order to enable the model to comprehensively and accurately perform the prediction operation of the PD-L1 expression status, multi-modal data can be used for training the network model.
[0058] Optionally, a training data set can be obtained. Among them, the training data set includes multiple training data pairs, and each training data pair includes the associated data of esophageal cancer patients and label information. Therefore, the target prediction model trained based on this training data set can accurately predict the PD-L1 expression status of esophageal cancer patients.
[0059] Among them, the associated data includes the first tomographic imaging and the second tomographic imaging of the tumor region of interest and the clinicopathological feature vector. The image parameters associated with the first tomographic imaging and the second tomographic imaging are the same, and the image parameters include voxel size, window level and width, image size, and numerical scale. The first tomographic imaging and the second tomographic imaging are 18F-FDG PET / CT images of the tumor region of interest.
[0060] The label information is used to indicate the expression status of programmed death ligand 1 in the tumor region of interest. Among them, the label information can be determined based on the surgical cases of esophageal cancer patients, and the label information can be a positive label or a negative label.
[0061] Step 202: Iteratively train a preset network model based on the training data set until the network model meets the preset convergence condition, and obtain a target prediction model, where the target prediction model is used to predict the predicted expression status of programmed death ligand 1 in the tumor region of interest.
[0062] In this embodiment, a network model can be preset. The network model can include a first feature extractor, a second feature extractor, and a feature classifier. The first feature extractor and the second feature extractor are used to extract features from the imaging data. Among them, the first feature extractor and the second feature extractor can be of the ResNet10 architecture, which consists of a 10-layer deep convolutional neural network.
[0063] Since the training data is multi-modal data, in order to implement the fusion operation of multi-modal data, the feature vectors output by the first feature extractor and the second feature extractor can be concatenated with the clinicopathological feature vector, and the concatenated vector is input into the feature classifier. The feature classifier is used to process the input vector to predict the status of programmed death ligand 1 in the tumor region of interest of the esophageal cancer patient. The feature classifier consists of two fully connected layers, and the output layer generates the class prediction probability value through the Sigmoid activation function.
[0064] Optionally, the data output by the feature classifier can be a value between 0 and 1, which represents the positive prediction probability value of programmed death ligand 1. Alternatively, the output result of the feature classifier can be negative, positive indicators, etc., and the present disclosure does not limit this.
[0065] Optionally, after obtaining the training data set, the preset network model can be iteratively trained based on the training data set until the network model meets the preset convergence condition to obtain the target prediction model. The target prediction model is used to predict the predicted expression status of programmed death ligand 1 in the tumor region of interest. Thus, the user can input the data to be recognized into the target prediction model, and can accurately determine the prediction operation of the expression status of programmed death ligand 1 without performing invasive biopsy.
[0066] The data processing method provided by the present disclosure pre-obtains a training data set, which can include multiple training data pairs. Each training data pair includes the associated data of the esophageal cancer patient and the label information, where the label information is used to indicate the true expression status of programmed death ligand 1 in the tumor region of interest of the esophageal cancer patient. Then, based on the training data set, the preset network model is trained. After the network model meets the preset convergence condition, the target prediction model can be obtained. Furthermore, after inputting the data to be recognized associated with the target patient, the target prediction model can accurately predict the predicted expression status of programmed death ligand 1 in the tumor region of interest of the target patient. In addition, since the associated data includes multi-modal data such as the first tomographic imaging, the second tomographic imaging, and the clinicopathological feature vector of the tumor region of interest, the trained target prediction model can combine different tomographic imaging and clinicopathological data of the patient and other multi-modal data to accurately identify the expression status of programmed death ligand 1 in the tumor region of interest, and there is no need to perform invasive biopsy on the target patient, effectively avoiding the invasive risk.
[0067] Figure 3 It is a schematic flowchart of the data processing method provided by another embodiment of the present disclosure. Based on any of the above embodiments, as Figure 3 shown, step 201 includes:
[0068] Step 301: Obtain the original dataset, where the original dataset includes multimodal data associated with multiple esophageal cancer patients. The multimodal data includes a first original image, a second original image, and clinicopathological data. The first original image and the second original image are 18F-FDG PET / CT images of esophageal cancer patients before surgery.
[0069] Step 302: For each multimodal data, perform a preprocessing operation on the multimodal data to obtain the training data.
[0070] Step 303: Construct the training dataset based on the training data corresponding to each multimodal data.
[0071] In this embodiment, the original dataset can be obtained. The original dataset includes multimodal data associated with multiple esophageal cancer patients. The multimodal data includes a first original image, a second original image, and clinicopathological data. The first original image and the second original image are 18F-FDG PET / CT images of each esophageal cancer patient before surgery.
[0072] For example, the first original image can be a PET image, and the second original image can be a CT image.
[0073] Since the data modalities of the multimodal data are all different, for the convenience of subsequent model training, for each multimodal data, a preprocessing operation needs to be performed on it to ensure the consistency and accuracy of the data and reduce the interference of non-tumor regions of interest. Therefore, the training dataset can be constructed based on the training data corresponding to each multimodal data.
[0074] For the data processing method provided by the present disclosure, since the original data includes multimodal data of a first tomographic image, a second tomographic image, and clinicopathological data, the multimodal data can be preprocessed in advance, so as to be able to fuse data of different modalities, give full play to the advantages of data of different modalities during the model training process, and enable the trained target prediction model to achieve a more comprehensive and accurate PD-L1 prediction.
[0075] Further, based on any of the above embodiments, step 302 includes:
[0076] Perform a first preprocessing operation on the first original image and the second original image respectively to obtain the first tomographic imaging and the second tomographic imaging.
[0077] Perform a second preprocessing operation on the clinicopathological data to obtain the clinicopathological feature vector.
[0078] In this embodiment, the multimodal data includes image data and pathological data, and the image data can be 18F-FDG PET / CT images.
[0079] In order to accurately perform preprocessing operations on different-modal data, different preprocessing methods can be adopted to preprocess different-modal data.
[0080] Optionally, a first preprocessing operation can be respectively performed on the first original image and the second original image to obtain a first tomogram and a second tomogram. A second preprocessing operation is performed on the clinical pathology data to obtain a clinical pathology feature vector.
[0081] The data processing method provided by the present disclosure can accurately perform processing and fusion operations on different-modal data by adopting different preprocessing methods for different-modal data, providing a basis for improving the training efficiency of the network model.
[0082] Figure 4 As shown in the flowchart of the data processing method provided by another embodiment of the present disclosure, based on any of the above embodiments, Figure 4 The step of respectively performing a first preprocessing operation on the first original image and the second original image to obtain the first tomogram and the second tomogram includes:
[0083] Step 401: Determine the tumor region of interest in the first original image and the second original image, and generate a mask binary image corresponding to the tumor region of interest.
[0084] Step 402: Respectively perform resampling operations on the first original image and the second original image to obtain a first intermediate image and a second intermediate image.
[0085] Step 403: Perform a resampling operation on the mask binary image to obtain a resampled mask binary image.
[0086] Step 404: Adjust the window level and window width of the second intermediate image based on preset window level and window width parameters to obtain a third intermediate image.
[0087] Step 405: Perform cropping operations on the first intermediate image and the third intermediate image based on the position information associated with the resampled mask binary image and a preset image size to obtain a fourth intermediate image and a fifth intermediate image.
[0088] Step 406: Perform data normalization operations and data augmentation operations on the fourth intermediate image and the fifth intermediate image to obtain the first tomogram and the second tomogram.
[0089] In this embodiment, during the data preprocessing process, in order to ensure the consistency and accuracy of the data and reduce the interference of non-tumor regions of interest, the regions of interest of the tumor in the first original image and the second original image can be determined. Among them, the regions of interest of the tumor (Region of Interest, abbreviated as ROI) in the image data can be manually outlined by doctors or professionals.
[0090] Optionally, the DICOM files of 18F-FDG PET / CT images of all esophageal cancer patients in the training data can be imported into the LIFEx software, and screened and checked by professional nuclear medicine doctors to ensure the integrity and usability of the image data. Subsequently, the lesion ROI is manually outlined on the first original image and the second original image, and a masked binary image is generated based on the outlined ROI.
[0091] Furthermore, since the spatial resolutions of the second original image (CT image) and the first original image (PET image) are different, in order to ensure data consistency, resampling operations need to be performed on the above-mentioned image data to make them have the same voxel size, and the first intermediate image and the second intermediate image are obtained.
[0092] After the resampling operations are respectively performed on the first original image and the second original image, the voxel sizes of the first original image and the second original image change. Therefore, a corresponding resampling operation needs to be performed on the masked binary image to obtain the resampled masked binary image.
[0093] In the medical field, the HU value (Hounsfield Unit) and the SUV value (Standardized Uptake Value) are two commonly used quantitative indicators in medical images, which are respectively used in CT and PET imaging to describe the physical characteristics and metabolic activities of tissues. Among them, the HU value reflects the degree of absorption of tissues by X-rays. Dark shadows represent low absorption areas, that is, low-density areas, such as the lungs with a lot of gas; white shadows represent high absorption areas, that is, high-density areas, such as bones. In order to better observe the tumor region, the window level and window width of the second intermediate image can be adjusted based on the preset window level and window width parameters to obtain the third intermediate image. For example, the threshold of the second original image (CT image) is adjusted to -1200 ≤ HU ≤ 1500, so as to improve the contrast and brightness of the second original image (CT image) and reduce the influence of tissues in non-tumor regions of interest.
[0094] After the above image processing is completed, in order to reduce the interference of non-tumor regions of interest, the first intermediate image and the third intermediate image can be cropped based on the position information associated with the resampled masked binary image and the preset image size to obtain the fourth intermediate image and the fifth intermediate image.
[0095] Further, in order to make the data scales of the multimodal data associated with different esophageal cancer patients consistent and increase the diversity of training samples, data normalization and data augmentation operations may also be performed on the fourth intermediate image and the fifth intermediate image to obtain the first tomographic image and the second tomographic image.
[0096] The data processing method provided by the present disclosure can extract more accurate feature information through steps such as resampling, window level and width adjustment, cropping, normalization, and data augmentation of the first original image and the second original image, and can improve the stability and generalization ability of the network model during the training process.
[0097] Further, based on any of the above embodiments, the first original image and the second original image are slice sequences stored in DICOM format. Step 402 includes:
[0098] Perform a merging operation on the switching sequences of the first original image and the second original image respectively to obtain the merged first original image and the merged second original image.
[0099] Convert the pixel values of the merged first original image into a standardized index for quantifying tracer uptake to obtain the converted first original image.
[0100] Convert the pixel values of the merged second original image into relative units for quantifying tissue density to obtain the converted second original image.
[0101] Perform data storage operations on the converted first original image and the converted second original image in accordance with the NIFTI format.
[0102] Perform resampling operations on the converted first original image and the converted second original image respectively by the bilinear interpolation method to obtain the first intermediate image and the second intermediate image.
[0103] In this embodiment, the first original image and the second original image are slice sequences stored in DICOM format. Therefore, first, the SimpleITK library in Python can be used to read the DICOM format image data, merge the slice sequences to construct a complete three-dimensional image, and obtain the merged first original image and the merged second original image.
[0104] Further, the HU value is a common quantization index of the second original image (CT image), and the SUV value is a common quantization index of the first original image (PET image). Therefore, the pixel values of the merged first original image can be converted into a standardized index for quantifying tracer uptake to obtain the converted first original image. The pixel values of the merged second original image are converted into relative units for quantifying tissue density to obtain the converted second original image. And data storage operations are performed on the converted first original image and the converted second original image in accordance with the NIFTI format.
[0105] Optionally, bilinear interpolation can be used to perform resampling operations on the converted first original image and the converted second original image respectively to obtain a first intermediate image and a second intermediate image, so that the voxel sizes of the first intermediate image and the second intermediate image are both 1mm×1mm×1mm.
[0106] The data processing method provided by the present disclosure can obtain a first intermediate image and a second intermediate image with the same voxel size by performing resampling operations on the first original image and the second original image, ensuring data consistency.
[0107] Further, based on any of the above embodiments, step 403 includes:
[0108] Performing a resampling operation on the binary mask image by the nearest neighbor interpolation method to obtain a resampled binary mask image.
[0109] In this embodiment, after performing resampling operations on the first original image and the second original image respectively, the voxel sizes of the first original image and the second original image change. Therefore, in order to accurately locate the tumor region of interest subsequently, it is necessary to control the binary mask image to be consistent with the first original image and the second original image.
[0110] Therefore, a resampling operation can be performed on the binary mask image by the nearest neighbor interpolation method to obtain a resampled binary mask image.
[0111] For example, the voxel sizes of all images can be adjusted to 1mm×1mm×1mm.
[0112] The data processing method provided by the present disclosure can ensure that the image size of the resampled binary mask image is the same as that of the first intermediate image and the second intermediate image by performing a resampling operation on the binary mask image, so that the tumor region of interest can be located and cropped more accurately subsequently.
[0113] Further, based on any of the above embodiments, the data augmentation operation includes one or more of random horizontal flipping, random vertical flipping, random rotation, and random translation.
[0114] Optionally, in order to increase the diversity of training samples, data augmentation operations may also be performed on the fourth intermediate image and the fifth intermediate image.
[0115] Among them, the data augmentation operations include one or more of random horizontal flipping, random vertical flipping, random rotation, and random translation.
[0116] The data processing method provided by the present disclosure can make the data scales of multi-modal data associated with different esophageal cancer patients consistent and increase the diversity of training samples through data normalization operations and data augmentation operations on the fourth intermediate image and the fifth intermediate image, thereby improving the prediction accuracy of the target prediction model obtained by training.
[0117] Further, on the basis of any of the above embodiments, the clinicopathological data includes numerical data and categorical data.
[0118] The second preprocessing operation on the clinicopathological data to obtain the clinicopathological feature vector includes:
[0119] Performing feature normalization processing on the numerical data to obtain a first vector.
[0120] Performing one-hot encoding processing on the categorical data to obtain a second vector.
[0121] Determining the first vector and the second vector as the clinicopathological feature vector.
[0122] In this embodiment, the clinicopathological data includes numerical data and categorical data. For example, the clinicopathological data includes numerical data such as age, BMI, tumor size, SUVmax, and SUVave, and the clinicopathological data includes categorical data such as gender, smoking history, drinking history, site grade, pathological type, histological grade, and tumor stage.
[0123] Different preprocessing methods can be adopted for different types of clinicopathological data.
[0124] Optionally, feature normalization processing may be performed on the numerical data to obtain a first vector. Ensure that the first vector has a consistent scale during training. Determine the first vector and the second vector as the clinicopathological feature vector.
[0125] The data processing method provided by the present disclosure can construct the clinicopathological data into a standardized input vector through vectorization operations, for subsequent feature fusion and learning by a deep learning model.
[0126] Optionally, based on any of the above embodiments, the network model includes a first feature extractor, a second feature extractor, and a feature classifier.
[0127] Figure 5 A schematic flowchart of a data processing method provided by another embodiment of the present disclosure. Based on any of the above embodiments, as Figure 5 shown, step 202 includes:
[0128] Step 501: For each piece of training data, input the first tomography into the first feature extractor to obtain a first feature vector.
[0129] Step 502: Input the second tomography into the second feature extractor to obtain a second feature vector.
[0130] Step 503: Perform a vector concatenation operation on the first feature vector, the second feature vector, and the clinicopathological feature vector to obtain a target feature vector.
[0131] Step 504: Input the target feature vector into the feature classifier to obtain a prediction result output by the feature classifier, where the prediction result is the predicted status of programmed death ligand 1 in the tumor region of interest.
[0132] Step 505: Calculate the prediction result and the label information through a preset loss function to obtain a loss value of the network model.
[0133] Step 506: Determine whether the network model meets a preset convergence condition based on the loss value.
[0134] Step 507: If it is satisfied, determine the network model as the target prediction model.
[0135] Step 508: If it is not satisfied, perform a reverse gradient adjustment on the parameters of the network model based on the loss value, and return to execute the step of inputting the first tomography into the first feature extractor for each piece of training data to obtain a first feature vector until the loss value meets the preset convergence condition to obtain the target prediction model.
[0136] In this embodiment, the network model includes a first feature extractor and a second feature extractor. The first feature extractor is used to extract features from the first tomographic image, and the second feature extractor is used to extract features from the second tomographic image.
[0137] Optionally, the input to the feature extractor is a tensor of size 96×96×96, adopting the ResNet10 architecture, consisting of a 10-layer deep convolutional neural network, which mainly contains 4 residual blocks. Each residual block contains multiple convolutional layers, batch normalization layers (Batch Normalization, BN), and activation functions (ReLU). At the same time, skip connections are introduced to avoid the vanishing gradient problem and enhance the expression ability of the network. Finally, the features are reduced to 1×1×512 through a global pooling layer and converted into a 1×512 one-dimensional feature vector, ready for subsequent feature fusion.
[0138] Therefore, for each training data, the first tomogram is input into the first feature extractor to obtain a 1×512 first feature vector. The second tomogram is input into the second feature extractor to obtain a 1×512 second feature vector.
[0139] In addition to the first tomogram and the second tomogram, the network model also combines clinical pathological information as an additional input. The clinical pathological information is constructed as a vector of size 1×47 and converted into a 1×512 feature vector through a fully connected layer.
[0140] Since the training data is multi-modal data, in order to perform the fusion operation on multi-modal data, the first feature vector, the second feature vector, and the clinical pathological feature vector can be concatenated to obtain a target feature vector. The target feature vector is input into the feature classifier to obtain the prediction result output by the feature classifier, and the prediction result is the predicted status of programmed death ligand 1 in the tumor region of interest.
[0141] Among them, the feature classifier consists of two fully connected layers, and the output layer generates the class prediction probability value through the Sigmoid activation function.
[0142] Furthermore, after obtaining the prediction result, the loss value of the network model can be calculated through a preset loss function for the prediction result and the label information.
[0143] Among them, the calculation of the loss value can be realized by using formula 1:
[0144]
[0145] Among them, σ(x) is the Sigmoid function, p i represents the probability that the sample x i is predicted as a positive example, y i represents the label information of the sample x i and p cis the weight of the positive sample category, N is the size of the Batchsize, and c is the number of categories.
[0146] Further, the convergence condition can be preset. Therefore, after determining the loss value, it can be determined whether it meets the convergence condition to determine whether to continue the iterative training of the model.
[0147] Among them, the convergence condition can be that the loss value is less than a preset threshold, or the convergence condition can be that the difference between the loss values of two rounds of training is less than a preset difference threshold, etc. The present disclosure does not limit this.
[0148] Therefore, if it is determined based on the loss value that the network model meets the preset convergence condition, the network model can be determined as the target prediction model. On the contrary, if it is determined based on the loss value that the network model does not meet the preset convergence condition, the parameters of the network model can be adjusted by backpropagation gradient based on the loss value, and the step of inputting the first tomogram into the first feature extractor to obtain the first feature vector for each training data can be returned and executed until the loss value meets the preset convergence condition to obtain the target prediction model.
[0149] As an implementable method, the training process of the network model can also adopt five-fold cross-validation, with 150 Epochs trained for each fold, the Batch size set to 8, the optimizer using SGD, and the initial learning rate being 0.001.
[0150] Users can select a suitable training method to train the network model according to actual needs, and the present disclosure does not limit this.
[0151] Figure 6 is the structural schematic diagram of the network model provided by the embodiment of the present disclosure, as Figure 6 shown, the network model includes a first feature extractor 61, a second feature extractor 62, and a feature classifier 63. Based on the above model architecture, the first tomogram 64 is input into the first feature extractor 61 to obtain the first feature vector 65. The second tomogram 66 is input into the second feature extractor 62 to obtain the second feature vector 67. The first feature vector 65, the second feature vector 67, and the clinicopathological feature vector 68 are subjected to a vector splicing operation to obtain the target feature vector 69. The target feature vector 69 is input into the feature classifier 63 to obtain the prediction result 610 output by the feature classifier.
[0152] The data processing method provided by the present disclosure can calculate the loss value of the network model during the training process, so as to be able to determine whether the network model reaches the convergence condition based on the loss value, and adjust the parameters of the network model when the convergence condition is not met, so as to be able to continuously train the network model to enable it to have the ability to accurately identify whether there is programmed death ligand 1 in the tumor region of interest and improve the prediction accuracy of the model.
[0153] Optionally, based on any of the above embodiments, after step 504, the method further includes:
[0154] Calculating the prediction result and the label information through a preset loss function to obtain a loss value of the network model.
[0155] Calculating the prediction result and the label information through at least one preset index calculation formula to obtain at least one model evaluation index, where the model evaluation index includes one or more of an area under the curve index, a sensitivity index, a specificity index, and an accuracy index.
[0156] Determining whether the network model meets a preset convergence condition based on the loss value and at least one model evaluation index.
[0157] If it is satisfied, determining the network model as the target prediction model.
[0158] If it is not satisfied, performing a reverse gradient adjustment on the parameters of the network model based on the loss value, and returning to execute the step of inputting the first tomographic imaging into the first feature extractor for each training data to obtain a first feature vector until the loss value meets the preset convergence condition to obtain the target prediction model.
[0159] In this embodiment, in addition to the loss value, the evaluation indexes of the network model may further include one or more of an area under the curve index (Area Under the Curve, abbreviated as AUC), a sensitivity index, a specificity index, and an accuracy index.
[0160] Among them, after inputting the target feature vector into the feature classifier, a prediction result output by the feature classifier can be obtained, where the prediction result is the predicted state of programmed death ligand 1 in the tumor region of interest.
[0161] Therefore, the loss value of the network model can be obtained by calculating the prediction result and the label information through a preset loss function. And at least one model evaluation index can be obtained by calculating the prediction result and the label information through at least one preset index calculation formula.
[0162] For example, the calculation of the sensitivity index, the specificity index, and the accuracy index can be implemented through Formula 2-4:
[0163]
[0164] Among them, TP (True Positive) is the number of true positives, FN (False Negative) is the number of false negatives, TN (True Negative) is the number of true negatives, and FP (False Positive) is the number of false positives. The number of true positives, the number of false negatives, the number of true negatives, and the number of false positives can be calculated based on the prediction results output by the model and the accurate label information pre-labeled.
[0165] And AUC (Area Under the Curve), that is, the area under the curve, refers to the area under the ROC curve (Receiver Operating Characteristic Curve). It comprehensively considers both false positives and true positives. Compared with individual sensitivity or specificity, it can evaluate the model performance more comprehensively. The ROC curve shows the performance of the model by plotting the relationship between the false positive rate (FPR, False Positive Rate) and the true positive rate (TPR, True Positive Rate, that is, sensitivity), reflecting the performance of the model at different thresholds. The AUC value is the area under the ROC curve, which represents the ability of the model to distinguish positive samples and negative samples. The value of AUC ranges from 0 to 1, and the closer it is to 1, the better the model performance.
[0166] Therefore, after calculating the loss value and at least one evaluation index respectively, the convergence condition of the network model can be comprehensively determined in combination with the above content.
[0167] For example, if the loss value of the current network model is very low, but the AUC value is small, it indicates that the model performance cannot meet the expectations. At this time, the network model can be continuously iteratively trained.
[0168] Or, if the current AUC value is close to 1, but the loss value of the network model is very high, it indicates that the prediction accuracy of the network model is not high. At this time, the network model can be continuously iteratively trained.
[0169] Or, different weight parameters can be set for different evaluation indexes. The weight parameter can be set by the user according to the actual usage scenario, or the weight parameter can be the default empirical value of the system. The present disclosure does not limit this. Weighted calculation is performed based on the weight parameter and at least one evaluation index. Determine whether both the weighted result and the loss value meet the convergence condition. If so, it indicates that the training of the network model is completed. Otherwise, continue to iteratively train the network model.
[0170] In practical applications, the user can set the convergence condition according to actual needs and select the corresponding evaluation index for model training according to actual needs. The present disclosure does not limit this.
[0171] Further, if it is determined based on the loss value and at least one model evaluation metric that the network model meets the preset convergence condition, the network model is determined as the target prediction model.
[0172] Conversely, if it is determined based on the loss value and at least one model evaluation metric that the network model does not meet the preset convergence condition, the parameters of the network model are adjusted by back-gradient based on the loss value, and the step of inputting the first tomographic imaging into the first feature extractor for each training data to obtain the first feature vector is returned, until the loss value meets the preset convergence condition to obtain the target prediction model.
[0173] Figure 7 The ROC curve schematic diagram provided by the embodiments of the present disclosure can use the data of 265 patients collected for the training operation of the model to be trained, and the data of 66 patients for the test operation of the model to be trained. As Figure 7 shown, the AUC71 of the training set (Train) is 0.93, and the AUC72 of the test set (Text) is 0.88, indicating the superiority of the target prediction model in the prediction performance of the PD-L1 status.
[0174] The data processing method provided by the present disclosure can further improve the prediction accuracy of the model by jointly training the network model based on one or more of the loss value, the area under the curve metric, the sensitivity metric, the specificity metric, and the accuracy metric.
[0175] Figure 8 The flowchart of the data processing method provided by the embodiments of the present disclosure is shown as Figure 8 shown, and the method includes:
[0176] Step 801, obtain the data to be recognized associated with the target patient, where the data to be recognized includes a first image, a second image, and target clinicopathological data, the first image and the second image are 18F-FDG PET / CT images associated with the target patient, and the target patient is a patient with esophageal cancer.
[0177] Step 802, perform a preprocessing operation on the data to be recognized to obtain target data, where the image parameters associated with the first image and the second image in the target data are consistent, and the image parameters include voxel size, window level and width, image size, and numerical scale.
[0178] Step 803, input the target data into a preset target prediction model to obtain a prediction result output by the target prediction model, where the prediction result is the predicted expression status of programmed death ligand 1 in the tumor region of interest of the target patient.
[0179] The target prediction model is obtained by training based on the data processing method described in any of the above embodiments.
[0180] The execution subject of this embodiment is a data processing device. The data processing device can be coupled to a server, and the server can communicate with a terminal device. After obtaining the data to be recognized associated with the target patient, the data to be recognized can be preprocessed, and the target data obtained by the preprocessing can be subjected to image processing through a pre-trained target prediction model to determine the predicted expression status of programmed death ligand 1 in the tumor region of interest of the target patient.
[0181] In this embodiment, a target prediction model can be pre-obtained by training according to any of the above embodiments. The prediction model can accurately predict whether programmed death ligand 1 exists in the tumor region of interest of the target patient without performing invasive biopsy on the target patient.
[0182] Optionally, data to be recognized associated with the target patient can be obtained, where the data to be recognized includes a first image, a second image, and target clinicopathological data. The first image and the second image are 18F-FDG PET / CT images associated with the target patient. The target patient can be an esophageal cancer patient.
[0183] For example, the first image can be a positron emission tomography image, and the second image can be a computed tomography image. The first image and the second image can be obtained by an imaging technique combining positron emission tomography (PET) and computed tomography (CT) using 18F-FDG as a tracer.
[0184] The target clinicopathological data includes but is not limited to numerical features such as age, BMI, tumor size, SUVmax, and SUVave, as well as categorical features such as gender, smoking history, drinking history, site grade, pathological type, histological grade, and tumor stage.
[0185] It should be noted that the above target clinicopathological data is obtained after obtaining the full authorization of the target patient.
[0186] Furthermore, since the first image, the second image, and the target clinicopathological data are data of different modalities respectively, in order to improve the accuracy of subsequent data processing, a preprocessing operation can be performed on the data to be recognized to obtain target data, where the image parameters associated with the first image and the second image in the target data are consistent, and the image parameters include voxel size, window level and width, image size, and numerical scale.
[0187] Among them, the preprocessing operation includes but is not limited to steps such as resampling, window level and width adjustment, cropping, and normalization. The specific implementation method can be seen in the above embodiments and will not be elaborated here.
[0188] It should be noted that during the model training process, preprocessing steps such as resampling, window level and window width adjustment, cropping, normalization, and data augmentation can be performed. During the model usage process, data augmentation operations can be not performed on the data to be recognized, and only preprocessing steps such as resampling, window level and window width adjustment, cropping, and normalization can be performed on it.
[0189] Optionally, after completing the preprocessing of the data to be recognized, the target data can be input into the target prediction model. The target prediction model can include two feature extractors and one feature classifier. After the target data is input into the target prediction model, the image features of the first image and the second image can be extracted by the two feature extractors respectively. In order to achieve the fusion of multimodal data, the image features of the first image and the second image and the target clinicopathological data features corresponding to the target clinicopathological data can be concatenated to obtain a concatenated vector. The concatenated vector is input into the feature classifier. Therefore, the feature classifier can accurately predict the predicted expression status of programmed death ligand 1 in the tumor region of interest of the target patient.
[0190] Among them, the feature extractor adopts the ResNet10 architecture, which consists of a 10-layer deep convolutional neural network and mainly contains 4 residual blocks. Each residual block contains multiple convolutional layers, batch normalization layers (BatchNormalization, BN), and activation functions (ReLU). At the same time, skip connections are introduced to avoid the problem of gradient disappearance and enhance the expression ability of the network. Finally, feature dimensionality reduction operations are performed through the global pooling layer to generate the image features of the first image and the second image. The feature classifier consists of two fully connected layers, and the output layer generates the class prediction probability value through the Sigmoid activation function.
[0191] The data processing method provided by the present disclosure can make the image parameters associated with the first image and the second image consistent by preprocessing the data to be recognized associated with the target patient, and then the target data can be input into the preset target prediction model, and the predicted expression status of programmed death ligand 1 in the tumor region of interest of the target patient can be accurately predicted based on the target prediction model. Therefore, the prediction of the PD-L1 expression status can be quickly and accurately achieved without performing invasive biopsy on the target patient.
[0192] Figure 9 The structural schematic diagram of the data processing device provided by the embodiment of the present disclosure is as Figure 9As shown, the device includes an acquisition module 91 and a training module 92. Among them, the acquisition module 91 is used to acquire a training data set, which includes a plurality of training data pairs. The training data pairs include the associated data of esophageal cancer patients and label information. The associated data includes the first tomographic image, the second tomographic image, and the clinicopathological feature vector of the tumor region of interest. The first tomographic image and the second tomographic image are associated with the same image parameters, and the image parameters include voxel size, window level and width, image size, and numerical scale. The label information is used to indicate the true expression state of programmed death ligand 1 in the tumor region of interest. The training module 92 is used to iteratively train a preset network model based on the training data set until the network model meets the preset convergence condition, and obtain a target prediction model. The target prediction model is used to predict the predicted expression state of programmed death ligand 1 in the tumor region of interest.
[0193] Further, based on any of the above embodiments, the acquisition module is used to: acquire an original data set, which includes multimodal data associated with a plurality of esophageal cancer patients. The multimodal data includes a first original image, a second original image, and clinicopathological data. The first original image and the second original image are 18F-FDG PET / CT images before surgery of esophageal cancer patients. For each multimodal data, perform a preprocessing operation on the multimodal data to obtain the training data. Construct the training data set based on the training data corresponding to each multimodal data.
[0194] Further, based on any of the above embodiments, the acquisition module is used to: perform a first preprocessing operation on the first original image and the second original image respectively to obtain the first tomographic image and the second tomographic image. Perform a second preprocessing operation on the clinicopathological data to obtain the clinicopathological feature vector.
[0195] Further, based on any of the above embodiments, the obtaining module is configured to: determine the regions of interest of tumors in the first original image and the second original image, and generate a binary mask image corresponding to the regions of interest of tumors. Respectively perform resampling operations on the first original image and the second original image to obtain a first intermediate image and a second intermediate image. Perform a resampling operation on the binary mask image to obtain a resampled binary mask image. Perform an adjustment operation on the window level and window width of the second intermediate image based on preset window level and window width parameters to obtain a third intermediate image. Perform cropping operations on the first intermediate image and the third intermediate image based on the position information associated with the resampled binary mask image and a preset image size to obtain a fourth intermediate image and a fifth intermediate image. Perform data normalization operations and data augmentation operations on the fourth intermediate image and the fifth intermediate image to obtain the first tomographic image and the second tomographic image.
[0196] Further, based on any of the above embodiments, the first original image and the second original image are slice sequences stored in DICOM format. The obtaining module is configured to: respectively perform a merging operation on the switching sequences of the first original image and the second original image to obtain a merged first original image and a merged second original image. Convert the pixel values of the merged first original image into a standardized index for quantifying tracer uptake to obtain a converted first original image. Convert the pixel values of the merged second original image into relative units for quantifying tissue density to obtain a converted second original image. Perform data storage operations on the converted first original image and the converted second original image in NIFTI format. Respectively perform resampling operations on the converted first original image and the converted second original image by bilinear interpolation to obtain the first intermediate image and the second intermediate image.
[0197] Further, based on any of the above embodiments, the obtaining module is configured to: perform a resampling operation on the binary mask image by nearest neighbor interpolation to obtain a resampled binary mask image.
[0198] Further, based on any of the above embodiments, the data augmentation operations include one or more of random horizontal flipping, random vertical flipping, random rotation, and random translation.
[0199] Further, based on any of the above embodiments, the clinical pathological data includes numerical data and categorical data. The obtaining module is configured to: perform feature normalization processing on the numerical data to obtain a first vector. Perform one-hot encoding processing on the categorical data to obtain a second vector. Determine the first vector and the second vector as the clinical pathological feature vector.
[0200] Further, based on any of the above embodiments, the network model includes a first feature extractor, a second feature extractor, and a feature classifier.
[0201] Further, based on any of the above embodiments, the training module is configured to: for each training data, input the first tomogram into the first feature extractor to obtain a first feature vector. Input the second tomogram into the second feature extractor to obtain a second feature vector. Perform a vector concatenation operation on the first feature vector, the second feature vector, and the clinicopathological feature vector to obtain a target feature vector. Input the target feature vector into the feature classifier to obtain a prediction result output by the feature classifier, where the prediction result is the predicted status of programmed death ligand 1 in the tumor region of interest. Calculate the prediction result and the label information through a preset loss function to obtain a loss value of the network model. Determine whether the network model meets a preset convergence condition based on the loss value. If it meets, determine the network model as the target prediction model. If not, perform a reverse gradient adjustment on the parameters of the network model based on the loss value, and return to execute the step of inputting the first tomogram into the first feature extractor for each training data to obtain a first feature vector until the loss value meets the preset convergence condition to obtain the target prediction model.
[0202] Further, based on any of the above embodiments, the training module is further configured to: calculate the prediction result and the label information through a preset loss function to obtain a loss value of the network model. Calculate the prediction result and the label information through at least one preset index calculation formula to obtain at least one model evaluation index, where the model evaluation index includes one or more of the area under the curve index, the sensitivity index, the specificity index, and the accuracy index. Determine whether the network model meets a preset convergence condition based on the loss value and the at least one model evaluation index. If it meets, determine the network model as the target prediction model. If not, perform a reverse gradient adjustment on the parameters of the network model based on the loss value, and return to execute the step of inputting the first tomogram into the first feature extractor for each training data to obtain a first feature vector until the loss value meets the preset convergence condition to obtain the target prediction model.
[0203] Figure 10 The structural schematic diagram of the data processing device provided by the embodiments of the present disclosure is as Figure 10As shown in the figure, the device includes: a data acquisition module 1001, a preprocessing module 1002, and a prediction module 1003. Among them, the data acquisition module 1001 is used to acquire the data to be recognized associated with the target patient. The data to be recognized includes a first image, a second image, and target clinicopathological data. The first image and the second image are 18F-FDG PET / CT images associated with the target patient, and the target patient is a patient with esophageal cancer. The preprocessing module 1002 is used to perform preprocessing operations on the data to be recognized to obtain target data. In the target data, the image parameters associated with the first image and the second image are consistent. The image parameters include voxel size, window level and width, image size, and numerical scale. The prediction module 1003 is used to input the target data into a preset target prediction model to obtain a prediction result output by the target prediction model. The prediction result is the predicted expression status of programmed death ligand 1 in the tumor region of interest of the target patient. The target prediction model is trained based on the data processing device described in any of the above embodiments.
[0204] The device provided in this embodiment can be used to execute the technical solutions of the above method embodiments. The implementation principles and technical effects are similar, and will not be elaborated here in this embodiment.
[0205] To implement the above embodiments, the present disclosure also provides a computer-readable storage medium. Computer-executable instructions are stored in the computer-readable storage medium. When the processor executes the computer-executable instructions, the information display method described in any of the above embodiments is implemented.
[0206] To implement the above embodiments, the present disclosure also provides a computer program product, including a computer program. When the computer program is executed by a processor, the information display method described in any of the above embodiments is implemented.
[0207] To implement the above embodiments, the present disclosure also provides an electronic device, including: a processor and a memory;
[0208] The memory stores computer-executable instructions;
[0209] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the information display method described in any of the above embodiments.
[0210] Figure 11Schematic structural diagram of the electronic device provided by the embodiments of the present disclosure. The electronic device 1100 may be a terminal device or a server. Among them, the terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers, portable media players (PMPs), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 11 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0211] As Figure 11 shown, the electronic device 1100 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 1101, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1102 or the program loaded from the storage device 1108 into the random access memory (RAM) 1103. In the RAM 1103, various programs and data required for the operation of the electronic device 1100 are also stored. The processing device 1101, the ROM 1102, and the RAM 1103 are connected to each other through a bus 1104. The input / output (I / O) interface 1105 is also connected to the bus 1104.
[0212] Generally, the following devices may be connected to the I / O interface 1105: an input device 1106 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1107 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1108 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1109. The communication device 1109 may allow the electronic device 1100 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 11 the electronic device 1100 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.
[0213] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowchart may be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network via the communication device 1109, or installed from the storage device 1108, or installed from the ROM 1102. When the computer program is executed by the processing device 1101, the above-described functions defined in the method of the embodiment of the present disclosure are performed.
[0214] It should be noted that the above-mentioned computer-readable medium in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0215] The above-mentioned computer-readable medium may be included in the above-mentioned electronic device; or may exist separately without being assembled into the electronic device.
[0216] The above-mentioned computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above embodiment.
[0217] Computer program code for performing the operations of this disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, execute as a stand-alone software package, execute partially on the user's computer and partially on a remote computer, or execute entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0218] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0219] The units involved in the embodiments described in this disclosure may be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases. For example, the first acquisition unit may also be described as "the unit for acquiring at least two Internet protocol addresses".
[0220] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), and so on.
[0221] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0222] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described apparatuses may refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0223] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: ROM, RAM, magnetic disk, or optical disk and other various media that can store program codes.
[0224] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure and are not intended to limit them; although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present disclosure.
Claims
1. A data processing method, characterized in that, Including: Obtain a training data set, where the training data set includes a plurality of training data pairs. Each training data pair includes associated data of esophageal cancer patients and label information. The associated data includes a first tomographic image, a second tomographic image, and a clinicopathological feature vector of a tumor region of interest. The first tomographic image and the second tomographic image are associated with the same image parameters. The image parameters include voxel size, window level and width, image size, and numerical scale. The label information is used to indicate the true expression state of programmed death ligand 1 in the tumor region of interest; Based on the training data set, perform iterative training on a preset network model until the network model meets the preset convergence condition, and obtain a target prediction model. The target prediction model is used to predict the predicted expression state of programmed death ligand 1 in the tumor region of interest.
2. The method according to claim 1, wherein The obtaining of the training data set includes: Obtain an original data set, where the original data set includes multimodal data associated with a plurality of esophageal cancer patients. The multimodal data includes a first original image, a second original image, and clinicopathological data. The first original image and the second original image are 18F-FDG PET / CT images before surgery of esophageal cancer patients; For each multimodal data, perform a first preprocessing operation on the first original image and the second original image respectively to obtain the first tomographic image and the second tomographic image, and perform a second preprocessing operation on the clinicopathological data to obtain the clinicopathological feature vector, so as to obtain the training data; Construct the training data set based on the training data corresponding to each multimodal data.
3. The method according to claim 2, characterized in that, The performing of the first preprocessing operation on the first original image and the second original image respectively to obtain the first tomographic image and the second tomographic image includes: Determine the tumor region of interest in the first original image and the second original image, and generate a mask binary image corresponding to the tumor region of interest; Perform resampling operations on the first original image and the second original image respectively to obtain a first intermediate image and a second intermediate image; Perform a resampling operation on the mask binary image to obtain a resampled mask binary image; Based on preset window level and width parameters, perform an adjustment operation on the window level and width of the second intermediate image to obtain a third intermediate image; Based on the position information associated with the resampled mask binary image and a preset image size, perform cropping operations on the first intermediate image and the third intermediate image to obtain a fourth intermediate image and a fifth intermediate image; Perform data normalization operations and data augmentation operations on the fourth intermediate image and the fifth intermediate image to obtain the first tomographic image and the second tomographic image.
4. The method according to claim 3, wherein The first original image and the second original image are slice sequences stored in DICOM format; The performing of the resampling operations on the first original image and the second original image respectively to obtain the first intermediate image and the second intermediate image includes: Perform a merging operation on the switching sequences of the first original image and the second original image respectively to obtain the merged first original image and the merged second original image; Convert the pixel values of the merged first original image into a standardized index for quantifying tracer uptake to obtain the converted first original image; Convert the pixel values of the merged second original image into relative units for quantifying tissue density to obtain the converted second original image; Perform a data storage operation on the converted first original image and the converted second original image in accordance with the NIFTI format; Perform a resampling operation on the converted first original image and the converted second original image respectively by means of bilinear interpolation to obtain the first intermediate image and the second intermediate image.
5. The method according to claim 2, wherein The clinical pathological data includes numerical data and categorical data; The performing a second preprocessing operation on the clinical pathological data to obtain the clinical pathological feature vector includes: Perform a feature normalization process on the numerical data to obtain a first vector; Perform a one-hot encoding process on the categorical data to obtain a second vector; Determine the first vector and the second vector as the clinical pathological feature vector.
6. The method according to any one of claims 1-5, characterized in that, The network model includes a first feature extractor, a second feature extractor, and a feature classifier; The iteratively training a preset network model based on the training data set until the network model meets a preset convergence condition to obtain a target prediction model includes: For each training data, input the first tomographic imaging into the first feature extractor to obtain a first feature vector; Input the second tomographic imaging into the second feature extractor to obtain a second feature vector; Perform a vector concatenation operation on the first feature vector, the second feature vector, and the clinical pathological feature vector to obtain a target feature vector; Input the target feature vector into the feature classifier to obtain a prediction result output by the feature classifier, and the prediction result is the predicted status of programmed death ligand 1 in the tumor region of interest; Calculate the prediction result and the label information by means of a preset loss function to obtain a loss value of the network model; Determine whether the network model meets the preset convergence condition based on the loss value; If it meets, determine the network model as the target prediction model; If it does not meet, perform a reverse gradient adjustment on the parameters of the network model based on the loss value, and return to execute the step of inputting the first tomographic imaging into the first feature extractor for each training data to obtain a first feature vector until the loss value meets the preset convergence condition to obtain the target prediction model.
7. A data processing method, characterized in that, Including: Obtain identification data associated with a target patient, where the identification data includes a first image, a second image, and target clinical pathological data, the first image and the second image are 18F-FDG PET / CT images associated with the target patient, and the target patient is a patient with esophageal cancer; Perform preprocessing operations on the data to be recognized to obtain target data, where the image parameters associated with the first image and the second image in the target data are consistent, and the image parameters include voxel size, window level and width, image size, and numerical scale; Input the target data into a preset target prediction model to obtain a prediction result output by the target prediction model, where the prediction result is the predicted expression status of programmed death ligand 1 in the tumor region of interest of the target patient; The target prediction model is trained based on the data processing method according to any one of claims 1-6.
8. A data processing device, characterized in that, It includes: An acquisition module for acquiring a training data set, where the training data set includes a plurality of training data pairs, the training data pairs include the associated data of esophageal cancer patients and label information, the associated data includes the first tomographic image and the second tomographic image of the tumor region of interest and a clinicopathological feature vector, the image parameters associated with the first tomographic image and the second tomographic image are consistent, and the image parameters include voxel size, window level and width, image size, and numerical scale, and the label information is used to indicate the true expression status of programmed death ligand 1 in the tumor region of interest; A training module for iteratively training a preset network model based on the training data set until the network model meets a preset convergence condition to obtain a target prediction model, where the target prediction model is used to predict the predicted expression status of programmed death ligand 1 in the tumor region of interest.
9. A data processing device, characterized in that, It includes: A data acquisition module for acquiring the data to be recognized associated with the target patient, where the data to be recognized includes a first image, a second image, and target clinicopathological data, the first image and the second image are 18F-FDG PET / CT images associated with the target patient, and the target patient is an esophageal cancer patient; A preprocessing module for performing preprocessing operations on the data to be recognized to obtain target data, where the image parameters associated with the first image and the second image in the target data are consistent, and the image parameters include voxel size, window level and width, image size, and numerical scale; A prediction module for inputting the target data into a preset target prediction model to obtain a prediction result output by the target prediction model, where the prediction result is the predicted expression status of programmed death ligand 1 in the tumor region of interest of the target patient; The target prediction model is trained based on the data processing device according to claim 8.
10. An electronic device, characterized in that, It includes: A processor and a memory; The memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory, so that the processor executes the data processing method according to any one of claims 1 to 7.
11. A computer-readable storage medium, characterized in that, Computer execution instructions are stored in the computer-readable storage medium, and when the processor executes the computer execution instructions, the data processing method according to any one of claims 1 to 7 is implemented.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the data processing method according to any one of claims 1 to 7 is implemented.