Head and neck tumor intelligent auxiliary analysis method and system based on multi-modal deep learning
The multimodal deep learning method is used to construct a multimodal head and neck tumor analysis model, which solves the problem of insufficient utilization of single mode data in the diagnosis and treatment of head and neck tumors, and realizes high-precision diagnosis and treatment strategy prediction, which improves the accuracy and efficiency of diagnosis and treatment.
Patent Information
- Application Number
- CN202510667980.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-08-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art has problems such as insufficient utilization of single modal data, difficulty in fusion of multimodal data, inefficient efficiency and difficulty in achieving personalized treatment in the diagnosis and treatment of head and neck tumors, resulting in insufficient accuracy of diagnosis and treatment.
The multimodal deep learning method is adopted, and a multimodal head and neck tumor analysis model is combined with efficient data set construction and labeling schemes, and a multimodal head and neck tumor analysis model is constructed to perform head and neck tumor feature recognition and diagnosis and treatment strategy prediction, including data acquisition, model pre-training, migration training and performance verification.
It improves the accuracy of diagnosis and treatment of head and neck tumors, can achieve high segmentation accuracy and stability on various types of head and neck tumor imaging data, has high prediction accuracy, and effectively distinguish different treatment methods and prognosis conditions.
Smart Images

Figure CN120496808A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent diagnosis and treatment technology, and in particular to an intelligent assisted analysis method and system for head and neck tumors based on multimodal deep learning. Background Art
[0002] Before the rise of multimodal deep learning technology, clinical research on the diagnosis and treatment of head and neck cancers primarily relied on single-modality data and traditional statistical and machine learning methods for disease diagnosis and treatment. These methods have significant limitations. First, the limitations of a single data source are very obvious. Clinical data can reflect basic patient information, but it is difficult to fully present the entire disease picture and accurately predict treatment efficacy and prognosis. Although medical imaging data can display anatomical structure and functional information, it is not sensitive enough to molecular changes in the disease. Genomic data can reveal genetic characteristics, but lacks anatomical and functional information, making it difficult to directly use in clinical decision-making.
[0003] Furthermore, traditional technologies face technical hurdles when integrating multiple data types. They typically rely on simple statistical methods and fail to fully exploit the complementarity and correlation between data, thus affecting the accuracy of diagnosis and treatment. Traditional methods have limited ability to remove redundant information and reduce noise when processing multimodal data, resulting in reduced model performance. Furthermore, traditional methods primarily rely on static data and population analysis, making it difficult to implement personalized treatment plans and dynamic monitoring, and failing to fully account for individual patient differences.
[0004] Finally, traditional methods are inefficient and time-consuming when processing large-scale, diverse medical data, making them unable to meet the clinical requirements of real-time and high efficiency. Consequently, single-modality technologies have significant limitations in clinical research, as they are unable to integrate multi-source data, resulting in insufficient information utilization and limiting the accuracy of diagnosis and treatment.
[0005] The development of multimodal deep learning technology has improved the accuracy and efficiency of diagnosis and treatment by integrating multiple data types, providing new possibilities for solving these problems. Summary of the Invention
[0006] In order to solve the technical problems existing in the prior art, the present invention provides the following technical solutions: In one aspect, a method for intelligent assisted analysis of head and neck tumors based on multimodal deep learning is provided. The method is implemented by an electronic device and includes: S1. Collect clinical examination data of head and neck cancer patients; S2. Inputting the head and neck tumor clinical examination data into a pre-deployed multimodal head and neck tumor analysis model, identifying the patient's head and neck tumor data features through the multimodal head and neck tumor analysis model, and outputting a tumor diagnosis and treatment label that matches the head and neck tumor data features; S3. Record the head and neck tumor data characteristics and tumor diagnosis and treatment labels, and write them into the HIS electronic medical record of the head and neck tumor patient to generate a corresponding head and neck tumor pre-diagnosis report.
[0007] Preferably, the method for generating the multimodal head and neck tumor analysis model comprises: Collect multi-dimensional and multimodal clinical data of several head and neck cancer patients and write them into the database to build a multi-dimensional and multimodal clinical database; Based on the Transformer segmentation model, pre-training learning of head and neck tumor data features is performed on a public head and neck image segmentation dataset, and a head and neck tumor segmentation model is constructed; wherein the head and neck tumor data features are annotated with tumor diagnosis and treatment labels corresponding to the features; Using the head and neck tumor data in the multi-dimensional and multi-modal clinical database, performing migration training on the pre-trained head and neck tumor segmentation model to construct a multi-modal head and neck tumor analysis model; Collect clinical data sets of head and neck patients to clinically validate the multimodal head and neck tumor analysis model: If the clinical performance test is qualified, the multimodal head and neck tumor analysis model will be put into clinical deployment and application; If the clinical performance test fails, the above steps are repeated to retrain and construct the multimodal head and neck tumor analysis model.
[0008] Preferably, the head and neck image segmentation public dataset adopts the HeckTOR 2022 dataset.
[0009] Preferably, the multi-dimensional and multi-modal clinical data of several head and neck cancer patients are collected and written into a library to construct a multi-dimensional and multi-modal clinical database, including: Parse the HeckTOR 2022 dataset to obtain the data structure in the dataset; According to the data dimensions and modalities in the data structure of the HeckTOR 2022 dataset, clinical examination data of several head and neck cancer patients in the corresponding data dimensions and modalities are collected; The clinical examination data of each head and neck cancer patient in the corresponding data dimensions and modalities are written into the preset MySQL database in sequence to construct the multi-dimensional and multi-modal clinical database.
[0010] Preferably, the tumor diagnosis and treatment label includes: Treatment modality labels, including: neoadjuvant chemotherapy strategy and extent of surgical resection after chemotherapy and targeted drugs; Prognostic signatures include: the range of downstaging of various clinical indicators after chemotherapy, survival rate, and one-year local control rate.
[0011] In another aspect, a multimodal deep learning-based intelligent assisted analysis system for head and neck tumors is provided. The multimodal deep learning-based intelligent assisted analysis system for head and neck tumors is used to implement the multimodal deep learning-based intelligent assisted analysis method for head and neck tumors described above. The system comprises: A data acquisition module for collecting clinical examination data of head and neck cancer patients; a tumor pre-diagnosis module, configured to input the head and neck tumor clinical examination data into a pre-deployed multimodal head and neck tumor analysis model, identify the patient's head and neck tumor data features through the multimodal head and neck tumor analysis model, and output a tumor diagnosis and treatment label that matches the head and neck tumor data features; The report generation module is used to record the head and neck tumor data characteristics and tumor diagnosis and treatment labels, and write them into the HIS electronic medical record of the head and neck tumor patient to generate the corresponding head and neck tumor pre-diagnosis report.
[0012] On the other hand, an electronic device is provided, comprising: a processor; and a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, any one of the above-mentioned intelligent assisted analysis methods for head and neck tumors based on multimodal deep learning is implemented.
[0013] On the other hand, a computer-readable storage medium is provided, wherein the storage medium stores at least one instruction, and the at least one instruction is loaded and executed by a processor to implement any one of the above-mentioned intelligent assisted analysis methods for head and neck tumors based on multimodal deep learning.
[0014] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least: This invention provides an intelligent, assisted analysis method for head and neck tumors based on multimodal deep learning. By incorporating a pretrained segmentation model, pretraining a multimodal deep learning model, an efficient dataset construction and annotation scheme, and comprehensive model training and performance verification methods, a multimodal head and neck tumor analysis model is constructed and used for clinical head and neck tumor feature identification and predictive recommendation of diagnosis and treatment strategies. This method aims to address the current problems of insufficient recognition accuracy and low automation in intelligent, assisted analysis methods for head and neck tumors.
[0015] This invention enables the construction of a precise treatment guidance model for head and neck tumor identification. Based on the head and neck tumor segmentation model, a classification model for guiding precise treatment can be established. This model achieves high segmentation accuracy and stability on various types of head and neck tumor imaging data. Furthermore, the multimodal head and neck tumor analysis model undergoes pre-training and multimodal, independent training, enabling multiple correction training to achieve high prediction accuracy and effectively distinguish between different treatments and prognoses. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0017] Figure 1 This is a flow chart of a method for intelligent assisted analysis of head and neck tumors based on multimodal deep learning provided by an embodiment of the present invention; Figure 2 This is a logical diagram of constructing a multimodal head and neck tumor analysis model provided by an embodiment of the present invention; Figure 3 This is a block diagram of a head and neck tumor intelligent assisted analysis system based on multimodal deep learning provided by an embodiment of the present invention; Figure 4 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0018] The technical solution of the present invention is described below in conjunction with the accompanying drawings.
[0019] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as an "exemplary" in the present invention should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of the word "exemplary" is intended to present concepts in a concrete manner. Furthermore, in the embodiments of the present invention, "and / or" can mean both or either of the two.
[0020] In the embodiments of the present invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, when the distinction is not emphasized, the meanings they convey are the same. The terms "of," "corresponding," and "corresponding" may sometimes be used interchangeably. It should be noted that, when the distinction is not emphasized, the meanings they convey are the same.
[0021] In the embodiments of the present invention, sometimes a subscript such as W1 may be mistakenly written as a non-subscript form such as W1. When the difference is not emphasized, the meanings to be expressed are the same.
[0022] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0023] The embodiment of the present invention provides a method for intelligent assisted analysis of head and neck tumors based on multimodal deep learning. The method can be implemented by an electronic device, which can be a terminal or a server. Figure 1 The flowchart of the intelligent assisted analysis method for head and neck tumors based on multimodal deep learning is shown. The processing flow of the method may include the following steps: S1. Collect clinical examination data of head and neck cancer patients; S2. Inputting the head and neck tumor clinical examination data into a pre-deployed multimodal head and neck tumor analysis model, identifying the patient's head and neck tumor data features through the multimodal head and neck tumor analysis model, and outputting a tumor diagnosis and treatment label that matches the head and neck tumor data features; S3. Record the head and neck tumor data characteristics and tumor diagnosis and treatment labels, and write them into the HIS electronic medical record of the head and neck tumor patient to generate a corresponding head and neck tumor pre-diagnosis report.
[0024] Clinical data of cancer patients, such as CT images of head and neck thyroid tumors (with corresponding tumor annotation information), ultrasound, pathology, and basic patient information, can be collected in combination with clinical collection methods and will not be detailed here.
[0025] Each inspection device and system can transmit data to the backend HIS system and save it in the patient's electronic medical record file.
[0026] A multimodal head and neck tumor analysis model is deployed in the background. It can identify characteristics of a patient's clinical data from different dimensions, such as tumor size and morphology in tumor images, and the presence of pathological lesions in pathological data. Integrating big data technology, it predicts and outputs tumor diagnosis and treatment strategies that match the patient's clinical characteristics, as well as relevant post-treatment prognoses, such as survival rates.
[0027] The following will describe in detail the construction method of the multimodal head and neck tumor analysis model.
[0028] like Figure 2 As shown, preferably, the method for generating the multimodal head and neck tumor analysis model includes: Collect multi-dimensional and multimodal clinical data of several head and neck cancer patients and write them into the database to build a multi-dimensional and multimodal clinical database; Based on the Transformer segmentation model, pre-training learning of head and neck tumor data features was performed on a public head and neck image segmentation dataset (using the HeckTOR 2022 dataset), and a head and neck tumor segmentation model was constructed; wherein the head and neck tumor data features were annotated with the corresponding tumor diagnosis and treatment labels; Using the head and neck tumor data in the multi-dimensional and multi-modal clinical database, performing migration training on the pre-trained head and neck tumor segmentation model to construct a multi-modal head and neck tumor analysis model; Collect clinical data sets of head and neck patients to clinically validate the multimodal head and neck tumor analysis model: If the clinical performance test is qualified, the multimodal head and neck tumor analysis model will be put into clinical deployment and application; If the clinical performance test fails, the above steps are repeated to retrain and construct the multimodal head and neck tumor analysis model.
[0029] This paper first pre-trains the segmentation model: a pre-training scheme based on public head and neck tumor image segmentation datasets (such as the HECKTOR2022 dataset) is proposed to build image segmentation models. These models include fully convolutional networks (FCNs), U-net, DeepLab, and Mask R-CNN. By comparing the pre-training results, the best model is selected to ensure that the imaging data can be effectively integrated into the multimodal database.
[0030] Multimodal deep learning model pre-training: A comprehensive multimodal deep learning model pre-training approach is implemented using head and neck cancer data from the TCIA-HNSCC database (a multi-dimensional, multimodal clinical database) of clinical patients. This approach avoids overlap between the training and test sets of the segmentation model. By identifying potential laboratory or imaging indicators associated with efficacy or prognosis during the pre-training phase, this lays the foundation for subsequent data validation.
[0031] Efficient dataset construction and annotation scheme: Propose an efficient dataset construction and annotation method to ensure the integrity and accuracy of the multimodal data required for deep learning model training.
[0032] Model training and performance verification methods: A complete set of model training and performance verification methods are provided to ensure the robustness of the algorithm and the feasibility of practical application. Through rigorous performance verification, the reliability and effectiveness of the model in clinical application are ensured.
[0033] Among them, preferably, the head and neck image segmentation public dataset adopts the HeckTOR 2022 dataset.
[0034] Preferably, the multi-dimensional and multi-modal clinical data of several head and neck cancer patients are collected and written into a library to construct a multi-dimensional and multi-modal clinical database, including: Parse the HeckTOR 2022 dataset to obtain the data structure in the dataset; According to the data dimensions and modalities in the data structure of the HeckTOR 2022 dataset, clinical examination data of several head and neck cancer patients in the corresponding data dimensions and modalities are collected; The clinical examination data of each head and neck cancer patient in the corresponding data dimensions and modalities are written into the preset MySQL database in sequence to construct the multi-dimensional and multi-modal clinical database.
[0035] Combine Figure 2 The data structure shown in the figure, the present invention collects the modal data of each dimension according to the data dimension of the HeckTOR 2022 dataset. Specifically: 1) Data Collection Patient inclusion criteria: Patients diagnosed with head and neck malignant tumors by histopathological or cytological examination.
[0036] According to RECIST 1.1 tumor evaluation criteria, there was at least one measurable lesion.
[0037] No treatment (including surgery, radiotherapy, chemotherapy, or hormonal therapy) had been received for head and neck lesions before imaging examination.
[0038] The image quality was good, with no obvious motion artifacts or magnetic susceptibility artifacts.
[0039] The patients or their legal representatives voluntarily agreed to participate in the study and signed the informed consent form.
[0040] Exclusion criteria: There is no imaging and clinical treatment data, or the imaging data is modality missing.
[0041] The patient had other malignant tumors.
[0042] Inability to understand or give informed consent due to dementia, altered mental status, or any mental illness.
[0043] Those deemed unsuitable for inclusion by the researchers.
[0044] Image data collection: Sources: CT (using X-ray layered scanning and computer-reconstructed tomographic images), PET-CT (combining PET (positron emission tomography) and CT anatomical positioning to achieve dual imaging of function and structure), MRI (based on magnetic fields and radio frequency waves to stimulate hydrogen proton vibrations to generate high-resolution tomographic images), pathological images, medical history, treatment plan, survival, molecular pathology.
[0045] Methods: Multidimensional and multimodal clinical data were obtained from the above sources and a corresponding database was established.
[0046] 2) Data processing Image preprocessing: ROI delineation: Professional radiologists mark and evaluate the tumor area.
[0047] Resampling: Select the continuous image slices containing the largest tumor lesion area for resampling.
[0048] Resizing: Crop and pad the image to a uniform size of 160x160.
[0049] Normalization: Calculate the standard deviation (std) and mean (mean) for normalization.
[0050] Image enhancement: Segmentation tasks: random rotation, cropping, flipping, affine transformation, elastic transformation.
[0051] Classification tasks: random rotation, cropping, horizontal flipping, and translation.
[0052] Baseline clinical data collection: Baseline clinicopathological data were collected, including age, sex, BMI, tumor location, maximum diameter, tumor type, blood parameter levels before neoadjuvant chemotherapy, and degree of differentiation.
[0053] Label allocation: Treatment tags: Neoadjuvant chemotherapy regimens: No, Carbo (carboplatin) + Taxol (paclitaxel), Carboplatin + Taxol + Ifosfamide (carboplatin + paclitaxel + ifosfamide), Carboplatin + Taxol + Cetuximab (carboplatin + paclitaxel + cetuximab), Cisplatin + Docetaxel (cisplatin + docetaxel), Cisplatin + 5-FU + Docetaxel (cisplatin + fluorouracil + docetaxel); extent of surgical resection after neoadjuvant chemotherapy; targeted drug regimens; Prognostic Label: Downstaging situation; Survival status (OS (Overall Survival), PFS (Progression-Free Survival), EFS (Event-Free Survival), DFS (Disease-Free Survival); One-year local control rate.
[0054] The clinical data of each patient in each dimension and modality are collected to obtain a multi-dimensional and multi-modal clinical database. The details are as follows: Combine Figure 2 As shown, in order to implement the scheme, parse the HeckTOR 2022 dataset and build a multi-dimensional and multimodal clinical database, the following steps can be followed: 1. Analyze the data structure of the HeckTOR 2022 dataset Get the dataset: Download the HeckTOR 2022 dataset from the official channel or designated location.
[0055] Data decompression and viewing: Unzip the dataset file and view the included files and directory structure.
[0056] Typically, a dataset contains image files (such as PET / CT images), label files (such as tumor segmentation masks), and possible metadata files (such as patient information, scan parameters, etc.).
[0057] Analyze data structure: Determine the format of the image data (such as NIfTI, DICOM, etc.).
[0058] Understand the format and meaning of labeled data (e.g., annotation of tumor regions).
[0059] Review the metadata file to see what type of information is included (e.g., patient ID, age, gender, scan date, etc.).
[0060] Record data dimensions and modality: For image data, record its spatial dimensions (e.g., width, height, depth) and modality (PET, CT).
[0061] For label data, record its corresponding relationship with image data.
[0062] For metadata, record its fields and types.
[0063] 2. Collect clinical examination data of patients with head and neck cancer Determine data collection criteria: Based on the data dimensions and modalities of the HeckTOR 2022 dataset, determine the types of clinical examination data that need to be collected.
[0064] Ensure that the collected data is consistent with the HeckTOR dataset in terms of dimension and modality.
[0065] Data Collection: Clinical examination data of patients with head and neck cancer were obtained from hospitals or research institutions.
[0066] Ensure that the data contains PET / CT images, corresponding tumor segmentation information (if available), and patient metadata.
[0067] Data preprocessing: The collected data are preprocessed to ensure that they are consistent with the HeckTOR dataset format.
[0068] This may include image format conversion, label data alignment, metadata organization, etc.
[0069] 3. Build a multi-dimensional and multi-modal clinical database Design database structure: Design the structure of the MySQL database based on the collected data types and dimensions.
[0070] Create tables to store image data, label data, and patient metadata.
[0071] Define appropriate fields and data types for each table.
[0072] Data is written to the database: Write a script or program to write the collected clinical examination data of head and neck cancer patients into a MySQL database.
[0073] For image data, you can consider storing it as a file and storing the file path in the database, or directly encoding the image data into a binary format for storage.
[0074] Ensure data integrity and consistency during data writing.
[0075] Through the above steps, the data structure of the HeckTOR 2022 dataset can be successfully parsed, and the clinical examination data of head and neck cancer patients can be collected based on this data structure, and finally a multi-dimensional and multimodal clinical database can be constructed.
[0076] 3) Data Analysis Transfer training model: Based on the general Transformer segmentation model SAM, we performed migration training on the HeckTOR2022 dataset for a head and neck tumor segmentation model. We used different migration and freezing strategies to pre-train the head and neck tumor segmentation model.
[0077] Multimodal training and performance testing: The above model is trained unimodally on a multi-dimensional and multimodal clinical database.
[0078] A scheme for pre-training a holistic multimodal deep learning model using head and neck cancer data from a multi-dimensional and multimodal clinical database to avoid overlap between the training and test sets of the segmentation model.
[0079] Multimodal transfer training and performance testing are performed on clinical datasets.
[0080] Based on the Transformer segmentation model, we pre-trained the head and neck tumor data features on the HeckTOR 2022 dataset and built a head and neck tumor segmentation model. The head and neck tumor data features were annotated with the corresponding tumor diagnosis and treatment labels. Please refer to the following steps: To implement the solution of pre-training head and neck tumor data features on the HeckTOR 2022 dataset based on the Transformer segmentation model and building a head and neck tumor segmentation model, you can follow the following steps: 1. Data Preparation and Preprocessing Get the HeckTOR 2022 dataset: Download the HeckTOR 2022 dataset from the official channel and ensure the integrity and correctness of the data.
[0081] Data decompression and organization: Unzip the dataset and organize the image data (PET / CT), label data (tumor segmentation mask), and patient metadata into corresponding folders or data structures.
[0082] Data preprocessing: Perform necessary preprocessing on image data, such as normalization, resizing, data augmentation (such as rotation, flipping, cropping, etc.), etc., to increase data diversity and the generalization ability of the model.
[0083] Ensure that the label data corresponds one-to-one with the image data and convert it into a format that the model can understand.
[0084] 2. Building a Transformer Segmentation Model Choose Model Architecture: There are many Transformer-based segmentation models, such as Vision Transformer (ViT), SwinTransformer, and UNETR (a Transformer for medical images). Choose the appropriate model architecture based on task requirements and computing resources. Administrators can select the appropriate model for training.
[0085] Model Implementation: Implement the chosen Transformer segmentation model using a deep learning framework such as PyTorch, TensorFlow, etc.
[0086] Make sure the model can receive preprocessed image data as input and output segmentation maps of the same size as the input image.
[0087] Loss function and optimizer: Choose an appropriate loss function such as cross entropy loss, Dice loss, or a combination of them to measure the difference between the model prediction and the true label.
[0088] Select an optimizer, such as Adam, SGD, etc., to update the model parameters.
[0089] 3. Pre-training learning Data loading and batch processing: Write a data loader to load preprocessed image data and label data into the model in batches.
[0090] Make sure that each batch of data has the same size and format.
[0091] Model training: The Transformer segmentation model is pre-trained using the preprocessed HeckTOR 2022 dataset.
[0092] Set reasonable training parameters, such as learning rate, number of training rounds, batch size, etc.
[0093] During the training process, monitor the changes in loss function and evaluation indicators (such as Dice coefficient, IoU, etc.) to evaluate the performance of the model.
[0094] Model Validation and Adjustment: Use the validation set to validate the model and evaluate its performance on unseen data.
[0095] Adjust model parameters, optimizer settings, or data preprocessing methods based on the validation results to improve model performance.
[0096] 4. Building a Head and Neck Tumor Segmentation Model Model saving and loading: After pre-training is completed, save the model parameters and architecture for subsequent use.
[0097] When needed, the pre-trained model can be loaded and fine-tuned or further trained.
[0098] Model Deployment: Deploy the trained head and neck tumor segmentation model to an appropriate computing platform (usually the oncology department's backend system), such as a server, cloud, or local device.
[0099] Write an interface or application that allows doctors or researchers to easily input new head and neck tumor image data and obtain segmentation results.
[0100] Model Evaluation and Improvement: An independent test set is used to comprehensively evaluate the model, including segmentation accuracy, generalization ability, computational efficiency, etc.
[0101] Based on the evaluation results, continuously improve the model architecture, training strategy, or data preprocessing method to improve the performance and reliability of the model.
[0102] Through the above steps, we can pre-train the Transformer segmentation model on the HeckTOR 2022 dataset to learn the characteristics of head and neck tumor data, and build a head and neck tumor segmentation model. This model can automatically segment new head and neck tumor images and annotate the corresponding tumor areas, providing strong support for doctors' diagnosis and treatment.
[0103] Using the head and neck tumor data of several patients in the multi-dimensional and multi-modal clinical database, performing migration training on the pre-trained head and neck tumor segmentation model to construct a multi-modal head and neck tumor analysis model; Collect clinical data sets of head and neck patients to clinically validate the multimodal head and neck tumor analysis model: If the clinical performance test is qualified, the multimodal head and neck tumor analysis model will be put into clinical deployment and application; If the clinical performance test fails, the above steps are repeated to retrain and construct the multimodal head and neck tumor analysis model.
[0104] Specifically: To implement the solution of using head and neck tumor data from a multi-dimensional and multimodal clinical database to transfer training of the pre-trained head and neck tumor segmentation model, build a multimodal head and neck tumor analysis model, and perform clinical validation and deployment, you can follow the following steps: 1. Transfer training of a multimodal head and neck tumor analysis model Select Patient Data: Head and neck tumor data of several patients were selected from a multi-dimensional and multi-modal clinical database. These data should include PET / CT images, tumor segmentation labels, and possible diagnosis and treatment related information (such as pathological type, stage, etc.).
[0105] Data preprocessing: Perform any necessary preprocessing on the selected data to ensure it is consistent with the pre-trained model input requirements. This may include image resizing, normalization, data augmentation, etc.
[0106] Transfer Training: Load the pre-trained head and neck tumor segmentation model. Assume that the pre-trained head and neck tumor segmentation model is .
[0107] The model after transfer training is : , Added new task-related parameters (such as multimodal fusion layer); M is the frozen mask matrix (Mi=1 means freezing the i-th layer, Mi= means trainable); ⊙ represents element-by-element multiplication, which is used to control the parameter update range.
[0108] The model is transferred and trained using selected patient data, that is, the model parameters are fine-tuned on the new dataset so that it can better adapt to multimodal data and diagnosis-related tasks.
[0109] During the training process, monitor the changes in loss functions and evaluation indicators and adjust the training strategy in a timely manner.
[0110] Model Save: After the migration training is completed, save the obtained multimodal head and neck tumor analysis model.
[0111] 2. Clinical Validation Collecting clinical datasets: Collect new clinical datasets of head and neck patients from hospitals or research institutions. These datasets should include PET / CT images, diagnosis and treatment information, and possible pathological results.
[0112] Data preprocessing and annotation: Perform necessary preprocessing on the clinical dataset, such as image format conversion, size adjustment, etc.
[0113] If possible, professional doctors are invited to annotate the tumor areas in the images for subsequent performance evaluation.
[0114] Model Testing: Testing a multimodal head and neck cancer analysis model using a clinical dataset.
[0115] Evaluate the model's performance indicators such as segmentation accuracy, recognition accuracy of diagnosis and treatment-related information, and computational efficiency.
[0116] Performance Evaluation: Based on the test results, the model's clinical performance is evaluated to see if it meets the requirements. This may include comparing the results with manual segmentation and diagnosis by doctors, as well as calculating indicators such as sensitivity, specificity, and accuracy.
[0117] 3. Clinical deployment and application decisions Decision-making basis: If the clinical performance test is qualified, that is, the model performance meets or exceeds the preset standards, the model will be considered for clinical deployment and application.
[0118] If the clinical performance test fails, that is, the model performance does not meet expectations, the reasons need to be analyzed, which may be problems with data quality, model architecture, training strategy, etc.
[0119] Repeat training or improvement: If the model performance is unsatisfactory, repeat the above steps based on the analysis results, reselect the data, adjust the model architecture or training strategy, and retrain to build a multimodal head and neck tumor analysis model.
[0120] It is also possible to consider introducing more clinical data or adopting more advanced algorithms to improve model performance.
[0121] Clinical deployment: If the model performance is satisfactory and fully validated, the multimodal head and neck cancer analysis model can be deployed in a clinical setting.
[0122] Write user-friendly interfaces or applications that enable doctors to easily use the model to assist in diagnosis and treatment decision-making.
[0123] Continuous Monitoring and Updates: After clinical deployment, continuously monitor the performance and usage of the model.
[0124] Based on feedback and new clinical data, the model is regularly updated to maintain its performance and accuracy.
[0125] Through the above steps, we can use head and neck tumor data from a multi-dimensional and multimodal clinical database to transfer training to the pre-trained head and neck tumor segmentation model, build a multimodal head and neck tumor analysis model, and conduct clinical validation and deployment. This will help improve the diagnosis and treatment of head and neck tumors and provide patients with a better medical experience.
[0126] During transfer training, different transfer and freezing strategies can be used to pre-train the head and neck tumor segmentation model. For example: First, choose a deep learning model that has been pre-trained on a large dataset (such as a public dataset for medical image segmentation) as a basis. This model can be an architecture such as U-Net or Transformer. The specific choice depends on the complexity of the task and the characteristics of the data.
[0127] Prepare a head and neck tumor dataset for transfer learning. This dataset should contain multimodal image data (such as CT, MRI, etc.) and corresponding tumor segmentation labels.
[0128] Data preprocessing, including image resizing, normalization, data augmentation, etc., is performed to ensure data quality and diversity.
[0129] Freeze some layers: Based on the migration and freezing policy (determined by the administrator), you can choose to freeze some layers of the base model, especially those layers that extract low-level features, because these layers have learned common image features on large datasets.
[0130] Fine-tune the remaining layers: For the unfrozen layers, fine-tune them using the head and neck tumor dataset to suit the specific segmentation task. This typically involves adjusting hyperparameters such as learning rate and optimizer settings.
[0131] Training and Validation: The preprocessed data is fed into the model for training, and the validation set is used to monitor the performance of the model to prevent overfitting.
[0132] Adjust the model parameters and training strategy based on the validation results until satisfactory performance is achieved.
[0133] Testing and Evaluation: The performance of the model is evaluated on an independent test set, using metrics such as the Dice coefficient and IoU (Intersection over Union) to measure the accuracy of segmentation.
[0134] If the performance meets expectations, further clinical validation can be performed; otherwise, it is necessary to return to step 3 or earlier steps for adjustment and optimization.
[0135] Throughout the entire process, the choice of transfer and freezing strategies depends on the specific application scenario and data characteristics. For example, if the head and neck tumor dataset is small, more layers of the base model may need to be frozen to reduce the risk of overfitting; if the dataset is large and significantly different from the dataset used to train the base model, more layers may need to be fine-tuned to adapt to the new task.
[0136] Therefore, by incorporating a pre-trained segmentation model, pre-training a multimodal deep learning model, an efficient dataset construction and annotation scheme, and comprehensive model training and performance verification methods, this invention constructs a multimodal head and neck tumor analysis model for clinical head and neck tumor feature recognition and predictive recommendation of treatment strategies. This aims to address the current issues of insufficient recognition accuracy and low automation in intelligent-assisted analysis methods for head and neck tumors.
[0137] When training the model, you can also use consistency regularization and pseudo-labeling strategies to build a semi-supervised learning model.
[0138] For unlabeled data, two data augmentation strategies are used: weak augmentation and strong augmentation. Strong augmentation uses the RandomAugment method combined with CutOut for enhancement.
[0139] This invention enables the construction of a precise treatment guidance model for head and neck tumor identification. Based on the head and neck tumor segmentation model, a classification model for guiding precise treatment can be established. This model achieves high segmentation accuracy and stability on various types of head and neck tumor imaging data. Furthermore, the multimodal head and neck tumor analysis model undergoes pre-training and multimodal, independent training, enabling multiple correction training to achieve high prediction accuracy and effectively distinguish between different treatments and prognoses.
[0140] Figure 3 This is a block diagram of a head and neck tumor intelligent assisted analysis system based on multimodal deep learning according to an exemplary embodiment. The system is used for a head and neck tumor intelligent assisted analysis method based on multimodal deep learning. Figure 3 The system includes a data acquisition module 310, a tumor pre-diagnosis module 320, and a report generation module 330. Among them: The data collection module 310 is used to collect clinical examination data of head and neck cancer patients; a tumor pre-diagnosis module 320 for inputting the head and neck tumor clinical examination data into a pre-deployed multimodal head and neck tumor analysis model, identifying the patient's head and neck tumor data features through the multimodal head and neck tumor analysis model, and outputting a tumor diagnosis and treatment label that matches the head and neck tumor data features; The report generation module 330 is used to record the head and neck tumor data characteristics and tumor diagnosis and treatment labels, and write them into the HIS electronic medical record of the head and neck tumor patient to generate a corresponding head and neck tumor pre-diagnosis report.
[0141] Please understand the above modules and their interactions in conjunction with the corresponding steps in the above method, and will not be repeated here.
[0142] Figure 4is a schematic structural diagram of an electronic device provided by an embodiment of the present invention, such as Figure 4 As shown, the electronic device may include the above Figure 3 The electronic device 410 may include a first processor 2001 .
[0143] Optionally, the electronic device 410 may further include a memory 2002 and a transceiver 2003 .
[0144] The first processor 2001, the memory 2002 and the transceiver 2003 may be connected via a communication bus.
[0145] The following combination Figure 4 The components of the electronic device 410 are described in detail. The first processor 2001 is the control center of the electronic device 410 and can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 can be one or more central processing units (CPUs), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs) or one or more field programmable gate arrays (FPGAs).
[0146] Optionally, the first processor 2001 can execute various functions of the electronic device 410 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.
[0147] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 4 CPU0 and CPU1 are shown in FIG.
[0148] In a specific implementation, as an embodiment, the electronic device 410 may also include multiple processors, such as Figure 4 1 and 2. The first processor 2001 and the second processor 2004 are shown in FIG. Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). A processor herein can refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).
[0149] The memory 2002 is used to store the software program for executing the solution of the present invention, and is controlled by the first processor 2001 for execution. The specific implementation method can refer to the above method embodiment and will not be repeated here.
[0150] Alternatively, the memory 2002 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, an optical disc storage (including a compact disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001 or exist independently and accessed through the interface circuit ( Figure 4 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.
[0151] The transceiver 2003 is used to communicate with a network device or a terminal device.
[0152] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 4 The receiver is used to implement a receiving function, and the transmitter is used to implement a sending function.
[0153] Optionally, the transceiver 2003 may be integrated with the first processor 2001, or may exist independently and communicate with the first processor 2001 through the interface circuit ( Figure 4 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.
[0154] It should be noted that Figure 4 The structure of the electronic device 410 shown in the figure does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0155] In addition, the technical effects of the electronic device 410 can refer to the technical effects of the intelligent assisted analysis method for head and neck tumors based on multimodal deep learning described in the above method embodiment, and will not be repeated here.
[0156] It should be understood that the first processor 2001 in the embodiment of the present invention may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor, or the processor may be any conventional processor, etc.
[0157] It should also be understood that the memory in the embodiments of the present invention may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0158] The above embodiments can be implemented in whole or in part via software, hardware (e.g., circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product comprises one or more computer instructions or computer programs. When loaded or executed on a computer, the processes or functions described in accordance with the embodiments of the present invention are fully or partially performed. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable system. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired means (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0159] It should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " as used herein generally indicates an "or" relationship between the associated objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.
[0160] In this disclosure, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, "at least one of a, b, or c" can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.
[0161] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0162] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0163] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0164] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, systems, and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of the system or unit, which can be electrical, mechanical or other forms.
[0165] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0166] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0167] If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical disks.
[0168] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. An intelligent assisted analysis method for head and neck tumors based on multimodal deep learning, characterized in that: The method comprises: S1. Collect clinical examination data of head and neck cancer patients; S2. Inputting the head and neck tumor clinical examination data into a pre-deployed multimodal head and neck tumor analysis model, identifying the patient's head and neck tumor data features through the multimodal head and neck tumor analysis model, and outputting a tumor diagnosis and treatment label that matches the head and neck tumor data features; S3. Record the head and neck tumor data characteristics and tumor diagnosis and treatment labels, and write them into the HIS electronic medical record of the head and neck tumor patient to generate a corresponding head and neck tumor pre-diagnosis report.
2. The intelligent assisted analysis method for head and neck tumors based on multimodal deep learning according to claim 1 is characterized in that: The method for generating the multimodal head and neck tumor analysis model comprises: Collect multi-dimensional and multimodal clinical data of several head and neck cancer patients and write them into the database to build a multi-dimensional and multimodal clinical database; Based on the Transformer segmentation model, pre-training learning of head and neck tumor data features is performed on a public head and neck image segmentation dataset, and a head and neck tumor segmentation model is constructed; wherein the head and neck tumor data features are annotated with tumor diagnosis and treatment labels corresponding to the features; Using the head and neck tumor data in the multi-dimensional and multi-modal clinical database, performing migration training on the pre-trained head and neck tumor segmentation model to construct a multi-modal head and neck tumor analysis model; Collect clinical data sets of head and neck patients to clinically validate the multimodal head and neck tumor analysis model: If the clinical performance test is qualified, the multimodal head and neck tumor analysis model will be put into clinical deployment and application; If the clinical performance test fails, the above steps are repeated to retrain and construct the multimodal head and neck tumor analysis model.
3. The intelligent assisted analysis method for head and neck tumors based on multimodal deep learning according to claim 2 is characterized in that: The head and neck image segmentation public dataset adopts the HeckTOR 2022 dataset.
4. The intelligent assisted analysis method for head and neck tumors based on multimodal deep learning according to claim 3 is characterized in that: The multi-dimensional and multi-modal clinical data of several patients with head and neck cancer are collected and written into a database to construct a multi-dimensional and multi-modal clinical database, including: Parse the HeckTOR 2022 dataset to obtain the data structure in the dataset; According to the data dimensions and modalities in the data structure of the HeckTOR 2022 dataset, clinical examination data of several head and neck cancer patients in the corresponding data dimensions and modalities are collected; The clinical examination data of each head and neck cancer patient in the corresponding data dimensions and modalities are written into the preset MySQL database in sequence to construct the multi-dimensional and multi-modal clinical database.
5. The intelligent assisted analysis method for head and neck tumors based on multimodal deep learning according to claim 4 is characterized in that: The tumor diagnosis and treatment label includes: Treatment modality labels, including: neoadjuvant chemotherapy strategy and extent of surgical resection after chemotherapy and targeted drugs; Prognostic signatures include: the range of downstaging of various clinical indicators after chemotherapy, survival rate, and one-year local control rate.
6. A multimodal deep learning-based intelligent assisted analysis system for head and neck tumors, wherein the multimodal deep learning-based intelligent assisted analysis system for head and neck tumors is used to implement the multimodal deep learning-based intelligent assisted analysis method for head and neck tumors as described in any one of claims 1 to 5, characterized in that: The system comprises: A data acquisition module for collecting clinical examination data of head and neck cancer patients; a tumor pre-diagnosis module, configured to input the head and neck tumor clinical examination data into a pre-deployed multimodal head and neck tumor analysis model, identify the patient's head and neck tumor data features through the multimodal head and neck tumor analysis model, and output a tumor diagnosis and treatment label that matches the head and neck tumor data features; The report generation module is used to record the head and neck tumor data characteristics and tumor diagnosis and treatment labels, and write them into the HIS electronic medical record of the head and neck tumor patient to generate the corresponding head and neck tumor pre-diagnosis report.
7. An electronic device, characterized in that: The electronic device comprises: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program code, which can be called by a processor to execute the method according to any one of claims 1 to 5.