A radiotherapy plan decision method and system based on a tumor classification annotation database
By establishing a tumor classification and annotation database, the problems of fragmented storage and lack of unified annotation for spinal tumor radiotherapy data have been solved, enabling precise radiotherapy plan recommendations and efficacy prediction, and promoting the transformation of spinal tumor radiotherapy towards data-driven precision diagnosis and treatment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PEKING UNIVERSITY THIRD HOSPITAL (THE THIRD CLINICAL MEDICAL SCHOOL OF PEKING UNIVERSITY)
- Filing Date
- 2025-11-20
- Publication Date
- 2026-07-21
AI Technical Summary
In the field of spinal tumor radiotherapy, there are problems with fragmented data storage and lack of unified labeling, which makes it difficult to use data efficiently. Clinicians rely on personal experience to formulate radiotherapy plans without data support.
A tumor classification and annotation database was established. By collecting and annotating historical spinal tumor radiotherapy data, multimodal fusion data was constructed. The medical terminology mapping matrix was used to associate the data with a pre-set data dictionary, and the information gain value was calculated to determine the radiotherapy plan.
This has enabled precise recommendations and efficacy predictions for spinal tumor radiotherapy plans, promoted the transformation of spinal tumor radiotherapy towards data-driven precision diagnosis and treatment, and enhanced the scientific nature and clinical reference value of decision-making.
Smart Images

Figure CN121171483B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of tumor treatment technology, and in particular to a method and system for radiotherapy planning decisions based on a tumor classification and annotation database. Background Technology
[0002] Spinal tumors are a type of disease with high clinical difficulty in diagnosis and treatment. Their treatment plans are complex, especially radiotherapy, which requires precise decision-making based on a comprehensive assessment of the patient's baseline condition, tumor pathological characteristics, imaging findings, radiotherapy plan details, and long-term follow-up results. However, the current field of spinal tumor radiotherapy faces significant bottlenecks in data management and utilization. Relevant data is generally stored in a fragmented manner, lacks unified labeling, and has inconsistent data standards across different systems, directly leading to inefficient data utilization. Massive amounts of clinical data cannot be systematically aggregated, analyzed, and transformed into knowledge that supports decision-making. Clinicians still heavily rely on personal experience to formulate radiotherapy plans, resulting in decisions lacking data support. Summary of the Invention
[0003] In view of the above, the present invention aims to provide a radiotherapy plan decision-making method and system based on a tumor classification and annotation database to solve the aforementioned technical problems.
[0004] The technical solution adopted in this invention is as follows:
[0005] This invention provides a radiotherapy protocol decision-making method based on a tumor classification and annotation database, including:
[0006] Collect the patient's historical spinal tumor radiotherapy data;
[0007] The historical spinal tumor radiotherapy data were annotated to obtain a classification and annotation database;
[0008] Historical text data, historical medical image data, and historical multi-source time-series data from multiple systems within the hospital were collected to obtain multimodal fusion data;
[0009] The multimodal fusion data is mapped to the classification and labeling database to obtain historical diagnosis and treatment data;
[0010] Collect current case data;
[0011] The current case data and the historical treatment data are matched to obtain the radiotherapy plan decision results.
[0012] Optionally, the historical spinal tumor radiotherapy data may be annotated, including:
[0013] Patient baseline information annotation, tumor core feature annotation, radiotherapy planning and execution annotation, efficacy and follow-up annotation.
[0014] Optionally, historical text data, historical medical image data, and historical multi-source time-series data from multiple systems within the hospital can be collected to obtain multimodal fusion data, including:
[0015] The historical text data is cleaned and key terms are extracted to obtain standard historical text data.
[0016] Extract the feature vectors from the historical medical image data to obtain image feature data;
[0017] The historical multi-source time-series data is time-series aligned to obtain time-series feature data;
[0018] The standard historical text data, image feature data, and time-series feature data are fused to obtain multimodal fusion data.
[0019] Optionally, the multimodal fusion data is mapped to the classification and labeling database to obtain historical medical data, including:
[0020] Construct a medical terminology mapping matrix;
[0021] The mapping words with the highest similarity are determined by comparing the medical terminology mapping matrix with a pre-set data dictionary;
[0022] Based on the mapping words with the highest similarity, data association rules are obtained;
[0023] Based on the data association rules, the multimodal fusion data is mapped to the classification and labeling database to obtain historical diagnosis and treatment data.
[0024] Optionally, the medical terminology mapping matrix is:
[0025] ;
[0026] in, is a medical terminology mapping matrix; i represents terms in the multimodal fusion data; j represents standard terms in the data dictionary; Word vector representations of terms in multimodal fusion data; This refers to the word vector representations of standard terms in the data dictionary.
[0027] Optionally, the current case data and the historical treatment data are matched to obtain the radiotherapy plan decision results, including:
[0028] The current case data and the historical treatment data are matched to obtain the matching degree;
[0029] Based on the matching degree, calculate the information gain value of a preset number of historical medical data;
[0030] The historical treatment data corresponding to the maximum information gain is determined as the decision result for radiotherapy plan.
[0031] Optionally, the current case data and the historical treatment data are matched to obtain the matching degree, including:
[0032] via Sim , thus obtaining the matching degree;
[0033] Among them, Sim For the matching degree, P is the feature vector of the current case data, and Q is the feature vector of the historical treatment data.
[0034] This invention also provides a radiotherapy protocol decision-making system based on a tumor classification and annotation database, comprising:
[0035] The classification and annotation database module is used to annotate and store historical spinal tumor radiotherapy data;
[0036] The intelligent preprocessing engine module is used to collect historical text data, historical medical image data, and historical multi-source time-series data from multiple systems within the hospital to obtain multimodal fusion data;
[0037] The multi-source data fusion module is used to map multimodal fused data with a classification and annotation database to obtain historical diagnosis and treatment data;
[0038] The decision support module is used to match current case data with historical treatment data to obtain radiotherapy plan decision results.
[0039] Optionally, the intelligent preprocessing engine module includes:
[0040] The text cleaning unit is used to clean historical text data and extract key terms to obtain standard historical text data.
[0041] The image preprocessing unit is used to extract feature vectors from historical medical image data to obtain image feature data.
[0042] The time-series alignment unit is used to perform time-series alignment on historical multi-source time-series data to obtain time-series feature data.
[0043] The fusion module is used to fuse standard historical text data, image feature data, and time-series feature data to obtain multimodal fused data.
[0044] Optionally, the radiotherapy protocol decision-making system based on the tumor classification and annotation database may also include: a visualization platform for displaying the radiotherapy protocol decision results in a multi-dimensional visualization format.
[0045] The above-described solution of the present invention has at least the following beneficial effects:
[0046] The above-mentioned solution of the present invention involves: collecting historical spinal tumor radiotherapy data from patients; annotating the historical spinal tumor radiotherapy data to obtain a classification and annotation database; collecting historical text data, historical medical image data, and historical multi-source time-series data from multiple systems within the hospital to obtain multimodal fusion data; mapping the multimodal fusion data to the classification and annotation database to obtain historical treatment data; collecting current case data; and matching the current case data with the historical treatment data to obtain radiotherapy plan decision results. The solution of the present invention is a radiotherapy plan decision-making method based on a tumor classification and annotation database. Addressing the current situation of "fragmented storage, lack of unified annotation, and difficulty in efficient utilization" of spinal tumor radiotherapy data, it establishes a scientific classification and annotation system. Through multi-source data fusion, intelligent preprocessing, and adaptive visualization algorithms, it achieves accurate recommendation and efficacy prediction of spinal tumor radiotherapy plans, promoting the transformation of spinal tumor radiotherapy towards "data-driven precision treatment." This effectively solves the problems of insufficient utilization of existing spinal tumor radiotherapy data and lack of data support for decision-making. Attached Figure Description
[0047] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described below with reference to the accompanying drawings, wherein:
[0048] Figure 1 A flowchart is provided for a radiotherapy plan decision-making method based on a tumor classification annotation database, which is an embodiment of the present invention. Detailed Implementation
[0049] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0050] This invention proposes an embodiment of a radiotherapy protocol decision-making method based on a tumor classification annotation database. Specifically, as follows: Figure 1 As shown, it includes:
[0051] Step 11: Collect the patient's historical spinal tumor radiotherapy data;
[0052] Step 12: Annotate the historical spinal tumor radiotherapy data to obtain a classification and annotation database;
[0053] This embodiment is applied to spinal tumor radiotherapy. In this step, the spinal tumor radiotherapy data is annotated according to the entire process of "patient-tumor-radiotherapy-follow-up," covering patient baseline information, core tumor features, radiotherapy planning and execution information, and efficacy and follow-up data. The annotation dimensions include patient baseline information annotation, core tumor feature annotation, radiotherapy planning and execution annotation, and efficacy and follow-up annotation.
[0054] Specifically, patient baseline information is labeled, including basic attributes (unique patient ID, gender, age, etc.) and spinal function baseline (JOA score, VAS score).
[0055] The core tumor feature annotation includes pathological type annotation, staging annotation (WHO classification standards, Enneking staging, etc.), and imaging feature annotation (imaging modality annotation, tumor imaging parameter annotation). Specifically, pathological type annotation distinguishes between primary spinal tumors (such as chordoma, giant cell tumor of bone, Ewing sarcoma, etc.) and metastatic spinal tumors according to the WHO classification standards for bone and soft tissue tumors (indicating the primary tumor site, such as lung cancer, breast cancer, prostate cancer, etc.), and records the basis for pathological diagnosis (puncture biopsy / postoperative pathology). Staging annotation uses the Enneking surgical staging system for primary tumors and the Tomita and Tokuhashi scores to indicate tumor burden and prognostic risk levels for metastatic tumors. It also indicates the spinal segment involved (e.g., C3, T12-L2, distinguishing specific locations such as vertebral body, pedicle, transverse process, etc.) and whether the spinal cord / nerve roots are invaded (indicating the source of imaging evidence, such as MRI / T2WI enhanced sequences). Image modality annotation: Differentiate between CT (plain / enhanced), MRI (T1WI, T2WI, DWI, enhanced sequences), PET-CT (annotating SUVmax value), and full-length spinal X-ray films, recording the examination time and equipment model. Tumor image parameter annotation: Manually or semi-automatically annotate the maximum diameter and volume of the tumor (measured using 3D image segmentation tools), and whether it is accompanied by vertebral compression fracture / spinal cord edema / epidural hematoma. The annotated data must be correlated with the corresponding image slice thickness and coordinate information.
[0056] The radiotherapy planning and execution annotations include the radiotherapy technique type, target volume and organ-at-risk parameters, and treatment execution information. Specifically, the radiotherapy technique type annotation distinguishes between intensity-modulated radiotherapy (IMRT), stereotactic body radiotherapy (SBRT), proton and heavy ion radiotherapy, and indicates the brand and model of the radiotherapy equipment. Precise radiotherapy parameter annotations include: Target volume annotation: According to ICRU Reporting Standard 83, distinguishing between gross tumor target volume (GTV), clinical target volume (CTV), and planning target volume (PTV), recording the volume, prescribed dose (e.g., GTV 60Gy / 30f), dose homogeneity index (HI), and conformity index (CI) for each target volume. Organ-at-risk (OAR) annotation: Annotating critical organs such as the spinal cord, nerve roots, esophagus, trachea, lung tissue, and heart, recording the dose limitation parameters for each OAR (e.g., maximum spinal cord dose ≤45Gy, lung V20 ≤30%) and the actual dose received. Treatment execution annotations: Recording the total number of radiotherapy sessions, the duration of each session, and whether treatment interruptions occurred (indicating the reason for interruption and the re-irradiation plan).
[0057] Efficacy and Follow-up Labeling: This includes short-term efficacy labeling (RECIST 1.1 criteria) and long-term follow-up labeling (recurrence status, survival status). Short-term efficacy labeling: 1-3 months after radiotherapy, labeling radiographic efficacy (differentiated according to RECIST 1.1 criteria: complete remission, partial remission, stable disease, progression), pain relief (VAS score changes), and adverse reactions (such as radiation esophagitis, myelosuppression, labeled using CTCAE 5.0 grading). Long-term follow-up labeling: Follow-up every 6-12 months, labeling tumor recurrence (labeling recurrence site and radiographic evidence), spinal stability (whether new fractures occur, internal fixation loosening), and patient survival status (disease-free survival / survival with tumor / death, recording the cause of death). The longest recommended follow-up period is ≥5 years.
[0058] By constructing a four-dimensional annotation system of "patient-tumor-radiotherapy-follow-up", the limitations of the existing technology's "general annotation dimension" are broken through, and the "annotated data and radiotherapy decision-making needs" are accurately matched.
[0059] Step 13: Collect historical text data, historical medical image data, and historical multi-source time-series data from multiple systems within the hospital to obtain multimodal fusion data;
[0060] In this step, historical text data, historical medical image data, and historical multi-source time-series data from multiple systems within the hospital are collected. The data is then cleaned, standardized, and feature extracted to achieve data fusion and obtain multimodal fused data. Specifically, a BERT pre-trained model is used for entity recognition of unstructured text (such as pathology reports). Term similarity calculation and the Levenshtein distance algorithm are used for terminology standardization and typo correction. Finally, the TF-IDF algorithm is used to extract key terms, ensuring the standardization and validity of the text data.
[0061] Step 131: Construct a semantic knowledge base containing 150,000 medical terms, establish a synonym mapping table, and define a term similarity calculation function:
[0062] ;
[0063] in, For terminology similarity, , This is a set of characteristic words for two medical terms.
[0064] The Levenshtein distance algorithm is used to identify and correct typos in text. The calculation formula is as follows:
[0065] ;
[0066] =min ;
[0067] ;
[0068] when and The function identifies 'a' as a typo for 'b' and replaces it. Here, 'a' and 'b' are characters in the text.
[0069] The TF-IDF algorithm is used to calculate term importance, and the formula is as follows:
[0070] TF(t, d) = Number of times term t appears in document d / Total number of terms in document d;
[0071] IDF(t) = log(total number of documents) / (number of documents containing term t + 1);
[0072] TF-IDF(t, d) = TF(t, d) ×IDF(t);
[0073] Retain the top 80% of key terms by TF-IDF value.
[0074] Step 132: Extract feature vectors from medical images using a convolutional neural network, perform distortion correction by combining imaging equipment parameters, and simultaneously use a 3D image segmentation tool to measure parameters such as the maximum diameter and volume of the tumor to provide image feature data for subsequent analysis.
[0075] Step 133: Based on the timestamp, the dynamic time warping algorithm is used to align the multi-source time series data (such as VAS scores from multiple follow-ups). For missing values, an LSTM model based on the attention mechanism is used for prediction to ensure the integrity and consistency of the time series data.
[0076] Alignment of multi-source time series data using a dynamic time warping algorithm based on timestamps, with specific steps including:
[0077] Given two time series X= Y= Calculate the distance matrix D, where D = ;
[0078] Construct a cumulative distance matrix C, where D ;
[0079] C ;
[0080] From C Backtrack to find the optimal path and complete the timing alignment;
[0081] For missing values, an attention-based LSTM model is used for prediction. The attention weights are calculated as follows:
[0082] ;
[0083] in, For LSTM hidden states, For query vector, Let k be the LSTM hidden state at the k-th time step in the sequence. .
[0084] Step 14: Map the multimodal fusion data to the classification and labeling database to obtain historical diagnosis and treatment data;
[0085] In this step, a medical terminology mapping matrix is constructed; the mapping words with the highest similarity are determined by comparing the medical terminology mapping matrix with a preset data dictionary; data association rules are obtained based on the mapping words with the highest similarity; and multimodal fusion data is mapped to a classification and labeling database based on the data association rules to obtain historical diagnosis and treatment data.
[0086] Specifically, construct a medical terminology mapping matrix:
[0087] ;
[0088] in, is a medical terminology mapping matrix; i represents terms in the multimodal fusion data; j represents standard terms in the data dictionary; Word vector representations of terms in multimodal fusion data; This refers to the word vector representations of standard terms in the data dictionary.
[0089] For each term, select the top 3 candidate mapping words with the highest similarity to form a mapping candidate set;
[0090] Here, medical experts can review and confirm the final mapping relationship, establishing a cross-system data association rule base. Mapping relationships are established by calculating semantic similarity of terms, ensuring that data from different systems can be accurately associated with the classification and labeling database.
[0091] Step 15: Collect current case data;
[0092] Step 16: Match the current case data with the historical treatment data to obtain the radiotherapy plan decision result.
[0093] In this step, the current case data and historical treatment data are matched to obtain the matching degree; based on the matching degree, the information gain value of a preset number of historical treatment data is calculated.
[0094] Specifically, a cosine similarity matching algorithm is used to match current case data with historical treatment data:
[0095] via Sim , thus obtaining the matching degree;
[0096] Among them, Sim For the matching degree, P is the feature vector of the current case data, and Q is the feature vector of the historical treatment data.
[0097] The top 5 historical cases with the highest similarity were selected, and a treatment decision tree was generated using the C4.5 algorithm. The information gain was calculated as follows:
[0098] ;
[0099] ;
[0100] ;
[0101] Where D is the dataset; This represents the proportion of samples of class i in the dataset; Entropy of the original dataset represents the degree of disorder before the data was partitioned; The weighted entropy after partitioning according to feature A represents the average disorder level of each subset; This is the information gain value.
[0102] The historical treatment data corresponding to the maximum information gain is used as the decision result for radiotherapy planning. This embodiment overcomes the limitations of the existing technology of "simple similarity ranking" and realizes the selection of the optimal plan from similar plans. By quantifying the "efficacy value of historical plans" into a decision indicator through "information gain", it avoids the drawbacks of "solely relying on similarity", ensuring that the recommended radiotherapy plan is not only "similar" but also "optimal", significantly improving the scientific nature and clinical reference value of decision-making, realizing accurate recommendation and efficacy prediction of spinal tumor radiotherapy plans, and promoting the transformation of spinal tumor radiotherapy towards "data-driven precision diagnosis and treatment".
[0103] This embodiment presents a radiotherapy planning decision-making method based on a tumor classification and annotation database. Addressing the current situation of fragmented storage, lack of unified annotation, and difficulty in efficient utilization of spinal tumor radiotherapy data, it establishes a scientific classification and annotation system. Through multi-source data fusion, intelligent preprocessing, and adaptive visualization algorithms, it achieves accurate recommendation of spinal tumor radiotherapy plans and prediction of efficacy, promoting the transformation of spinal tumor radiotherapy towards "data-driven precision diagnosis and treatment." This effectively solves the problems of insufficient utilization of existing spinal tumor radiotherapy data and lack of data support for decision-making.
[0104] This embodiment also provides a radiotherapy planning decision system based on a tumor classification and annotation database, including a classification and annotation database module, a multi-source data fusion module, an intelligent preprocessing engine, an auxiliary decision-making module, and a visualization platform.
[0105] The classification and annotation database module is used to annotate and store historical spinal tumor radiotherapy data;
[0106] The intelligent preprocessing engine module is used to collect historical text data, historical medical image data, and historical multi-source time-series data from multiple systems within the hospital to obtain multimodal fusion data;
[0107] The multi-source data fusion module is used to map multimodal fused data with a classification and annotation database to obtain historical diagnosis and treatment data;
[0108] The decision support module is used to match current case data with historical treatment data to obtain radiotherapy plan decision results.
[0109] A visualization platform is used to present the results of radiotherapy plan decisions in a multi-dimensional and visual format.
[0110] In this embodiment, the classification and annotation database module is used to store spinal tumor radiotherapy data annotated according to the entire process of "patient-tumor-radiotherapy-follow-up". Specifically, it includes: patient baseline information annotation: covering basic attributes (unique patient ID, gender, age, etc.) and spinal function baseline (JOA score, VAS score); tumor core feature annotation: including pathology and staging (according to WHO classification standards, Enneking staging, etc.) and imaging features (imaging modality, tumor imaging parameters); radiotherapy planning and execution annotation: including radiotherapy technology type, target area and organ at risk parameters, and treatment execution information; efficacy and follow-up annotation: including short-term efficacy (RECIST 1.1 criteria) and long-term follow-up (recurrence status, survival status).
[0111] The intelligent preprocessing engine module includes a text cleaning unit, an image preprocessing unit, and a temporal alignment unit. These units are used to clean, standardize, and extract features from the data. Specifically, the text cleaning unit uses a BERT pre-trained model to perform entity recognition on unstructured text (such as pathology reports), standardizes terms and corrects typos through term similarity calculation and the Levenshtein distance algorithm, and then uses the TF-IDF algorithm to extract key terms, ensuring the standardization and validity of the text data. The image preprocessing unit extracts feature vectors from medical images using convolutional neural networks, performs distortion correction in conjunction with imaging equipment parameters, and uses 3D image segmentation tools to measure parameters such as the maximum diameter and volume of tumors, providing image feature data for subsequent analysis. The temporal alignment unit uses a dynamic time warping algorithm based on timestamps to align multi-source time-series data (such as VAS scores from multiple follow-ups), and uses an attention-based LSTM model to predict missing values, ensuring the integrity and consistency of the time-series data.
[0112] The multi-source data fusion module is used to construct a unified data dictionary based on multimodal fusion data (such as patient information from the HIS system and image data from the PACS system) and through semantic mapping to achieve cross-system data association. Its core is to establish mapping relationships by calculating the semantic similarity of terms to ensure that data from different systems can be accurately associated with the classification and labeling database.
[0113] The decision support module is used to match historical treatment data with current case characteristics using a cosine similarity matching algorithm, and generate a visualized treatment plan decision tree to provide intuitive radiotherapy plan recommendations for clinicians.
[0114] The visualization platform features dynamic dimension switching, an anomaly warning module, and a clinical pathway recommendation unit. The dynamic dimension switching function allows medical staff to switch between individual patient views and group statistical views; the anomaly warning module automatically triggers visual highlighting and audio prompts when analysis results exceed preset thresholds, displaying radiotherapy plan decision results in a multi-dimensional visual format.
[0115] It should be noted that this system is the same as the method described above. All implementation methods in the above method embodiments are applicable to the embodiments of this device and can achieve the same technical effect.
[0116] An embodiment of the present invention also provides a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described in the above embodiments. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.
[0117] In this embodiment of the invention, a computer-readable storage medium is also provided, storing instructions that, when executed on a computer, cause the computer to perform the method described in the above embodiments. All implementations of the methods described in the above embodiments are applicable to this embodiment and can achieve the same technical effect.
[0118] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0119] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0120] In the embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0121] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0122] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0123] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0124] Furthermore, it should be noted that in the apparatus and method of the present invention, it is obvious that the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered equivalent solutions of the present invention. Moreover, the steps performing the above series of processes can naturally be executed in the order described, but are not necessarily required to be executed in chronological order; some steps can be executed in parallel or independently of each other. Those skilled in the art will understand that all or any step or component of the method and apparatus of the present invention can be implemented in any computing device (including processors, storage media, etc.) or network of computing devices, in hardware, firmware, software, or a combination thereof. This is something that those skilled in the art can achieve by using their basic programming skills after reading the description of the present invention.
[0125] Therefore, the object of the present invention can also be achieved by running a program or a set of programs on any computing device. The computing device can be a known general-purpose device. Therefore, the object of the present invention can also be achieved simply by providing a program product containing program code implementing the method or apparatus. That is, such a program product also constitutes the present invention, and the storage medium storing such a program product also constitutes the present invention. Obviously, the storage medium can be any known storage medium or any storage medium developed in the future. It should also be noted that in the apparatus and method of the present invention, it is obvious that the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered equivalent to the present invention. Furthermore, the steps performing the above series of processes can naturally be performed in the order described, but are not necessarily required to be performed in chronological order. Some steps can be performed in parallel or independently of each other.
[0126] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A radiotherapy protocol decision-making method based on a tumor classification annotation database, characterized in that, include: Collect the patient's historical spinal tumor radiotherapy data; The historical spinal tumor radiotherapy data were annotated to obtain a classification and annotation database; Historical text data, historical medical image data, and historical multi-source time-series data from multiple systems within the hospital were collected to obtain multimodal fusion data; The multimodal fusion data is mapped to the classification and labeling database to obtain historical diagnosis and treatment data; Collect current case data; The current case data and the historical treatment data are matched to obtain the radiotherapy plan decision results; The historical spinal tumor radiotherapy data is annotated, including: Patient baseline information annotation, tumor core feature annotation, radiotherapy planning and execution annotation, efficacy and follow-up annotation; Specifically, the multimodal fusion data is mapped to the classification and labeling database to obtain historical medical data, including: Construct a medical terminology mapping matrix; The mapping words with the highest similarity are determined by comparing the medical terminology mapping matrix with a pre-set data dictionary; Based on the mapping words with the highest similarity, data association rules are obtained; Based on the data association rules, the multimodal fusion data is mapped to the classification and labeling database to obtain historical diagnosis and treatment data; The medical terminology mapping matrix is as follows: ; in, is a medical terminology mapping matrix; i represents terms in the multimodal fusion data; j represents standard terms in the data dictionary; These are word vectors for terms in multimodal fusion data. These are the word vectors for standard terms in the data dictionary; In this process, the mapping words for each term are sorted from high to low similarity, and the top 3 mapping words in the similarity ranking are selected to form a mapping candidate set. By reviewing the mapping candidate set, the data association rules are confirmed, and a cross-system data association rule library is established. The process of matching the current case data with the historical treatment data to obtain the radiotherapy plan decision results includes: The current case data and the historical treatment data are matched to obtain the matching degree; Based on the matching degree, calculate the information gain value of a preset number of historical medical data; The historical treatment data corresponding to the maximum information gain is determined as the decision result for radiotherapy plan; The matching of the current case data and the historical treatment data to obtain the matching degree includes: via Sim , thus obtaining the matching degree; Among them, Sim For matching degree, P is the feature vector of the current case data, and Q is the feature vector of the historical treatment data; The top 5 historical cases with the highest similarity were selected, and a treatment decision tree was generated using the C4.5 algorithm. The information gain was calculated as follows: ; ; ; Where D is the dataset; This represents the proportion of samples of class i in the dataset; Entropy of the original dataset represents the degree of disorder before the data was partitioned; The weighted entropy after partitioning according to feature A represents the average disorder level of each subset; This is the information gain value.
2. The radiotherapy plan decision-making method based on a tumor classification annotation database according to claim 1, characterized in that, Historical text data, historical medical image data, and historical multi-source time-series data were collected from multiple systems within the hospital to obtain multimodal fusion data, including: The historical text data is cleaned and key terms are extracted to obtain standard historical text data. Extract the feature vectors from the historical medical image data to obtain image feature data; The historical multi-source time-series data is time-series aligned to obtain time-series feature data; The standard historical text data, image feature data, and time-series feature data are fused to obtain multimodal fusion data.
3. A radiotherapy protocol decision-making system based on a tumor classification and annotation database, characterized in that, include: The classification and annotation database module is used to annotate and store historical spinal tumor radiotherapy data; The intelligent preprocessing engine module is used to collect historical text data, historical medical image data, and historical multi-source time-series data from multiple systems within the hospital to obtain multimodal fusion data; The multi-source data fusion module is used to map multimodal fused data with a classification and annotation database to obtain historical diagnosis and treatment data; The decision support module is used to match current case data with historical treatment data to obtain radiotherapy plan decision results; The historical spinal tumor radiotherapy data is annotated, including: Patient baseline information annotation, tumor core feature annotation, radiotherapy planning and execution annotation, efficacy and follow-up annotation; Specifically, the multimodal fusion data is mapped to the classification and labeling database to obtain historical medical data, including: Construct a medical terminology mapping matrix; The mapping words with the highest similarity are determined by comparing the medical terminology mapping matrix with a pre-set data dictionary; Based on the mapping words with the highest similarity, data association rules are obtained; Based on the data association rules, the multimodal fusion data is mapped to the classification and labeling database to obtain historical diagnosis and treatment data; The medical terminology mapping matrix is as follows: ; in, is a medical terminology mapping matrix; i represents terms in the multimodal fusion data; j represents standard terms in the data dictionary; These are word vectors for terms in multimodal fusion data. These are the word vectors for standard terms in the data dictionary; In this process, the candidate mapping words for each term are sorted from high to low similarity, and the top 3 candidate mapping words in the similarity ranking are selected to form a mapping candidate set; the final mapping relationship is confirmed through review, and a cross-system data association rule base is established. The process of matching the current case data with the historical treatment data to obtain the radiotherapy plan decision results includes: The current case data and the historical treatment data are matched to obtain the matching degree; Based on the matching degree, calculate the information gain value of a preset number of historical medical data; The historical treatment data corresponding to the maximum information gain is determined as the decision result for radiotherapy plan; The matching of the current case data and the historical treatment data to obtain the matching degree includes: via Sim , thus obtaining the matching degree; Among them, Sim For matching degree, P is the feature vector of the current case data, and Q is the feature vector of the historical treatment data; The top 5 historical cases with the highest similarity were selected, and a treatment decision tree was generated using the C4.5 algorithm. The information gain was calculated as follows: ; ; ; Where D is the dataset; This represents the proportion of samples of class i in the dataset; Entropy of the original dataset represents the degree of disorder before the data was partitioned; The weighted entropy after partitioning according to feature A represents the average disorder level of each subset; This is the information gain value.
4. The radiotherapy protocol decision-making system based on a tumor classification annotation database according to claim 3, characterized in that, The intelligent preprocessing engine module includes: The text cleaning unit is used to clean historical text data and extract key terms to obtain standard historical text data. The image preprocessing unit is used to extract feature vectors from historical medical image data to obtain image feature data. The time-series alignment unit is used to perform time-series alignment on historical multi-source time-series data to obtain time-series feature data. The fusion module is used to fuse standard historical text data, image feature data, and time-series feature data to obtain multimodal fused data.
5. The radiotherapy protocol decision-making system based on a tumor classification and annotation database according to claim 3, characterized in that, Also includes: A visualization platform is used to present the results of radiotherapy plan decisions in a multi-dimensional and visual format.