Tumor patient clinical test matching system and method based on large language model and OCR technology
By constructing a three-dimensional knowledge graph and combining a large language model with OCR technology into a multi-module system, the shortcomings of existing clinical trial matching models in terms of accuracy and applicability are solved, achieving precise matching of cancer patients and efficient generation of matching reports.
Patent Information
- Application Number
- CN202511044021.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-07-28
AI Technical Summary
Existing AI-based clinical trial matching models have shortcomings in accuracy and applicability. In particular, they are difficult to reflect individualized information when facing different hospitals and patient groups. Furthermore, the accuracy of information extraction is highly dependent on the quality and completeness of structured data, resulting in low matching accuracy.
A multi-module collaborative system based on a large language model and OCR technology is constructed. By building a three-dimensional knowledge graph of "tumor type-stage-treatment line number", clinical data is analyzed and structured data is generated. Combined with OCR processing of image data, multimodal analysis and rule verification are performed to generate a matching report containing confidence level and risk warning.
It improves the accuracy and efficiency of clinical trial matching for cancer patients, generates precise trial recruitment tools, can process unstructured text and image data, dynamically adjust inclusion and exclusion criteria, and provide visualized matching reports.
Smart Images

Figure CN120913728A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of medical data processing, in particular to a tumor patient clinical trial matching system and method based on a large language model and OCR technology. BACKGROUND
[0002] Existing AI-based clinical trial matching models mostly perform natural language processing on patient information and clinical trial information, extract key information from both ends, and then perform matching to select matching patients. However, the accuracy of information extraction is highly dependent on the quality and completeness of structured data, and any omission or error of information will directly affect the accuracy of matching. Especially for some complex inclusion and exclusion conditions, these conditions are often difficult to accurately correspond through existing matching methods. Therefore, the matching accuracy is not high, errors are prone to occur, and all potential inclusion or exclusion factors cannot be covered.
[0003] In addition, existing AI analysis methods mostly rely on structured data in medical record systems, and these data are often limited by the framework of hospital information systems and the standardization degree of medical records. This limits the applicability and generalization of AI models, especially when facing different hospitals and different patient groups, which may not fully reflect the individualized information of patients.
[0004] It should be noted that the information disclosed in the above background section is only used to strengthen the understanding of the background of the present disclosure, and therefore includes information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY
[0005] The present application aims to provide a tumor patient clinical trial matching system and method based on a large language model and OCR technology, which at least to some extent overcomes the problems existing in the prior art, and improves the efficiency and accuracy of tumor clinical trial matching through multi-module cooperation. A three-dimensional knowledge graph of "tumor type-staging-treatment line number" is constructed, and dynamic rules are upgraded. The large language model analyzes clinical data to generate structured data, and the OCR processes image data to generate supplementary information. After multi-modal analysis, rule verification and model optimization, a report containing confidence, basis and risk prompts is finally generated, providing a precise tool for trial recruitment.
[0006] Other characteristics and advantages of the present application will become apparent from the following detailed description, or will be learned partly through practice of the present application.
[0007] According to an aspect of the present application, a tumor patient clinical trial matching method based on a large language model and OCR technology is provided, comprising: obtaining clinical data of a tumor patient and clinical trial information, preprocessing based on a knowledge graph construction module, constructing a "tumor type-stage-treatment line number" three-dimensional knowledge graph, integrating NCCN / CSCO guidelines and real-time literature data, and upgrading static judgment conditions to dynamic correlation rules; based on a large language model preprocessing module, analyze the clinical data and trial information, process unstructured text, combine a time sequence neural network and a RECIST1.1 standard time axis analysis model to accurately stage, and generate structured clinical feature data through context correlation analysis; simultaneously process paper and image data through an OCR technology analysis module, extract key data and correct through double verification, and generate structured data supplement information; based on a multi-modal processing pipeline module, enhance the structured clinical feature data and supplement information, extract image quantitative indicators, analyze immunohistochemical results, and generate a comprehensive matching score; through a rule engine and semantic similarity calculation, realize item-by-item comparison of patient features and entry and exclusion conditions, and generate a preliminary matching result; based on a large language model optimization module, correct edge case misjudgments through context-aware multi-round reasoning, adjust the order in combination with clinical trial priority weight, and generate an optimized clinical trial matching list; input the structured clinical feature data, supplement information, preliminary matching result, and optimized clinical trial matching list into a clinical decision support module, visually display the matching full path, and generate a clinical trial matching report containing matching confidence, key basis, and risk prompts.
[0008] In another aspect of the present application, a tumor patient clinical trial matching device based on a large language model and an OCR technology comprises: an acquisition module configured to acquire clinical data of a tumor patient and clinical trial information, pre-process based on a knowledge graph construction module, construct a "tumor type-stage-treatment line number" three-dimensional knowledge graph, integrate NCCN / CSCO guidelines and real-time literature data, and upgrade static judgment conditions to dynamic correlation rules; and a processing module configured to parse the clinical data and the trial information based on a large language model preprocessing module, process unstructured text, accurately stage in combination with a time sequence neural network and a RECIST1.1 standard time axis analysis model, generate structured clinical feature data through context correlation analysis, process paper and image data through an OCR technology analysis module, extract key data and correct through double verification, and generate structured data supplement information; a multi-modal processing pipeline module is used to enhance the structured clinical feature data and the supplement information, extract image quantitative indicators, analyze immunohistochemical results, and generate a comprehensive matching score; a rule engine and semantic similarity calculation are used to realize item-by-item comparison of patient features and entry and exclusion conditions, and generate a preliminary matching result; a large language model optimization module is used to correct edge case misjudgments through context perception multi-round reasoning, adjust the order in combination with clinical trial priority weight, and generate an optimized clinical trial matching list; and the structured clinical feature data, the supplement information, the preliminary matching result, and the optimized clinical trial matching list are input into a clinical decision support module, a matching full path is visualized and displayed, and a clinical trial matching report containing matching confidence, key basis, and risk prompts is generated.
[0009] According to another aspect of the present application, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a second processor to implement the above-mentioned tumor patient clinical trial matching method based on a large language model and an OCR technology.
[0010] The tumor patient clinical trial matching system and method based on a large language model and an OCR technology provided by the present application improve the efficiency and accuracy of tumor patient clinical trial matching through multi-module cooperation. First, a "tumor type-stage-treatment line number" three-dimensional knowledge graph is constructed, guidelines and literature are integrated, and static conditions are upgraded to dynamic rules. A large language model is used to analyze clinical data and accurately stage, and structured data is generated; paper and image data are processed through an OCR technology, and supplement information is generated through double verification. A multi-modal module extracts image indicators, analyzes immunohistochemistry, and generates a matching score. A rule engine and semantic calculation are used to obtain a preliminary result, which is then optimized and sorted by a large language model, and finally a report containing confidence, basis, and risk prompts is generated through a clinical decision support module, providing a precise tool for trial recruitment.
[0011] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory and are not restrictive of the disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0012] Figure 1 A flowchart of a tumor patient clinical trial matching method based on a large language model and an OCR technology according to an embodiment of the present application is shown.
[0013] Figure 2 A structural schematic diagram of a tumor patient clinical trial matching device based on a large language model and an OCR technology according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0014] The preferred embodiments of the present application are described below in conjunction with the accompanying drawings, and it should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application.
[0015] The tumor patient clinical trial matching method based on a large language model and an OCR technology according to the exemplary embodiments of the present application is described below in conjunction with Figure 1 The present application also proposes a tumor patient clinical trial matching system and method based on a large language model and an OCR technology in one embodiment:
[0016] In S101, the clinical data and clinical trial information of a tumor patient are obtained, preprocessed based on a knowledge graph construction module, a three-dimensional knowledge graph of "tumor type-stage-treatment line number" is constructed, NCCN / CSCO guidelines and real-time literature data are integrated, and static judgment conditions are upgraded to dynamic association rules.
[0017] In one embodiment, the clinical data and clinical trial information of several tumor patients are collected. Taking one of the patients as an example, the tumor patient has a current medical history (diagnosed with endometrial cancer for more than 2 years, liver progression in 2023-3, and recheck liver progression on 2024-2-23); treatment history (2021-8-9, albumin paclitaxel + carboplatin→2022-2, sedili monoclonal antibody→2023-3, albumin paclitaxel + carboplatin→2023-9, bevacizumab→2023-11, bevacizumab + letrozole); symptoms (diabetic ketoacidosis in 2022-9, improved after symptomatic treatment, and currently insulin treatment); gene detection results are NGS (this to) show KRAS G12C mutation, MSS. The score / basic information is ECOG = 1; age = 30; sex = male; stage = advanced. The treatment line number is determined as first line (2021-8-9, albumin paclitaxel + carboplatin) and second line (2023-3, albumin paclitaxel + carboplatin), the current treatment is second line, and the next step is third line.
[0018] The clinical trial information covers the inclusion criteria (such as age range, ECOG requirements, tumor stage adaptation, genetic test (KRAS G12C, MSS) restrictions, etc.), exclusion criteria (such as acute phase of diabetic ketoacidosis, recent major surgery, etc.), trial purpose (such as late-stage endometrial cancer post-line treatment), indications (advanced metastatic endometrial cancer), etc. The core elements of the clinical data structure are extracted, and the medical record elements are: tumor type (endometrial cancer), stage (advanced), treatment line number (first line, second line, and third line), genetic test (KRAS G12C mutation, MSS), and key events (liver metastasis progression, diabetic ketoacidosis). The trial key conditions: need to match the inclusion stage (advanced), treatment line number (such as whether to open the third line after the failure of the second line), and genetic characteristics (KRAS G12C mutation, MSS adaptation to the trial), etc., lay a good foundation for knowledge graph association.
[0019] Based on actual medical record data, the "endometrial cancer-advanced-second / third line treatment" node is associated. The treatment plan in this scenario (such as bevacizumab + letrozole, etc. existing plan), genetic characteristics (KRAS G12C mutation, MSS corresponding treatment response evidence), and guideline adaptation (NCCN / CSCO guideline for advanced endometrial cancer post-line treatment recommendation) are associated.
[0020] Inclusion of NCCN / CSCO endometrial cancer guidelines, regarding advanced / metastatic cancer staging criteria, treatment line number definition (such as changing line after disease progression), and indication matching specification (such as inclusion adaptation of patients with specific genetic characteristics). Inclusion of real-time literature, KRAS G12C mutation, MSS phenotype in endometrial cancer treatment (such as immunotherapy, targeted therapy), and response research, supplement the basis for the "endometrial cancer-advanced-post-line treatment" node in the knowledge graph.
[0021] Dynamic inclusion and exclusion criteria, such as clinical trial "exclude patients with acute diabetic ketoacidosis", combined with "patient had ketoacidosis but was successfully treated", knowledge graph dynamically determines whether it is adapted by associating "ketoacidosis relief safety data"; for treatment line number, based on "postoperative / treatment progression time-line number mapping model", combined with patient "2023-3 due to liver progression from first line to second line", dynamically correct the line number determination, support the trial inclusion line number condition matching.
[0022] S102, based on the large language model preprocessing module, analyze clinical data and trial information, process unstructured text, and accurately stage by combining time sequence neural network and RECIST1.1 standard time axis analysis model, generate structured clinical feature data through context association analysis; simultaneously process paper and image data through OCR technology analysis module, extract key data and correct through double verification, generate structured data supplement information.
[0023] In one implementation, the large language model preprocessing module analyzes the unstructured text in the clinical data and test information, extracts the patient's condition description and treatment history, and combines the time series neural network and the RECIST1.1 standard time axis analysis model to accurately determine the patient's disease stage, generating preliminary staging feature data. The patient's medical record states: "Diagnosed with endometrial cancer in March 2022, underwent total hysterectomy + bilateral adnexectomy, postoperative pathology showed low differentiated adenocarcinoma, invasion of deep muscle layer, cervical stroma involvement, no lymph node metastasis (0 / 12). In May 2023, a review found a liver S5 metastatic lesion (2.3 cm in diameter), with no other distant metastasis."
[0024] The large language model extracts the condition description (initial diagnosis, postoperative recurrence) and treatment history (surgical resection), combines the time series neural network and the RECIST1.1 standard, and analyzes the time axis (initial diagnosis in March 2022 → liver metastasis in May 2023) to determine the initial stage as IB and the recurrence stage as IVB, generating preliminary staging feature data.
[0025] The preliminary staging feature data is analyzed in context to eliminate text ambiguity, map key clinical indicators to standardized features, and generate structured clinical feature data, which includes core medical record elements such as tumor type, treatment line number, stage, and ECOG score. The descriptions in the medical record, such as "6 cycles of postoperative adjuvant chemotherapy" and "current physical condition is good, can take care of daily activities," are analyzed in context to determine that "postoperative adjuvant chemotherapy" corresponds to the treatment line number (first line) and "physical condition is good" maps to an ECOG score of 0. The structured data content is as follows: tumor type: endometrial cancer (low differentiated adenocarcinoma); treatment line number: first line (postoperative adjuvant chemotherapy); stage: IVB (liver metastasis); ECOG score: 0.
[0026] Simultaneously, the OCR technology analysis module identifies text from paper and image data, extracts key values and terms from medical records and examination reports, and combines text box position information to perform content splicing and format arrangement, generating original recognition text. When performing OCR recognition on the patient's paper pathology report photo, the system first locates the key areas in the report (such as the "immunohistochemical results" and "gene detection results" content blocks below the title) using a multi-modal text recognition algorithm, and accurately extracts key values and terms such as "ER (+, 80%)" and "PR (+, 60%)" that are scattered in different text boxes, based on common formats such as tables and bulleted lists in pathology reports.
[0027] Due to the tight layout, blurred handwriting, or scanning deviation of paper reports, the OCR technology combines the position coordinate information of the text box (such as the horizontal axis and vertical axis coordinate difference) to determine the content relevance: if the vertical axis coordinate difference between the text boxes of "ER(+, 80%)" and "PR(+, 60%)" is less than a preset threshold (such as 5 pixels), it is determined that they are adjacent items, and are spliced in the original text order as "immunohistochemistry: ER(+, 80%), PR(+, 60%), p53(wild type), Ki-67(index 70%)"; for the "gene detection" part, "MSH2, MSH6 protein expression normal" and "no clear pathogenic mutation" are also spliced into a coherent text through position association, ensuring that the original recognized text fully reflects the report logic. This process not only preserves the original expression of each indicator in the pathology report, but also eliminates content breaks caused by scanning or layout through format arrangement, providing complete and coherent basic data for subsequent character-level and semantic-level checking.
[0028] The original recognized text is subjected to character-level and semantic-level double checking rules to identify and correct abnormal text, and the data format is standardized to generate structured data supplement information containing immunohistochemistry information and gene detection results. When checking the OCR recognition results at the character level and semantic level, the processing of "Ki-67(index 70%)" reflects the character-level checking logic: through the preset medical terminology dictionary comparison, the system finds that the word "index" is often simplified in the pathology report, and combined with the format of adjacent indicators (such as ER, PR) (only the core value is retained), it is determined that "index" is a redundant expression, which is a misreading of the original text layout during OCR recognition, and is corrected to "Ki-67(70%)", ensuring uniformity with other immunohistochemistry indicators and meeting the technical requirements of "character-level correction of misrecognized text" in Appendix 1.
[0029] For semantic analysis of "MSH2, MSH6 protein expression normal", the semantic-level checking rule is followed: the system calls the judgment standard of "microsatellite stability (MSS)" in the medical knowledge graph - when the DNA mismatch repair protein (such as MSH2, MSH6) is expressed normally, the phenotype is microsatellite stable. Through the association of this medical knowledge, the vague expression "protein expression normal" in the original text is accurately converted to the standardized term "microsatellite stable (MSS)", which not only preserves the core information, but also meets the standardization requirements of clinical trials for gene detection results. The final generated standardized supplement information, through format unification (such as removing redundant terms, using standard abbreviations) and semantic standardization (such as replacing with industry-standard terms), makes the immunohistochemistry and gene detection data meet the input requirements of the subsequent multi-modal processing module, providing accurate and standardized feature data for the calculation of the comprehensive matching score, ensuring the coherence with the subsequent matching process.
[0030] S103, based on the multi-modal processing pipeline module, the structured clinical feature data and supplementary information are enhanced, image quantitative indicators are extracted, and immunohistochemical results are analyzed to generate a comprehensive matching score.
[0031] In an embodiment, based on the multi-modal processing pipeline module, the structured clinical feature data and supplementary information are analyzed, the lesion size and density in the image data are extracted, the protein expression level in the immunohistochemical result is analyzed, and a unique identifier is assigned to each type of indicator. From the patient's abdominal enhanced CT image, the liver metastasis lesion characteristics are extracted: maximum diameter 2.3 cm (identifier IM-001), lesion density 35 HU (identifier IM-002), and no other distant metastasis is recorded. From the structured data supplementary information, the protein expression level is extracted: ER (+, 80%, identifier IH-001), PR (+, 60%, identifier IH-002), Ki-67 (70%, identifier IH-003), and microsatellite stability (MSS, identifier IH-004) is associated. The unique identifier ensures that the indicators can be traced back during subsequent data association and calculation.
[0032] The image quantitative indicators and immunohistochemical analysis results are standardized, the dimensional differences between different indicators are eliminated through data normalization, and standardized feature data are generated. The maximum diameter of the liver metastasis lesion (2.3 cm) is converted to a 0-1 interval value using min-max standardization: the metastasis lesion maximum diameter threshold in the clinical trial is 5 cm, and the standardized value is 2.3 / 5 = 0.46; the lesion density (35 HU) is referenced to the normal liver tissue density (50 HU), and the standardized value is 35 / 50 = 0.7.
[0033] The protein expression positive rate is directly normalized to a percentage value: ER (80%→0.8), PR (60%→0.6), and Ki-67 (70%→0.7); MSS is a classification variable, with a value of 1 (stable) and 0 (unstable). The standardized feature data are generated: IM-001 = 0.46, IM-002 = 0.7, IH-001 = 0.8, IH-002 = 0.6, IH-003 = 0.7, and IH-004 = 1.
[0034] The standardized feature data are associated with the structured clinical feature data and supplementary information, the correlation between the indicators is established, and the correspondence between each indicator and the patient's condition is determined. Specifically, the image indicator association is that the maximum diameter of the liver metastasis lesion (IM-001) is directly related to "advanced endometrial cancer liver metastasis", and the lesion density (IM-002) assists in judging the lesion property (solid metastasis).
[0035] Immunohistochemical correlation for ER / PR positive (IH-001, IH-002) suggests hormone sensitivity, high expression of Ki-67 (IH-003) suggests active tumor proliferation, and MSS (IH-004) excludes the immunotherapy advantage group. Correlate the above indicators with "stage IVB" and "ECOG 0 points" in structured data to determine the corresponding relationship between "high proliferation activity + liver metastasis" and the risk of disease progression.
[0036] Based on the correlation characteristics, a weighted algorithm is used to fuse and calculate the image quantification indicators, immunohistochemical results and clinical characteristics to generate a preliminary matching score. Combine the key points of endometrial cancer clinical trials and assign weights: liver metastasis size (20%), ER expression (25%), PR expression (20%), Ki-67 (20%), and MSS (15%). Preliminary matching score = (0.46 x 20%) + (0.8 x 25%) + (0.6 x 20%) + (0.7 x 20%) + (1 x 15%) = 0.092 + 0.2 + 0.12 + 0.14 + 0.15 = 0.702 (full score 1 point).
[0037] The preliminary matching score is checked and adjusted, the weight of the enrollment condition is corrected to correct the score deviation, and the final comprehensive matching score is generated to measure the matching degree of the patient and the clinical trial. The weight of the enrollment condition is corrected to be higher for "ER / PR positive" (weight increased to 30%) and lower for "MSS" (weight reduced to 5%) in a certain endometrial cancer trial. Recalculate the corrected score = (0.46 x 20%) + (0.8 x 30%) + (0.6 x 20%) + (0.7 x 20%) + (1 x 5%) = 0.092 + 0.24 + 0.12 + 0.14 + 0.05 = 0.642. The comprehensive matching score 0.642 indicates that the matching degree of the patient and the trial is 64.2%, which can be used for subsequent trial priority sorting.
[0038] In S104, the rule engine and semantic similarity calculation are used to realize the item-by-item comparison of patient characteristics and enrollment conditions, and generate a preliminary matching result.
[0039] In one embodiment, the rule engine is used to structure the comparison of patient characteristics and clinical trial enrollment conditions, extract the tumor type, stage, and treatment line number of the patient, and match them one by one with the preset judgment conditions of the clinical trial to generate a compliance label for each judgment condition. Taking a patient with advanced endometrial cancer as an example, the extracted "tumor type: endometrial cancer (poorly differentiated adenocarcinoma)", "stage: stage IVB (liver metastasis)", and "treatment line number: second line (planned third line treatment)" from the structured clinical characteristic data. These characteristics have been converted to structured data that can be directly compared through the standardization processing of the pre-stage large language model preprocessing and OCR analysis module.
[0040] The preset judgment conditions of the selected late-stage endometrial carcinoma phase III trial are based on the basic requirements for "tumor specificity, disease progression, and treatment experience" in the clinical trial enrollment criteria, consistent with the structured screening logic of "20 preset judgment conditions" in Appendix 1. Among them, "Condition 1: Tumor is endometrial carcinoma" directly corresponds to the tumor characteristics of the patient, ensuring that the trial matches the patient's disease; "Condition 2: Stage is ⅢB and above (including Ⅳ)" covers the patient's disease progression of ⅣB, meeting the requirements of the late-stage tumor trial for the stage of the disease; "Condition 3: Treatment line number is two and above" matches the current treatment stage of the patient, meeting the enrollment basis of the late-line treatment trial. The generation of the "+" label in the matching result follows the principle of "complete match between features and conditions, then determine compliance", providing a structured basis for subsequent semantic similarity calculation and correlation matrix construction.
[0041] The semantic similarity calculation module processes the patient's unstructured characteristics and the specific conditions of the clinical trial. By analyzing the semantic connotation of the enrollment and exclusion conditions through a large language model, the semantic matching degree between the patient's medical record description and the condition text is calculated, and a similarity score is generated. The patient's medical record "the patient has received albumin paclitaxel + carboplatin chemotherapy without serious adverse reactions, currently with ECOG 0 points, and liver metastasis is stable" belongs to unstructured text, and expressions such as "no serious adverse reactions" are not directly covered by structured features, and potential information needs to be mined through semantic analysis.
[0042] In the selected clinical trial specific conditions, "previous chemotherapy without 3 or above hematologic toxicity" and "existence of uncontrollable abdominal effusion" belong to the trial-specific requirements not covered by the preset judgment conditions. The large language model first analyzes the semantic connotation of the enrollment conditions and converts "3 or above hematologic toxicity" into the standardized judgment standard of "chemotherapy toxicity ≤ 2", and then performs semantic comparison with the patient's "no serious adverse reactions" - since "no serious adverse reactions" usually implies "toxicity reaction ≤ 2" in the clinical context, the matching degree is 90% (score 0.9); For exclusion conditions, the model analyzes that its core is "abdominal effusion needs to be drained repeatedly", and the patient's medical record does not mention abdominal effusion-related content. According to the logic of "not mentioned by default, not meeting the exclusion conditions", the matching degree is 100% (score 1.0). Through the semantic understanding ability of the large language model, the limitations of structured comparison in handling ambiguous expressions are overcome, and the generated similarity score will be combined with the compliance label of the previous structured matching to provide quantitative support for the subsequent correlation matrix construction, ensuring the coherence of the previous and subsequent processes.
[0043] The association matrix of patient characteristics and clinical trial conditions is established by combining the preset judgment condition matching label and the semantic similarity score, and the matching relationship of each characteristic and condition is determined. The association matrix established by combining the preset judgment condition matching label and the semantic similarity score needs to clearly present the matching relationship of patient characteristics and clinical trial conditions in text form. Taking endometrial cancer patients as an example, in the tumor type dimension, the endometrial cancer suffered by the patient corresponds to the condition of “tumor type is endometrial cancer” required by the trial, the preset judgment condition matching label is “+”, indicating that they are completely matched, which belongs to the direct correspondence of structured characteristics; in the staging dimension, the patient's ⅣB stage meets the condition of the trial “staging is ⅢB stage and above (including Ⅳ stage)”, the preset judgment condition matching label is “+”, which embodies the matching of structured characteristics; in the treatment line number dimension, the patient is currently treated as the second line and plans to be treated as the third line, which meets the requirement of the trial “treatment line number is the second line and above”, the preset judgment condition matching label is “+”, which is also the matching of structured characteristics.
[0044] For the unstructured characteristic of chemotherapy toxicity, the trial requires “no 3rd grade or above hematological toxicity occurred in previous chemotherapy”, and the description of “no serious adverse reactions” in the patient's medical record is highly matched with the determination standard of “chemotherapy toxicity ≤ 2nd grade” after semantic analysis, with a semantic similarity score of 0.9; in terms of ascites, the trial excludes “existence of uncontrollable ascites”, and the patient's medical record does not mention related content, so it is defaulted to not meet the exclusion condition, with a semantic similarity score of 1.0, which embodies the semantic association of unstructured characteristics and trial specific conditions. Through such text description, it can be determined that the structured characteristics of the patient meet the basic enrollment conditions of the trial, and the semantic analysis of the unstructured characteristics also achieves a high matching degree, providing comprehensive quantitative basis for the generation of the subsequent preliminary matching list.
[0045] Based on the association matrix, the patient and the clinical trial are preliminarily matched, and the trial plan that meets the preset judgment condition and achieves the semantic similarity score is retained to generate the preliminary matching list. Taking endometrial cancer patients as an example, the screening criteria are “structured conditions all meet (label “+”)” and “semantic similarity score ≥ 0.8”. Among them, the structured conditions correspond to core dimensions such as tumor type, staging, and treatment line number, which need to be completely matched with the preset judgment conditions of the clinical trial (i.e. the label is “+”); the semantic similarity score is for unstructured characteristics (such as chemotherapy toxicity, ascites, etc.), which needs to reach the “high matching” threshold (≥ 0.8). For a certain late-stage endometrial cancer Ⅲ trial, the structured characteristics of the patient (endometrial cancer, ⅣB stage, second-line treatment) all meet the preset judgment conditions of the trial (the label is “+”), and the semantic similarity scores of the unstructured characteristics (chemotherapy toxicity 0.9, ascites 1.0) all exceed the threshold of 0.8, meeting the double screening criteria, so it is included in the preliminary matching list.
[0046] The other trial required "only 1 chemotherapy regimen", but the patient actually received 2 or more chemotherapy regimens such as albumin paclitaxel + carboplatin, cediranib, etc. The number of treatment regimens in the structured condition does not match (label "-"), and it does not pass the first layer of screening and is directly excluded. By strictly implementing the screening criteria, it not only ensures that the trials that meet the core enrollment conditions are retained, but also filters low matching degree specific conditions through semantic similarity threshold, providing a reliable candidate list for subsequent verification steps.
[0047] The preliminary matching list is checked to filter suspected mispairs caused by data missing or ambiguity, generating a preliminary matching result containing matching items, non-compliant items and judgment basis. For the preliminary matching list of endometrial cancer patients, the verification content focuses on two dimensions: data integrity: check whether the patient's core information is complete, such as "KRAS G12C mutation" in the gene test result is clearly recorded, there is no missing key indicator, which meets the basic requirements of biomarker information for clinical trials, avoiding misjudgment due to incomplete data. Logical consistency: verify the internal consistency of structured features and unstructured descriptions. The structured annotation of "second-line treatment" of the patient is consistent with the medical record description of "planned third-line treatment" (the former is the current treatment stage, and the latter is the next plan), and the semantic similarity score (chemotherapy toxicity 0.9, ascites 1.0) completely corresponds to the original content in the medical record "no serious adverse reactions" "no ascites mentioned", without logical conflicts.
[0048] The preliminary matching result generated after verification clearly distinguishes between two types of trials: matching items: a certain phase III trial for advanced endometrial cancer meets all three structured conditions (tumor type, stage, and number of treatment lines) (label "+"), and the semantic similarity scores are all ≥0.8, so it is determined to be a match, and its basis is directly associated with the association matrix and screening criteria in the previous text. Non-compliant items: a trial that only includes patients in the first line of treatment is explicitly listed as non-compliant because the patient is currently in the second line of treatment, and the structured condition of the number of treatment lines does not match (label "-"). The judgment basis directly corresponds to the structured comparison result. This verification process eliminates the risk of data missing and logical conflicts, providing high-quality input data for the subsequent large language model optimization module, ensuring the rigor of the matching results.
[0049] S105, based on the large language model optimization module, corrects the misjudgment of edge cases through context-aware multi-round reasoning, adjusts the ranking in combination with the priority weight of clinical trials, and generates an optimized clinical trial matching list.
[0050] In one implementation, the large language model optimization module analyzes the edge cases in the preliminary matching results, extracts the fuzzy matching items in the patient characteristics and clinical trial conditions, corrects the misjudgments through context-aware multi-round reasoning, and generates corrected matching labels. In the preliminary matching results, a certain endometrial cancer patient was excluded by a certain trial due to "abdominal fluid (not drained) for nearly 1 month" recorded in the medical record (the exclusion condition of this trial is "uncontrollable abdominal fluid"). The large language model optimization module analyzes the semantic connotation of "not drained" through context-aware multi-round reasoning - combining clinical common sense, the abdominal fluid not drained usually does not meet the "uncontrollable" standard, so the original "not eligible" label is corrected to "eligible" (i.e. not triggering the exclusion condition), and the corrected matching label is generated.
[0051] The priority weight parameters of the clinical trial settings cover the trial phase, the degree of indication matching, and the dimension of potential benefit, and each trial plan is assigned a weight value. The priority weights of the three endometrial cancer clinical trials included in the preliminary matching list cover three dimensions: trial phase (10% for phase I, 20% for phase II, and 30% for phase III): reflecting the maturity of the trial, phase III trial has the highest weight due to more sufficient data; degree of indication matching (up to 100%): assigning values according to the degree of agreement of tumor type, stage, and genetic characteristics, such as a trial specifically targeting "advanced endometrial cancer with KRAS G12C mutation" with a matching degree of 90%; potential benefit (up to 100%): combining drug mechanism and patient genetic characteristics for evaluation, such as assigning a value of 80% to the potential benefit of KRAS G12C mutation patients in a targeted drug trial.
[0052] Finally, each trial is assigned a comprehensive weight value: Trial A (phase III, indication matching degree 90%, potential benefit 80%) has a weight of 30% x 0.3 + 90% x 0.5 + 80% x 0.2 = 73%; Trial B (phase II, indication matching degree 85%, potential benefit 70%) has a weight of 20% x 0.3 + 85% x 0.5 + 70% x 0.2 = 65.5%; Trial C (phase I, indication matching degree 80%, potential benefit 60%) has a weight of 10% x 0.3 + 80% x 0.5 + 60% x 0.2 = 55%.
[0053] The association model of the matching result and the weight is established by combining the revised matching label and the priority weight value, and the adjustment basis of each test scheme is clear. The revised matching label is associated with the weight value, and the sorting logic is clear. Specifically, for test A: the revised label is "consistent", the weight is 73%, the priority is 1, and the basis is "phase III test + high matching degree of indication + significant potential target benefit"; for test B: the label is "consistent", the weight is 65.5%, the priority is 2, and the basis is "phase II test + high matching degree of indication"; for test C: the label is "consistent", the weight is 55%, the priority is 3, and the basis is "phase I test + moderate potential benefit".
[0054] The preliminary matching result is sorted and adjusted based on the association model, the test scheme with high weight and consistent with the condition after revision is preferentially retained, and a sorted candidate list is generated. The test schemes consistent with the condition after revision are retained in descending order of weight value, and a candidate list is generated: test A (phase III targeted therapy, matching degree 92%, priority 1); test B (phase II immunotherapy, matching degree 88%, priority 2); test C (phase I chemotherapy combination, matching degree 80%, priority 3).
[0055] The final verification is performed on the candidate list, the repeated items and the schemes not meeting the core enrollment conditions are filtered, and an optimized clinical test matching list containing the test name, matching degree and priority is generated. The candidate list is double-checked, specifically, repeated item filtering: no repeated test, no need to process; core enrollment condition review: confirming that all tests meet the core conditions such as "advanced endometrial cancer, second-line and above treatment, KRAS G12C mutation". Finally, an optimized matching list is generated, test A (phase III targeted therapy): matching degree 92%, priority 1, key basis is "consistent with the exclusion condition after revision, highest weight and complete indication matching"; test B (phase II immunotherapy): matching degree 88%, priority 2, key basis is "high matching degree of indication, mature test phase".
[0056] The revised label solves the misjudgment caused by semantic ambiguity in the preliminary matching (such as the degree definition of "ascites"); the priority weight combines the semantic analysis results of the "specific condition matching model", ensuring that the sorting is consistent with the actual value of the clinical test; the final list provides a precise test recommendation order for subsequent matching report generation, which meets the "patient-centered" matching logic.
[0057] S106, input the structured clinical feature data, supplementary information, preliminary matching result and optimized clinical test matching list into the clinical decision support module, visualize the matching whole path, and generate a clinical test matching report containing matching confidence, key basis and risk prompt.
[0058] In one embodiment, based on the clinical decision support module processing the structured clinical feature data, supplementary information, preliminary matching results and the optimized clinical trial matching list, the multi-source data is integrated by a data correlation algorithm to construct a matching relationship graph between patient characteristics and clinical trials, which contains a patient core feature set, a trial key condition set and a matching quantization value. The clinical decision support module integrates the structured clinical feature data of endometrial cancer patients (tumor type: endometrial cancer IVB stage, treatment line number: second line, ECOG score 0), supplementary information (immunohistochemical ER+80%, KRAS G12C mutation), preliminary matching results (3 trials meet the core conditions) and the optimized matching list (trial A priority 1, trial B priority 2). Through the data correlation algorithm, the patient core characteristics (such as "KRAS G12C mutation" and "liver metastasis") are associated with the trial key conditions (such as "late-stage endometrial cancer post-line treatment" and "KRAS G12C mutation positive"), and a relationship graph containing a matching quantization value is generated, for example, the matching degree of trial A and the patient is 92%, corresponding to the associated nodes such as "targeted drug adaptation KRAS G12C mutation" and "stage completely consistent".
[0059] The matching full path is presented using a visual rendering technology, and the complete process from data extraction to optimization sorting is displayed in the form of a time axis. Interactive nodes are set to allow backtracking to view the basis for judging each link, and a dynamic matching path graph is generated. The matching full process is displayed in the form of a time axis using a visual rendering technology. Specifically, the starting point (0 minutes): OCR identifies medical record pictures to extract core features such as "endometrial cancer with liver metastasis and second-line treatment"; 3 minutes: rule engine structured comparison to generate labels such as "tumor type, stage, and treatment line number all meet"; 5 minutes: semantic similarity calculation to obtain "chemotherapy toxicity matching degree 0.9, no ascites matching degree 1.0"; 8 minutes: large language model correction of edge cases (such as "un-drained ascites ≠ uncontrollable"); 10 minutes: priority weight sorting to generate an optimized list. Interactive buttons are set at each node of the time axis, and clicking on "semantic similarity calculation" allows backtracking to view the semantic alignment basis for "no serious adverse reactions" and "chemotherapy toxicity ≤ grade 2".
[0060] A report generation matrix is constructed to calculate the confidence parameters of each matching indicator based on the comprehensive matching score and the clinical priority weight, wherein the matching confidence uses a Bayesian probability correction model, and the key basis uses a text highlighting annotation algorithm. Based on the comprehensive matching score (trial A 64.2 points) and the clinical priority weight (Ⅲ stage trial weight 30%), a report generation matrix is constructed: the matching confidence uses a Bayesian probability correction model, combined with historical matching data (the accuracy of the same patient enrolling in trial A is 91%), to calculate the matching confidence of trial A as 94%.
[0061] By the text highlighting algorithm, highlight "KRAS G12C mutation" "liver metastasis" and other content directly matched with the inclusion criteria of test A in the medical record, highlight "targeting KRAS G12C" "second line and above treatment" and other key conditions in the test plan, to ensure traceability according to the matching basis.
[0062] Based on the report generation result and the preset risk assessment standard, combined with the characteristics of the test type, a clinical trial matching report containing matching confidence, key matching basis and risk prompt is generated, wherein the risk prompt is associated with the clinical trial adverse event database for annotation. Based on the preset risk assessment standard (such as "targeted drug common skin rash adverse reaction"), combined with the characteristics of the test type (I / II / III phase), the final report is generated, as follows, matching confidence: test A 94%, test B 88%; key basis: test A is listed as the first choice due to "KRAS G12C mutation adaptation + mature Ⅲ phase data"; risk prompt: reference clinical trial adverse event database, mark that test A may have "3 level skin rash (occurrence rate 12%)", test B needs to be alert to "immune related pneumonia (occurrence rate 8%)". The report also notes that "the final enrollment needs to review blood routine and imaging examination", which meets the requirement of "further examination after screening" in the research design, and provides complete reference for clinical decision making.
[0063] The present application improves the matching efficiency and accuracy of clinical trials for tumor patients through multi-module collaborative processing. First, a "tumor type-stage-treatment line number" three-dimensional knowledge graph is constructed, integrating guideline and literature data, and upgrading static judgment conditions to dynamic association rules. Use a large language model to analyze clinical data, combined with a time series neural network to accurately stage, generate structured clinical feature data; simultaneously process paper and image data using OCR technology, and generate supplementary information after double checking.
[0064] The multi-modal processing pipeline module enhances data, extracts image quantitative indicators, and analyzes immunohistochemical results to generate a comprehensive matching score. Rule engine and semantic similarity calculation realize item-by-item comparison to generate preliminary matching results. The large language model optimization module corrects misjudgments, adjusts the order combined with priority weights, and generates an optimized list. The clinical decision support module integrates data, visualizes the matching path, constructs a report generation matrix, and finally generates a report containing matching confidence, key basis and risk prompt, providing an efficient and accurate tool for tumor clinical trial recruitment.
[0065] In one embodiment, as shown in Figure 2 The present application also provides a tumor patient clinical trial matching device based on a large language model and OCR technology, comprising:
[0066] The acquisition module 201 is configured to acquire clinical data and clinical trial information of a tumor patient, pre-process based on a knowledge graph construction module, construct a "tumor type-stage-treatment line number" three-dimensional knowledge graph, integrate NCCN / CSCO guidelines and real-time literature data, and upgrade static judgment conditions to dynamic correlation rules.
[0067] The processing module 202 is configured to analyze the clinical data and trial information based on a large language model preprocessing module, process unstructured text, combine a time sequence neural network and a RECIST1.1 standard time axis analysis model to accurately stage, and generate structured clinical feature data through context correlation analysis. The paper and image data are processed by an OCR technology analysis module to extract key data and correct them through double verification, and generate structured data supplement information. Based on a multi-modal processing pipeline module, the structured clinical feature data and the supplement information are enhanced to extract image quantitative indicators and analyze immunohistochemical results, and generate a comprehensive matching score. Through a rule engine and semantic similarity calculation, the patient features and the entry and exit conditions are compared item by item to generate a preliminary matching result. Based on a large language model optimization module, edge case misjudgments are corrected through context-aware multi-round reasoning, and the ordering is adjusted in combination with clinical trial priority weight to generate an optimized clinical trial matching list. The structured clinical feature data, the supplement information, the preliminary matching result, and the optimized clinical trial matching list are input into a clinical decision support module to visually display the matching path and generate a clinical trial matching report containing matching confidence, key basis, and risk prompts.
[0068] Each of the embodiments in the present application is described in a related manner, and the same and similar parts between the embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, for the evaluation of the tumor patient clinical trial matching method, the electronic device, the electronic equipment, and the readable storage medium embodiments based on the large language model and the OCR technology, since they are basically similar to the above-mentioned tumor patient clinical trial matching method embodiments based on the large language model and the OCR technology, the description is relatively simple, and the relevant parts can be referred to the above-mentioned tumor patient clinical trial matching method embodiments based on the large language model and the OCR technology.
Claims
1. A method for matching a tumor patient with a clinical trial based on a large language model and OCR technology, characterized in that, Comprise: Obtain the clinical data and clinical trial information of tumor patients, preprocess based on the knowledge graph construction module, construct a "tumor type-stage-treatment line number" three-dimensional knowledge graph, integrate NCCN / CSCO guidelines and real-time literature data, and upgrade static judgment conditions to dynamic correlation rules; Based on the large language model preprocessing module, analyze the clinical data and trial information, process unstructured text, and combine the time series neural network and RECIST1.1 standard time axis analysis model to accurately stage, and generate structured clinical feature data through context association analysis; Simultaneously, through the OCR technology analysis module, process paper and image data, extract key data and correct through double verification, and generate structured data supplement information; Based on the multi-modal processing pipeline module, enhance the structured clinical feature data and supplementary information, extract image quantitative indicators, and analyze immunohistochemical results to generate a comprehensive matching score; Through the rule engine and semantic similarity calculation, the patient's characteristics and the entry and exclusion conditions are compared item by item to generate a preliminary matching result; Based on the large language model optimization module, through context-aware multi-round reasoning to correct edge case misjudgment, combined with clinical trial priority weight adjustment sorting, generate an optimized clinical trial matching list; Input the structured clinical feature data, supplementary information, preliminary matching result, and optimized clinical experiment matching list into the clinical decision support module, visualize the matching path, and generate a clinical trial matching report containing matching confidence, key basis, and risk prompts.
2. The method of claim 1, wherein, Based on the large language model preprocessing module, analyze the clinical data and trial information, process unstructured text, and combine the time series neural network and RECIST1.1 standard time axis analysis model to accurately stage, and generate structured clinical feature data through context association analysis; Simultaneously, through the OCR technology analysis module, process paper and image data, extract key data and correct through double verification, and generate structured data supplement information, including: Based on the large language model preprocessing module, analyze the non-structured text in the clinical data and trial information, extract the patient's condition description and treatment history, combine the time series neural network and RECIST1.1 standard time axis analysis model to accurately stage the patient's condition, and generate preliminary staging feature data; Perform context association analysis on the preliminary staging feature data to eliminate text ambiguity, map key clinical indicators to standardized features, and generate structured clinical feature data, which includes core medical record elements such as tumor type, treatment line number, stage, and ECOG score; Simultaneously, through the OCR technology analysis module, perform text recognition on paper and image data, extract key values and terms from medical records and examination reports, and combine text box position information for content splicing and format arrangement to generate original recognition text; Apply character-level and semantic-level double verification rules to the original recognition text, identify and correct abnormal text, perform data format standardization processing, and generate structured data supplement information including immunohistochemical information and gene detection results.
3. The method of claim 2, wherein, Based on the multi-modal processing pipeline module, the structured clinical feature data and supplementary information are enhanced, the image quantitative indicators are extracted, the immunohistochemical results are analyzed, and the comprehensive matching scores are generated, including: Based on the multi-modal processing pipeline module, the structured clinical feature data and supplementary information are analyzed, the lesion size and density in the image data are extracted, and the protein expression level in the immunohistochemical results is analyzed, and a unique identifier is assigned to each type of indicator; Standardize the image quantitative indicators and immunohistochemical analysis results, eliminate the dimensional differences between different indicators through data normalization, and generate standardized feature data; Correlation analysis is performed on the standardized feature data, structured clinical feature data, and supplementary information to establish the correlation characteristics between indicators and clarify the corresponding relationship between each indicator and the patient's condition; Based on the correlation characteristics, a weighted algorithm is used to fuse and calculate the image quantitative indicators, immunohistochemical results, and clinical features to generate preliminary matching scores; The preliminary matching scores are verified and adjusted, the score bias is corrected by combining the weight of the inclusion criteria of the clinical trial, and the final comprehensive matching scores are generated to measure the matching degree of the patient and the clinical trial.
4. The method of claim 1, wherein, Through the rule engine and semantic similarity calculation, the patient features are compared with the inclusion and exclusion conditions one by one to generate preliminary matching results, including: Based on the rule engine, the patient features are compared with the inclusion and exclusion conditions of the clinical trial, the tumor type, stage, and treatment line number of the patient are extracted, and they are matched one by one with the pre-set judgment conditions of the clinical trial to generate a compliance label for each judgment condition; The semantic similarity calculation module processes the patient's unstructured features and the specific conditions of the clinical trial, analyzes the semantic connotation of the inclusion and exclusion conditions through a large language model, calculates the semantic matching degree between the patient's medical record description and the condition text, and generates a similarity score; Based on the pre-set judgment condition compliance label and the semantic similarity score, an association matrix of patient features and clinical trial conditions is established to clarify the matching relationship between each feature and condition; Based on the association matrix, the patient and the clinical trial are preliminarily matched, and the trial plan that meets the pre-set judgment condition and has a satisfactory semantic similarity is retained to generate a preliminary matching list; The preliminary matching list is verified to filter out suspected misfits caused by data missing or ambiguity, and a preliminary matching result containing matching items, non-compliance items, and judgment basis is generated.
5. The method of claim 1, wherein, Based on the large language model optimization module, edge case misjudgments are corrected through context-aware multi-round reasoning, and the order is adjusted based on the priority weight of the clinical trial to generate an optimized clinical trial matching list, including: Based on the large language model optimization module, the edge cases in the preliminary matching result are analyzed, the fuzzy matching items in the patient features and the clinical trial conditions are extracted, and the misjudgments are corrected through context-aware multi-round reasoning to generate corrected matching labels; Set priority weight parameters for the clinical trial, covering trial phase, indication matching degree, and potential benefit dimensions, and assign a weight value to each trial plan; Based on the corrected matching labels and priority weight values, an association model of matching results and weights is established to clarify the adjustment basis for each trial plan; Based on the correlation model, the preliminary matching results are sorted and adjusted, high-weight and corrected test programs that meet the conditions are preferentially retained, and a sorted candidate list is generated; The candidate list is finally checked to filter out duplicate items and programs that do not meet the core enrollment conditions, and an optimized clinical trial matching list containing the test name, matching degree, and priority is generated.
6. The method of claim 5, wherein, Input the structured clinical feature data, supplementary information, preliminary matching results, and optimized clinical trial matching list into the clinical decision support module, visualize the matching full path, and generate a clinical trial matching report containing matching confidence, key basis, and risk prompts, including: Based on the clinical decision support module, the structured clinical feature data, supplementary information, preliminary matching results, and optimized clinical trial matching list are processed, multi-source data is integrated through data correlation algorithms, and a matching relationship graph between patient characteristics and clinical trials is constructed, which contains the patient core feature set, the test key condition set, and the matching quantitative value; Use visual rendering technology to present the matching full path, display the complete process from data extraction to optimized sorting in the form of a time axis, set interactive nodes to allow backtracking to view the judgment basis at each link, and generate a dynamic matching path graph; Build a report generation matrix, combine the comprehensive matching score and the clinical priority weight, calculate the confidence parameters of each matching index, and use the Bayesian probability correction model for matching confidence and the text highlighting algorithm for key basis; Based on the report generation results and the preset risk assessment standard, combined with the characteristics of the test type, a clinical trial matching report containing matching confidence, key matching basis, and risk prompts is generated, and the risk prompts refer to the clinical trial adverse event database for association annotation. 7.A tumor patient clinical trial matching device based on a large language model and an OCR technology, characterized in that, The device comprises: An acquisition module is configured to acquire clinical data and clinical trial information of a tumor patient, pre-process based on a knowledge graph construction module, construct a "tumor type-stage-treatment line number" three-dimensional knowledge graph, integrate NCCN / CSCO guidelines and real-time literature data, and upgrade static judgment conditions to dynamic correlation rules; The processing module is used for analyzing clinical data and test information based on a large language model preprocessing module, processing unstructured text, combining a time sequence neural network and a RECIST1.1 standard time axis analysis model for accurate staging, and generating structured clinical feature data through context correlation analysis; paper and image data are processed by an OCR technology analysis module to extract key data and correct them through double verification to generate structured data supplement information; the structured clinical feature data and the supplement information are enhanced based on a multi-modal processing pipeline module to extract image quantitative indicators and analyze immunohistochemical results, and a comprehensive matching score is generated; through a rule engine and semantic similarity calculation, item-by-item comparison of patient features and entry and exit conditions is realized to generate a preliminary matching result; based on a large language model optimization module, edge case misjudgments are corrected through context-aware multi-round reasoning, and the ordering is adjusted in combination with clinical trial priority weight to generate an optimized clinical trial matching list; the structured clinical feature data, the supplement information, the preliminary matching result and the optimized clinical trial matching list are input into a clinical decision support module to visually display the matching full path and generate a clinical trial matching report containing matching confidence, key basis and risk prompts.
8. An electronic device, comprising: It comprises: a first processor; and a memory for storing executable instructions of the first processor; wherein the first processor is configured to execute the executable instructions to perform the tumor patient clinical trial matching method based on a large language model and an OCR technology according to any one of claims 1-6.
Citation Information
Patent Citations
Device for constructing tumor knowledge base, computer readable storage medium and application
CN118098343A
Tumor standardized whole-course management system based on digitization
CN119108074A
Clinical test matching method, device, equipment, medium and product
CN119153008A
TNM (Tumor Necrosis Model) staging prediction method and system in clinical text of large language model
CN119830907A
Clinical auxiliary decision-making system based on big data
CN120108696A
Cited By
Image-text medical examination report generation method and system based on deep learning
CN121171460A
Patient matching method and system driven by natural language and used for scientific research and appointment and appointment
CN121439271A
A natural language driven scientific research patient matching method and system
CN121439271B
Intelligent drug clinical test scheme generation method and system
CN122091265A