A Clinical Trial Matching System and Method for Cancer Patients Based on Large Language Model and OCR Technology
By constructing a three-dimensional knowledge graph and combining a large language model with OCR technology into a multi-module system, the shortcomings of existing clinical trial matching models in terms of accuracy and applicability are solved, achieving efficient and accurate matching of clinical trials for cancer patients.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2026-03-10
AI Technical Summary
Existing AI-based clinical trial matching models have shortcomings in accuracy and applicability, especially when targeting different hospitals and patient groups. They cannot fully reflect individualized information, and the accuracy of information extraction is highly dependent on the quality and completeness of structured data, resulting in low matching accuracy.
A multi-module collaborative system based on a large language model and OCR technology is constructed. By building a three-dimensional knowledge graph of "tumor type-stage-treatment line number", clinical data is analyzed and structured data is generated. Combined with OCR processing of image data, multimodal analysis and rule verification are performed to generate a matching report containing confidence level and risk warning.
It improves the accuracy and efficiency of clinical trial matching for cancer patients, and generates precise matching reports that include matching degree, key evidence and risk warnings, supporting precise recruitment for clinical trials.
Smart Images

Figure CN120913728B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical data processing, and in particular to a clinical trial matching system and method for tumor patients based on large language models and OCR technology. Background Technology
[0002] Existing AI-based clinical trial matching models primarily employ natural language processing to extract key information from both patient and clinical trial data before matching and selecting compatible patients. However, the accuracy of information extraction heavily relies on the quality and completeness of the structured data; any omission or error directly impacts matching accuracy. This is especially true for complex inclusion and exclusion criteria, which are often difficult to accurately match using existing methods. Consequently, matching accuracy is low, prone to errors, and fails to cover all potential inclusion or exclusion factors.
[0003] Furthermore, most existing AI analysis methods rely on structured data from medical record systems, which are often limited by the framework of hospital information systems and the degree of standardization of medical records. This restricts the applicability and generalization of AI models, especially when targeting different hospitals and patient groups, and may fail to fully reflect individualized patient information.
[0004] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore includes information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] The purpose of this application is to provide a matching system and method for oncology patient clinical trials based on a large language model and OCR technology, which at least to some extent overcomes the problems of existing technologies and improves the efficiency and accuracy of oncology clinical trial matching through multi-module collaboration. A three-dimensional knowledge graph of "tumor type-stage-treatment line number" is constructed, and dynamic rules are upgraded. The large language model parses clinical data to generate structured data, and OCR processes imaging data to generate supplementary information. Through multimodal analysis, rule validation, and model optimization, a report containing confidence levels, supporting evidence, and risk warnings is finally generated, providing a precise tool for trial recruitment.
[0006] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part by practice of the invention.
[0007] According to one aspect of this application, a method for matching clinical trials of cancer patients based on a large language model and OCR technology is provided, comprising: acquiring clinical data and clinical trial information of cancer patients; preprocessing the data based on a knowledge graph construction module to construct a three-dimensional knowledge graph of "tumor type-stage-treatment line number"; integrating NCCN / CSCO guidelines and real-time literature data; and upgrading static judgment conditions to dynamic association rules; parsing the clinical data and trial information based on a large language model preprocessing module, processing unstructured text, accurately staging the data by combining a temporal neural network and the RECIST 1.1 standard time axis analysis model, and generating structured clinical feature data through contextual association analysis; and simultaneously processing paper and imaging data through an OCR technology parsing module, extracting key data and correcting it through double verification. The system generates supplementary structured data; enhances the structured clinical feature data and supplementary information based on the multimodal processing pipeline module, extracts image quantification indicators, analyzes immunohistochemical results, and generates a comprehensive matching score; compares patient features with inclusion and exclusion conditions item by item through a rule engine and semantic similarity calculation to generate preliminary matching results; optimizes the clinical trial matching list by correcting misjudgments of marginal cases through context-aware multi-round reasoning based on the large language model optimization module, and adjusts the ranking by combining clinical trial priority weights; inputs the structured clinical feature data, supplementary information, preliminary matching results, and optimized clinical trial matching list into the clinical decision support module, visualizes the entire matching path, and generates a clinical trial matching report that includes matching confidence, key evidence, and risk warnings.
[0008] Another aspect of this application discloses a tumor patient clinical trial matching device based on a large language model and OCR technology, characterized by comprising: an acquisition module for acquiring clinical data and clinical trial information of tumor patients, performing preprocessing based on a knowledge graph construction module to construct a three-dimensional knowledge graph of "tumor type-stage-treatment line number," integrating NCCN / CSCO guidelines and real-time literature data, and upgrading static judgment conditions to dynamic association rules; a processing module for parsing clinical data and trial information based on the large language model preprocessing module, processing unstructured text, accurately staging the tumor by combining a temporal neural network and the RECIST 1.1 standard time axis analysis model, and generating structured clinical feature data through contextual association analysis; and simultaneously processing paper and image data through an OCR technology parsing module to extract key data. After double verification and correction, supplementary structured data information is generated. The structured clinical feature data and supplementary information are enhanced using a multimodal processing pipeline module, extracting quantitative imaging indicators, analyzing immunohistochemical results, and generating a comprehensive matching score. A rule engine and semantic similarity calculation are used to compare patient characteristics with inclusion and exclusion criteria item by item, generating preliminary matching results. Based on a large language model optimization module, context-aware multi-round reasoning corrects misjudgments of marginal cases, and the ranking is adjusted by combining clinical trial priority weights to generate an optimized clinical trial matching list. The structured clinical feature data, supplementary information, preliminary matching results, and optimized clinical trial matching list are input into the clinical decision support module, which visualizes the entire matching path and generates a clinical trial matching report including matching confidence, key evidence, and risk warnings.
[0009] According to another aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a second processor, implements the above-described method for matching tumor patients in clinical trials based on a large language model and OCR technology.
[0010] This application provides a tumor patient clinical trial matching system and method based on a large language model and OCR technology, which improves the efficiency and accuracy of tumor patient clinical trial matching through multi-module collaboration. First, a three-dimensional knowledge graph of "tumor type-stage-treatment line number" is constructed, integrating guidelines and literature, and upgrading static conditions to dynamic rules. The large language model is used to analyze clinical data and accurately stage it, generating structured data; simultaneously, OCR technology is used to process paper and imaging data, generating supplementary information after double verification. A multimodal module extracts imaging indicators and analyzes immunohistochemistry to generate matching scores. Preliminary results are obtained through a rule engine and semantic calculations, then optimized, corrected, and ranked by the large language model. Finally, a clinical decision support module generates a report containing confidence levels, supporting evidence, and risk warnings, providing a precise tool for trial recruitment.
[0011] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0012] Figure 1 This document illustrates a flowchart of a clinical trial matching method for cancer patients based on a large language model and OCR technology, provided in an embodiment of this application.
[0013] Figure 2 This illustration shows a schematic diagram of a tumor patient clinical trial matching device based on a large language model and OCR technology, provided in one embodiment of this application. Detailed Implementation
[0014] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0015] The following is combined Figure 1 This application describes a method for matching tumor patients in clinical trials based on a large language model and OCR technology, according to exemplary embodiments thereof. In one embodiment, this application also proposes a system and method for matching tumor patients in clinical trials based on a large language model and OCR technology:
[0016] S101 acquires clinical data and clinical trial information of cancer patients, performs preprocessing based on the knowledge graph construction module, constructs a three-dimensional knowledge graph of "tumor type-stage-number of treatment lines", integrates NCCN / CSCO guidelines and real-time literature data, and upgrades static judgment conditions to dynamic association rules.
[0017] In one implementation, clinical data and clinical trial information of several cancer patients were collected. Taking one patient as an example, the patient's current medical history (diagnosed with endometrial cancer more than 2 years ago, liver progression in March 2023, and a follow-up liver progression test on February 23, 2024); treatment history (from August 9, 2021, paclitaxel + carboplatin → February 2022, sintilimab → March 2023, paclitaxel + carboplatin → September 2023, bevacizumab → November 2023, bevacizumab + letrozole); symptoms (diabetic ketoacidosis occurred in September 2022, improved after symptomatic treatment, currently on insulin therapy). Genetic testing results showed NGS (Ben-Zhi) with KRASG12C mutation and MSS. The patient's score / basic information was ECOG = 1; age = 30; sex = male; stage = advanced. The treatment lines were determined to be first-line (2021-08-09 albumin-bound paclitaxel + carboplatin) and second-line (2023-03 albumin-bound paclitaxel + carboplatin) based on analysis. The current treatment is second-line, and the next step is planned to be third-line treatment.
[0018] Clinical trial information covers enrollment criteria (e.g., age range, ECOG requirements, tumor stage matching, gene testing (KRASG12C, MSS) limitations), exclusion criteria (e.g., acute diabetic ketoacidosis, recent major surgery), trial objectives (e.g., for second-line treatment of advanced endometrial cancer), and indications (advanced metastatic endometrial cancer). Core elements are extracted from the structured clinical data, including: tumor type (endometrial cancer), stage (advanced), treatment lines (first-line, second-line, planned third-line), gene testing (KRASG12C mutation, MSS), and key events (liver metastasis progression, diabetic ketoacidosis). Key trial conditions include: matching enrollment stage (applicable to advanced stages), treatment lines (e.g., whether third-line enrollment is after second-line failure), and gene characteristics (KRASG12C mutation, MSS suitability for the trial), laying a foundation for knowledge graph association.
[0019] Based on actual medical record data, a node linking "endometrial cancer - advanced stage - second-line / third-line treatment" is established. This link includes treatment regimens in this scenario (such as existing regimens like bevacizumab + letrozole), genetic characteristics (KRASG12C mutation, treatment response evidence corresponding to MSS), and guideline adaptation (NCCN / CSCO guidelines recommending later-line treatment for advanced endometrial cancer).
[0020] The guidelines for endometrial cancer, including NCCN / CSCO guidelines, address staging criteria for advanced / metastatic cancer, definitions of treatment lines (e.g., changing treatment plans after disease progression), and indication matching standards (e.g., enrollment and adaptation of patients with specific genetic characteristics). Real-time literature studies on the response of KRASG12C mutations and MSS phenotypes in endometrial cancer treatment (e.g., immunotherapy, targeted therapy) are also included, supplementing the knowledge graph's "endometrial cancer - advanced - later-line treatment" node.
[0021] The inclusion and exclusion criteria are dynamic. For example, in clinical trials, patients with acute diabetic ketoacidosis are excluded. The knowledge graph dynamically determines whether the patient is suitable by associating the patient's medical records with "the patient had ketoacidosis but has recovered with symptomatic treatment". For the treatment line, the line determination is dynamically revised based on the "postoperative / post-treatment progression time-line number mapping model" and the patient's timeline of "transition from first-line to second-line treatment due to liver progression in March 2023", which supports the matching of the trial's inclusion line conditions.
[0022] S102, based on the large language model preprocessing module, parses clinical data and trial information, processes unstructured text, combines temporal neural networks and RECIST 1.1 standard time axis analysis model for accurate staging, and generates structured clinical feature data through contextual association analysis; simultaneously, it processes paper and image data through the OCR technology parsing module, extracts key data and performs double verification and correction, and generates supplementary structured data information.
[0023] In one implementation, a large language model preprocessing module parses unstructured text from clinical data and trial information to extract the patient's condition description and treatment history. This is then combined with a temporal neural network and the RECIST 1.1 standard time axis analysis model to accurately determine the patient's condition stage and generate preliminary staging feature data. The patient's medical record states: "Diagnosed with endometrial cancer in March 2022, underwent total hysterectomy and bilateral salpingo-oophorectomy. Postoperative pathology indicated poorly differentiated adenocarcinoma invading the deep myometrium, with cervical stroma involvement, and no lymph node metastasis (0 / 12). A follow-up examination in May 2023 revealed a metastatic lesion in liver segment S5 (2.3 cm in diameter), with no other distant metastases."
[0024] The large language model extracts disease descriptions (initial diagnosis, postoperative recurrence) and treatment history (surgical resection). Combined with temporal neural networks and RECIST 1.1 criteria, the initial stage is determined to be stage IB and the stage after recurrence is stage IVB through time axis analysis (initial diagnosis in March 2022 → liver metastasis in May 2023), generating preliminary staging feature data.
[0025] Contextual analysis was performed on the preliminary staging feature data to eliminate textual ambiguity and map key clinical indicators to standardized features, generating structured clinical feature data. This data includes core medical record elements such as tumor type, treatment line number, stage, and ECOG score. Descriptions in the medical record such as "6 cycles of adjuvant chemotherapy after surgery" and "current physical condition is good, able to manage daily activities" were clarified through contextual analysis to correspond to the treatment line number (first-line) for "adjuvant chemotherapy after surgery" and to map "good physical condition" to an ECOG score of 0. The structured data content is as follows: Tumor type: Endometrial carcinoma (poorly differentiated adenocarcinoma); Treatment line number: First-line (adjuvant chemotherapy after surgery); Stage: Stage IVB (liver metastasis); ECOG score: 0.
[0026] Simultaneously, the OCR technology parsing module performs text recognition on paper and imaging documents, extracting key values and terms from medical records and examination reports. It then combines this with text box location information to splice and format the content, generating the original recognized text. When performing OCR recognition on images of patients' paper pathology reports, the system first uses a multimodal text recognition algorithm to locate key areas in the report (such as the content blocks below the titles "Immunohistochemistry Results" and "Gene Testing Results"). For common formats in pathology reports, such as tables and bulleted lists, it accurately extracts key values and terms scattered across different text boxes, such as "ER(+, 80%)" and "PR(+, 60%)".
[0027] Because paper reports may have issues such as compact layout, blurry text, or scanning errors, OCR technology uses the positional coordinates of text boxes (e.g., the difference between the horizontal and vertical axes) to determine content relevance. If the difference in the vertical axis coordinates of the text boxes for "ER(+, 80%)" and "PR(+, 60%)" is less than a preset threshold (e.g., 5 pixels), they are considered adjacent items and concatenated in the original order as "Immunohistochemistry: ER(+, 80%), PR(+, 60%), p53 (wild type), Ki-67 (index 70%)". For the "Gene Detection" section, the positional correlation is similarly used to concatenate "MSH2 and MSH6 protein expression is normal" and "no clear pathogenic mutations" into coherent text, ensuring that the original recognized text fully reflects the report's logic. This process preserves the original descriptions of various indicators in the pathology report and eliminates content breaks caused by scanning or layout through formatting, providing complete and coherent basic data for subsequent character-level and semantic-level verification.
[0028] The original recognized text is subject to dual character-level and semantic-level verification rules to identify and correct misidentified and omitted abnormal text. Data format standardization processing is performed to generate structured supplementary data information containing immunohistochemical information and gene detection results. When performing dual character-level and semantic-level verification on the OCR recognition results, the handling of "Ki-67 (index 70%)" reflects the character-level verification logic: the system compares the term "index" with a preset medical terminology dictionary and finds that the term "index" is often simplified in pathology reports. Combined with the format of adjacent indicators (such as ER and PR) (only the core value is retained), "index" is determined to be a redundant expression, which is a misreading of the original text layout during OCR recognition. It is corrected to "Ki-67 (70%)" to ensure that the format is consistent with other immunohistochemical indicators and meets the technical requirements of "character-level verification and correction of misidentified text" in Appendix 1.
[0029] For the semantic analysis of "normal expression of MSH2 and MSH6 proteins," a semantic-level validation rule was followed: the system invoked the "microsatellite stability (MSS)" criterion from the medical knowledge graph—when DNA mismatch repair proteins (such as MSH2 and MSH6) are expressed normally, the corresponding phenotype is microsatellite stability. By associating with this medical knowledge, the vague expression "normal protein expression" in the original text was accurately transformed into the standardized term "microsatellite stability (MSS)," preserving the core information while meeting the standardization requirements of clinical trials for gene testing results. The finally generated standardized supplementary information, through format unification (such as removing redundant terms and using standardized abbreviations) and semantic standardization (such as replacing with industry-standard terms), ensures that the immunohistochemistry and gene testing data meet the input requirements of the subsequent multimodal processing module, providing accurate and standardized feature data for the calculation of the comprehensive matching score, and ensuring consistency with the subsequent matching process.
[0030] S103 enhances structured clinical feature data and supplementary information based on a multimodal processing pipeline module, extracts image quantification indicators, analyzes immunohistochemical results, and generates a comprehensive matching score.
[0031] In one implementation, a multimodal processing pipeline module analyzes structured clinical feature data and supplementary information, extracting lesion size and density from imaging data, and analyzing protein expression levels from immunohistochemical results, assigning a unique identifier to each indicator. Liver metastasis features are extracted from the patient's abdominal enhanced CT images: maximum diameter 2.3cm (identified as IM-001), lesion density 35HU (identified as IM-002), while recording the absence of other distant metastases. Protein expression levels are extracted from the supplementary information of the structured data: ER (+, 80%, identified as IH-001), PR (+, 60%, identified as IH-002), Ki-67 (70%, identified as IH-003), and associated with microsatellite stability (MSS, identified as IH-004). The unique identifiers ensure traceability of indicators during subsequent data association and calculation.
[0032] The quantitative indicators of imaging and the results of immunohistochemical analysis were standardized. Data normalization was used to eliminate dimensional differences between different indicators and generate standardized feature data. The maximum diameter of liver metastases (2.3 cm) was converted into a 0-1 interval value using min-max standardization: the threshold for the maximum diameter of metastases in clinical trials was 5 cm, and the standardized value was 2.3 / 5 = 0.46; the lesion density (35 HU) was referenced to the density of normal liver tissue (50 HU), and the standardized value was 35 / 50 = 0.7.
[0033] Protein expression positivity rates were directly normalized to percentage values: ER (80% → 0.8), PR (60% → 0.6), Ki-67 (70% → 0.7); MSS was used as a categorical variable, assigned values of 1 (stable) and 0 (unstable). Standardized feature data were generated: IM-001 = 0.46, IM-002 = 0.7, IH-001 = 0.8, IH-002 = 0.6, IH-003 = 0.7, IH-004 = 1.
[0034] Correlation analysis was performed on standardized feature data, structured clinical feature data, and supplementary information to establish the correlation characteristics between indicators and clarify the correspondence between each indicator and the patient's condition. Specifically, the imaging indicators were found to be directly related to the maximum diameter of liver metastases (IM-001) and "advanced endometrial cancer with liver metastases," while lesion density (IM-002) was used to help determine the nature of the lesions (solid metastases).
[0035] Immunohistochemical correlations showed that ER / PR positivity (IH-001, IH-002) indicated hormone sensitivity, high Ki-67 expression (IH-003) indicated active tumor proliferation, and MSS (IH-004) excluded patients with a high risk of immunotherapy. Correlating these indicators with "Stage IVB" and "ECOG 0 score" in structured data clarified the relationship between "high proliferative activity + liver metastasis" and the risk of disease progression.
[0036] Based on correlation characteristics, a weighted algorithm was used to fuse and calculate the quantitative indicators of imaging, immunohistochemical results, and clinical characteristics to generate a preliminary matching score. Weights were assigned based on key aspects of the endometrial cancer clinical trial: liver metastasis size (20%), ER expression (25%), PR expression (20%), Ki-67 (20%), and MSS (15%). The preliminary matching score was calculated as follows: (0.46 × 20%) + (0.8 × 25%) + (0.6 × 20%) + (0.7 × 20%) + (1 × 15%) = 0.092 + 0.2 + 0.12 + 0.14 + 0.15 = 0.702 (out of 1 point).
[0037] The initial matching score was verified and adjusted. The score deviation was corrected by incorporating the weighting of the clinical trial's inclusion criteria to generate a final comprehensive matching score, which measures the degree of match between the patient and the clinical trial. The inclusion criteria weighting was adjusted so that a certain endometrial cancer trial had a higher requirement for "ER / PR positive" (weight increased to 30%) and a lower weight for "MSS" (5%). The corrected score was recalculated as (0.46 × 20%) + (0.8 × 30%) + (0.6 × 20%) + (0.7 × 20%) + (1 × 5%) = 0.092 + 0.24 + 0.12 + 0.14 + 0.05 = 0.642. The comprehensive matching score of 0.642 indicates a 64.2% match between the patient and the trial, which can be used for subsequent trial prioritization.
[0038] S104 uses a rule engine and semantic similarity calculation to compare patient features with in and out criteria item by item, generating preliminary matching results.
[0039] In one implementation, a rule engine is used to perform a structured comparison between patient characteristics and clinical trial inclusion / exclusion criteria. This extracts the patient's tumor type, stage, and treatment line number, and matches these against the pre-defined criteria of the clinical trial, generating a matching label for each criterion. Taking a patient with advanced endometrial cancer as an example, the structured clinical feature data extracts "tumor type: endometrial cancer (poorly differentiated adenocarcinoma)," "stage: stage IVB (liver metastasis)," and "treatment line number: second-line (proposed third-line treatment)." These features, through preprocessing using a large language model and standardization by the OCR parsing module, have been transformed into directly comparable structured data.
[0040] The pre-defined criteria for a selected phase III trial of advanced endometrial cancer were based on the fundamental requirements of "tumor specificity, disease progression, and treatment history" in clinical trial enrollment criteria, consistent with the structured screening logic of the "20 pre-defined criteria" in Appendix 1. Specifically, "Criterion 1: Tumor type is endometrial cancer" directly corresponds to the patient's tumor type, ensuring a match between the trial and the patient's disease; "Criterion 2: Stage IIIB or higher (including stage IV)" covers the patient's stage IVB disease progression, meeting the requirements for disease stage in advanced tumor trials; and "Criterion 3: Second-line or higher treatment" matches the patient's current treatment stage, satisfying the enrollment criteria for later-line treatment trials. The generation of the "+" label in the matching results follows the rule engine's principle of "a complete match between features and conditions indicates compliance," providing a structured foundation for subsequent semantic similarity calculations and association matrix construction.
[0041] The semantic similarity calculation module processes unstructured patient features and clinical trial-specific conditions. It uses a large language model to parse the semantic connotations of inclusion and exclusion conditions, calculates the semantic matching degree between patient medical record descriptions and conditional texts, and generates a similarity score. For example, the patient's medical record statement, "The patient previously received albumin-bound paclitaxel + carboplatin chemotherapy, with no serious adverse reactions; currently, the ECOG score is 0, and liver metastases are stable," is unstructured text. Phrases such as "no serious adverse reactions" are not directly covered by structured features and require semantic analysis to uncover potential information.
[0042] Among the selected clinical trial-specific criteria, "no previous chemotherapy with grade 3 or higher hematologic toxicity" and "presence of uncontrollable ascites" are trial-specific requirements not covered by the pre-defined criteria. The large language model first analyzes the semantic connotation of the inclusion criteria, transforming "grade 3 or higher hematologic toxicity" into the standardized criterion of "chemotherapy toxicity ≤ grade 2," and then compares it semantically with the patient's "no serious adverse reactions"—since "no serious adverse reactions" in clinical context usually implies "toxicity reaction ≤ grade 2," a matching degree of 90% (score 0.9) is calculated. For the exclusion criteria, the model analyzes its core as "ascites requiring repeated drainage," but the patient's medical record does not mention anything related to ascites. Based on the logic of "if not mentioned, it is assumed that the exclusion criteria are not met," the matching degree is 100% (score 1.0). By leveraging the semantic understanding capabilities of large language models, the limitations of structured comparison in handling ambiguous expressions are overcome. The generated similarity score is combined with the matching labels from the preceding structured matching, providing quantitative support for the subsequent construction of the association matrix and ensuring consistency with the preceding and following processes.
[0043] By combining pre-defined criteria matching labels and semantic similarity scores, an association matrix is established between patient characteristics and clinical trial conditions to clarify the matching relationship between each characteristic and condition. This association matrix, based on pre-defined criteria matching labels and semantic similarity scores, must clearly present the matching relationship between patient characteristics and clinical trial conditions in text form. Taking an endometrial cancer patient as an example, in the tumor type dimension, the patient's endometrial cancer corresponds to the trial requirement of "tumor type is endometrial cancer," and the pre-defined criteria matching label is "+", indicating a complete match and a direct correspondence of structured features. In the stage dimension, the patient's stage IVB meets the trial's condition of "stage IIIB or above (including stage IV)," and the pre-defined criteria matching label is "+", reflecting a matching of structured features. In the treatment line dimension, the patient is currently undergoing second-line treatment and plans to undergo third-line treatment, meeting the trial's requirement of "second-line or above treatment," and the pre-defined criteria matching label is "+," again demonstrating a matching of structured features.
[0044] Regarding the unstructured characteristic of chemotherapy toxicity, the trial required "no previous chemotherapy with grade 3 or higher hematological toxicity." The description of "no serious adverse reactions" in the patient's medical record, after semantic analysis, highly matched the criterion of "chemotherapy toxicity ≤ grade 2," with a semantic similarity score of 0.9. Regarding ascites, the trial excluded "uncontrollable ascites," but the patient's medical record did not mention this, so it was assumed not to meet the exclusion criteria, with a semantic similarity score of 1.0, reflecting the semantic association between the unstructured characteristic and the trial-specific conditions. Through such textual descriptions, it is clear that the patient's structured characteristics all met the basic inclusion criteria of the trial, and the unstructured characteristics also achieved a high degree of matching through semantic analysis, providing a comprehensive quantitative basis for the subsequent generation of the preliminary matching list.
[0045] Based on the association matrix, patients and clinical trials are initially matched. Trial protocols that meet the preset criteria and achieve the required semantic similarity are retained, generating a preliminary matching list. Taking endometrial cancer patients as an example, the screening criteria are clearly defined as "all structured conditions are met (label '+')" and "semantic similarity score ≥ 0.8". The structured conditions correspond to core dimensions such as tumor type, stage, and number of lines of treatment, and must completely match the preset criteria of the clinical trial (i.e., all labels are "+"). The semantic similarity score targets unstructured features (such as chemotherapy toxicity, ascites, etc.) and must reach a "high match" threshold (≥ 0.8). For a phase III trial of advanced endometrial cancer, the patient's structured features (endometrial cancer, stage IVB, second-line treatment) all meet the trial's preset criteria (all labels are "+"), and the semantic similarity scores of unstructured features (chemotherapy toxicity 0.9, ascites 1.0) both exceed the threshold of 0.8, meeting the dual screening criteria, and therefore were included in the preliminary matching list.
[0046] Another trial, which required "only one chemotherapy regimen received," failed the first-level screening and was directly excluded because patients had actually received two or more chemotherapy regimens, such as albumin-bound paclitaxel + carboplatin or sintilimab. The "number of treatment regimens" in the structured criteria did not match (labeled "-"). By strictly enforcing the screening criteria, trials meeting the core inclusion criteria were retained, while specific conditions with low matching scores were filtered out using semantic similarity thresholds, providing a reliable candidate list for subsequent validation.
[0047] The preliminary matching list is validated, filtering out suspected mismatches due to missing or ambiguous data, and generating preliminary matching results containing matching items, non-matching items, and the basis for judgment. For the preliminary matching list of endometrial cancer patients, the validation focuses on two dimensions: Data completeness: verifying whether the patient's core information is complete, such as the clear recording of "KRASG12C mutation" in the gene testing results, with no missing key indicators, meeting the basic requirements of clinical trials for biomarker information, and avoiding misjudgments due to incomplete data. Logical consistency: verifying the inherent consistency between structured features and unstructured descriptions, ensuring that the structured labeling of the patient's "second-line treatment" and the medical record description of "proposed third-line treatment" are consistent (the former is the current treatment stage, and the latter is the next step plan), and that the semantic similarity scores (chemotherapy toxicity 0.9, ascites 1.0) completely correspond to the original content in the medical record of "no serious adverse reactions" and "no mention of ascites," with no logical conflicts.
[0048] The preliminary matching results generated after verification clearly distinguish between two types of trials: Matching: A stage III trial of advanced endometrial cancer was judged to be a match because it met all three structured criteria (tumor type, stage, and number of lines of treatment) (labeled "+") and had a semantic similarity score ≥0.8. This was based on a direct correlation with the aforementioned association matrix and screening criteria. Disqualifying: A trial that only included patients receiving first-line treatment was explicitly listed as a disqualifying trial because the patients were currently receiving second-line treatment and the structured criteria for the number of lines of treatment did not match (labeled "-"). This judgment directly corresponds to the structured comparison results. This verification process, by eliminating the risk of missing data and logical contradictions, provides high-quality input data for the subsequent large language model optimization module, ensuring the rigor of the matching results.
[0049] S105, based on the large language model optimization module, corrects misjudgments of edge cases through context-aware multi-round reasoning, and generates an optimized clinical trial matching list by adjusting the ranking based on clinical trial priority weights.
[0050] In one implementation, a large language model optimization module analyzes marginal cases in the initial matching results, extracts fuzzy matching items between patient characteristics and clinical trial conditions, corrects misjudgments through context-aware multi-round reasoning, and generates corrected matching labels. In the initial matching results, a patient with endometrial cancer was excluded by a trial due to the medical record stating "abdominal effusion (undrained) in the past month" (the trial's exclusion criterion was "uncontrollable ascites"). The large language model optimization module, through context-aware multi-round reasoning, analyzes the semantic connotation of "undrained"—combining clinical common sense, undrained ascites usually does not meet the "uncontrollable" criterion, thus correcting the original "does not meet" label to "meets" (i.e., does not trigger the exclusion criterion), and generating corrected matching labels.
[0051] Priority weighting parameters are set for clinical trials, covering dimensions such as trial stage, indication matching degree, and potential benefit, and a weight value is assigned to each trial protocol. Priority weighting is set for the three endometrial cancer clinical trials included in the preliminary matching list, covering three dimensions: Trial stage (Phase I 10%, Phase II 20%, Phase III 30%): reflecting trial maturity, with Phase III trials having the highest weight due to more comprehensive data; Indication matching degree (maximum 100%): assigned based on the degree of similarity in tumor type, stage, and genetic characteristics, such as a trial specifically targeting "advanced endometrial cancer with KRASG12C mutation," with a matching degree of 90%; Potential benefit (maximum 100%): assessed in conjunction with drug mechanism and patient genetic characteristics, such as a targeted drug trial assigning an 80% potential benefit value to patients with KRASG12C mutation.
[0052] Ultimately, each trial was assigned a weighted average: Trial A (Phase III, 90% indication matching, 80% potential benefit) had a weight of 30% × 0.3 + 90% × 0.5 + 80% × 0.2 = 73%; Trial B (Phase II, 85% indication matching, 70% potential benefit) had a weight of 20% × 0.3 + 85% × 0.5 + 70% × 0.2 = 65.5%; and Trial C (Phase I, 80% indication matching, 60% potential benefit) had a weight of 10% × 0.3 + 80% × 0.5 + 60% × 0.2 = 55%.
[0053] By combining the revised matching labels and priority weights, a correlation model between matching results and weights is established to clarify the basis for adjusting each trial protocol. The revised matching labels are associated with the weight values to clarify the ranking logic. Specifically, for Trial A: the revised label is "Matching", weight 73% → priority 1, based on "Phase III trial + high indication matching degree + significant potential benefit"; for Trial B: the label is "Matching", weight 65.5% → priority 2, based on "Phase II trial + relatively high indication matching degree"; for Trial C: the label is "Matching", weight 55% → priority 3, based on "Phase I trial + moderate potential benefit".
[0054] The preliminary matching results are sorted and adjusted based on the association model, prioritizing high-weighted trial protocols that meet the revised criteria, generating a sorted candidate list. Trials are sorted from highest to lowest weight, retaining those that meet the revised criteria, generating the following candidate list: Trial A (Phase III targeted therapy, 92% match, priority 1); Trial B (Phase II immunotherapy, 88% match, priority 2); Trial C (Phase I chemotherapy combination, 80% match, priority 3).
[0055] The candidate list undergoes final verification, filtering out duplicates and protocols that do not meet the core inclusion criteria, generating an optimized clinical trial matching list that includes trial name, matching degree, and priority. The candidate list is double-verified: duplicate filtering (no duplicate trials, no processing required); core inclusion criterion verification (confirming all trials meet the core criteria such as "advanced endometrial cancer, second-line or later treatment, KRASG12C mutation"). The final optimized matching list is generated as follows: Trial A (Phase III targeted therapy): matching degree 92%, priority 1, key criteria are "meets the exclusion criteria after correction, highest weight, and complete indication match"; Trial B (Phase II immunotherapy): matching degree 88%, priority 2, key criteria are "high indication match, second highest trial maturity".
[0056] The revised labels resolve misjudgments caused by semantic ambiguity in the initial matching (such as the degree definition of "ascites"); the priority weights combine the semantic analysis results of the "specific condition matching model" to ensure that the ranking is consistent with the actual value of clinical trials; the final list provides a precise trial recommendation order for the subsequent generation of matching reports, which conforms to the "patient-centered" matching logic.
[0057] S106 inputs structured clinical feature data, supplementary information, preliminary matching results, and optimized clinical trial matching list into the clinical decision support module, visualizes the entire matching path, and generates a clinical trial matching report that includes matching confidence level, key evidence, and risk warnings.
[0058] In one implementation, the clinical decision support module processes structured clinical feature data, supplementary information, preliminary matching results, and an optimized clinical trial matching list. A data association algorithm integrates multi-source data to construct a matching relationship map between patient characteristics and clinical trials, including a core set of patient features, a set of key trial conditions, and quantifiable matching values. Specifically, the clinical decision support module integrates the structured clinical feature data (tumor type: stage IVB endometrial cancer, treatment line: second-line, ECOG score 0), supplementary information (immunohistochemistry ER+80%, KRASG12C mutation), preliminary matching results (3 trials meeting the core conditions), and an optimized matching list (trial A priority 1, trial B priority 2) of endometrial cancer patients. Through data association algorithms, the core characteristics of patients (such as "KRASG12C mutation" and "liver metastasis") are associated with key trial conditions (such as "second-line treatment for advanced endometrial cancer" and "KRASG12C mutation positive"), generating a relationship map containing matching quantification values. For example, the matching degree between trial A and the patient is 92%, corresponding to association nodes such as "targeted drug matches KRASG12C mutation" and "complete staging match".
[0059] The system employs visualization rendering technology to present the entire matching process, displaying the complete workflow from data extraction to optimized sorting in a timeline format. Interactive nodes allow users to review the judgment criteria for each step and generate a dynamic matching path map. Specifically, the process begins at 0 minutes: OCR recognition of medical record images extracts core features such as "endometrial cancer with liver metastasis, second-line treatment"; at 3 minutes: a rule engine performs structured comparison, generating a label that "tumor type, stage, and treatment line all match"; at 5 minutes: semantic similarity calculation yields "chemotherapy toxicity matching score 0.9, no ascites matching score 1.0"; at 8 minutes: a large language model corrects edge cases (e.g., "no drainage of ascites ≠ uncontrollable"); at 10 minutes: sorting by priority weights generates an optimized list. Interactive buttons are provided at each node of the timeline; clicking "semantic similarity calculation" allows users to review the semantic alignment criteria between "no serious adverse reactions" and "chemotherapy toxicity ≤ grade 2".
[0060] A report generation matrix was constructed, and the confidence parameters of each matching indicator were calculated by combining the overall matching score and the clinical priority weight. The matching confidence was determined using a Bayesian probability-corrected model, and key evidence was highlighted using a text highlighting algorithm. Combining the overall matching score (Trial A 64.2 points) and the clinical priority weight (30% weight for Phase III trials), the report generation matrix was constructed: For matching confidence, a Bayesian probability-corrected model was used, and combined with historical matching data (the accuracy rate for enrolling similar patients in Trial A was 91%), the matching confidence of Trial A was calculated to be 94%.
[0061] By using a text highlighting algorithm, the medical records are highlighted with information that directly matches the inclusion criteria of Trial A, such as "KRASG12C mutation" and "liver metastasis". The trial protocol is highlighted with key conditions such as "targeting KRASG12C" and "second-line or above treatment" to ensure traceability of the evidence.
[0062] Based on the generated report and pre-defined risk assessment criteria, and considering the characteristics of the trial type, a clinical trial matching report is generated, including matching confidence level, key matching evidence, and risk warnings. The risk warnings are annotated using the clinical trial adverse event database. Based on pre-defined risk assessment criteria (e.g., "common adverse reactions to targeted drugs: rash"), and considering the characteristics of the trial type (Phase I / II / III), the final report is generated as follows: Matching confidence level: Trial A 94%, Trial B 88%; Key evidence: Trial A is the preferred choice due to "KRASG12C mutation compatibility + mature Phase III data"; Risk warnings: The clinical trial adverse event database is cited, indicating that Trial A may have "grade 3 rash (incidence 12%)", and Trial B requires vigilance for "immune-associated pneumonia (incidence 8%)". The report also notes that "final enrollment requires verification of complete blood count and imaging examinations," meeting the requirement of "further testing after initial screening" in the study design, providing a complete reference for clinical decision-making.
[0063] This application improves the efficiency and accuracy of clinical trial matching for cancer patients through multi-module collaborative processing. First, a three-dimensional knowledge graph of "tumor type-stage-treatment line number" is constructed, integrating guideline and literature data to upgrade static judgment conditions to dynamic association rules. Clinical data is analyzed using a large language model, combined with a temporal neural network for precise staging, generating structured clinical feature data. Simultaneously, OCR technology is used to process paper and imaging data, and supplementary information is generated after double verification.
[0064] The multimodal processing pipeline module enhances data, extracts quantitative indicators from images, analyzes immunohistochemical results, and generates a comprehensive matching score. A rule engine and semantic similarity calculation perform item-by-item comparisons, generating preliminary matching results. The large language model optimization module corrects misjudgments, adjusts the ranking based on priority weights, and generates an optimized list. The clinical decision support module integrates data, visualizes the matching path, constructs a report generation matrix, and ultimately generates a report containing matching confidence levels, key evidence, and risk warnings, providing an efficient and accurate tool for recruiting candidates for oncology clinical trials.
[0065] In one implementation, such as Figure 2 As shown, this application also provides a tumor patient clinical trial matching device based on a large language model and OCR technology, comprising:
[0066] The acquisition module 201 is used to acquire clinical data and clinical trial information of cancer patients. Based on the knowledge graph construction module, it performs preprocessing to construct a three-dimensional knowledge graph of "tumor type-stage-number of treatment lines", integrates NCCN / CSCO guidelines and real-time literature data, and upgrades static judgment conditions to dynamic association rules.
[0067] The processing module 202 is used to parse clinical data and trial information based on the large language model preprocessing module, process unstructured text, accurately stage the data by combining a temporal neural network and the RECIST 1.1 standard time axis analysis model, and generate structured clinical feature data through context association analysis. Simultaneously, it processes paper and image data through the OCR technology parsing module, extracts key data and corrects it through double verification, and generates supplementary structured data information. Based on the multimodal processing pipeline module, it enhances the structured clinical feature data and supplementary information, extracts image quantitative indicators, parses immunohistochemical results, and generates a comprehensive matching score. Through the rule engine and semantic similarity calculation, it compares patient characteristics with inclusion and exclusion conditions item by item to generate preliminary matching results. Based on the large language model optimization module, it corrects the misjudgment of marginal cases through context-aware multi-round reasoning, and adjusts the ranking by combining the priority weight of clinical trials to generate an optimized clinical trial matching list. The structured clinical feature data, supplementary information, preliminary matching results, and optimized clinical trial matching list are input into the clinical decision support module to visualize the entire matching path and generate a clinical trial matching report including matching confidence, key evidence, and risk warnings.
[0068] The various embodiments in this application are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for evaluating the matching method, electronic device, electronic device, and readable storage medium for tumor patients based on large language models and OCR technology are basically similar to the above-described embodiments of the matching method for tumor patients based on large language models and OCR technology, and therefore are described relatively simply. Relevant parts can be referred to in the descriptions of the above-described embodiments of the matching method for tumor patients based on large language models and OCR technology.
Claims
1. A method for matching a tumor patient with a clinical trial based on a large language model and OCR technology, characterized in that, Comprise: Obtain the clinical data and clinical trial information of tumor patients, preprocess based on the knowledge graph construction module, construct a "tumor type-staging-treatment line number" three-dimensional knowledge graph, integrate NCCN / CSCO guidelines and real-time literature data, and upgrade static judgment conditions to dynamic correlation rules; Based on the large language model preprocessing module, analyze the clinical data and trial information, process unstructured text, combine the time series neural network with the RECIST1.1 standard time axis analysis model to accurately stage, and generate structured clinical feature data through context correlation analysis; At the same time, process paper and image data through the OCR technology analysis module, extract key data and correct it through double verification, and generate structured data supplement information; Based on the multi-modal processing pipeline module, enhance the structured clinical feature data and supplementary information, extract image quantitative indicators, and analyze immunohistochemical results to generate a comprehensive matching score, including standardizing image quantitative indicators and immunohistochemical analysis results, eliminating the dimensional differences between different indicators through data normalization to generate standardized feature data; Correlation analysis of standardized feature data, structured clinical feature data, and supplementary information establishes the correlation between indicators, and clarifies the corresponding relationship between each indicator and the patient's condition; Based on the correlation characteristics, a weighted algorithm is used to fuse and calculate the image quantitative indicators, immunohistochemical results, and clinical characteristics to generate a preliminary matching score; Verify and adjust the preliminary matching score, correct the score deviation by combining the enrollment condition weight of the clinical trial, and generate the final comprehensive matching score to measure the matching degree of the patient and the clinical trial; Through the rule engine and semantic similarity calculation, the patient's features and the enrollment and exclusion conditions are compared item by item to generate a preliminary matching result; Based on the large language model optimization module, edge case misjudgments are corrected through context-aware multi-round reasoning, and the ranking is adjusted by combining the priority weight of the clinical trial to generate an optimized clinical trial matching list; Input the structured clinical feature data, supplementary information, preliminary matching result, and optimized clinical experiment matching list into the clinical decision support module to visualize the matching path and generate a clinical trial matching report containing matching confidence, key evidence, and risk prompts.
2. The method of claim 1, wherein, Based on the large language model preprocessing module, analyze the clinical data and trial information, process unstructured text, combine the time series neural network with the RECIST1.1 standard time axis analysis model to accurately stage, and generate structured clinical feature data through context correlation analysis; At the same time, process paper and image data through the OCR technology analysis module, extract key data and correct it through double verification, and generate structured data supplement information, including: Based on the large language model preprocessing module, analyze the non-structured text in the clinical data and trial information, extract the patient's condition description and treatment history, combine the time series neural network with the RECIST1.1 standard time axis analysis model to accurately stage the patient's condition, and generate preliminary staging feature data; Contextual correlation analysis of preliminary staging feature data, eliminating text ambiguity, mapping key clinical indicators to standardized features, generating structured clinical feature data, including core medical record elements such as tumor type, treatment line number, stage, and ECOG score; Synchronously analyze paper and image data through the OCR technology recognition module to extract key values and terms from medical records and examination reports, and combine text box position information to perform content splicing and format arrangement to generate original recognition text; Apply character-level and semantic-level double verification rules to the original recognition text to identify and correct abnormal text, perform data format standardization processing, and generate structured data supplement information including immunohistochemistry information and gene test results.
3. The method of claim 2, wherein, Based on the multi-modal processing pipeline module, structured clinical feature data and supplementary information are enhanced to extract image quantitative indicators and analyze immunohistochemistry results, generating comprehensive matching scores, including: Based on the multi-modal processing pipeline module, structured clinical feature data and supplementary information are analyzed to extract lesion size and density from image data and analyze protein expression levels from immunohistochemistry results, assigning a unique identifier to each type of indicator.
4. The method of claim 1, wherein, Through rule engine and semantic similarity calculation, the patient's features are compared with the inclusion and exclusion criteria to generate preliminary matching results, including: Based on the rule engine, the patient's features are compared with the clinical trial inclusion and exclusion criteria to extract the patient's tumor type, stage, and treatment line number, and match them with the pre-set judgment conditions to generate a compliance label for each judgment condition. The semantic similarity calculation module processes the patient's unstructured features and the specific conditions of the clinical trial. The semantic content of the inclusion and exclusion conditions is analyzed by a large language model, and the semantic similarity between the patient's medical record description and the condition text is calculated to generate a similarity score. Based on the pre-set judgment condition compliance label and the semantic similarity score, an association matrix of patient features and clinical trial conditions is established to determine the matching relationship between each feature and condition. Based on the association matrix, the patient and the clinical trial are preliminarily matched, and the trial plan that meets the pre-set judgment conditions and has a satisfactory semantic similarity is retained to generate a preliminary matching list. The preliminary matching list is verified to filter out suspected mispairing items caused by data missing or ambiguity, and a preliminary matching result containing matching items, non-compliance items, and judgment basis is generated.
5. The method of claim 1, wherein, Based on the large language model optimization module, edge case misjudgment is corrected through context-aware multi-round reasoning, and the order is adjusted based on the priority weight of the clinical trial to generate an optimized clinical trial matching list, including: Based on the large language model optimization module, the edge case in the preliminary matching result is analyzed to extract the ambiguous matching items in the patient's features and the clinical trial conditions, and the misjudgment is corrected through context-aware multi-round reasoning to generate a corrected matching label. Set priority weight parameters for the clinical trial, including trial phase, indication matching degree, and potential benefit dimensions, and assign a weight value to each trial plan. Based on the corrected matching label and the priority weight value, a matching result and weight association model is established to determine the adjustment basis for each trial plan. Based on the correlation model, the preliminary matching results are sorted and adjusted, high-weight and corrected test schemes that meet the conditions are preferentially retained, and a sorted candidate list is generated; The candidate list is finally checked to filter out duplicate items and schemes that do not meet the core enrollment conditions, and an optimized clinical trial matching list containing the test name, matching degree, and priority is generated.
6. The method of claim 5, wherein, Input the structured clinical feature data, supplementary information, preliminary matching results, and optimized clinical trial matching list into the clinical decision support module, visualize the matching full path, and generate a clinical trial matching report containing matching confidence, key basis, and risk prompts, including: Based on the clinical decision support module, the structured clinical feature data, supplementary information, preliminary matching results, and optimized clinical trial matching list are processed, multi-source data is integrated through data correlation algorithms, and a matching relationship graph between patient characteristics and clinical trials is constructed, which contains the patient core feature set, the test key condition set, and the matching quantitative value; Use visual rendering technology to present the matching full path, display the complete process from data extraction to optimized sorting in the form of a time axis, set interactive nodes to allow backtracking to view the judgment basis at each link, and generate a dynamic matching path graph; Build a report generation matrix, combine the comprehensive matching score and the clinical priority weight, calculate the confidence parameters of each matching indicator, and use the Bayesian probability correction model for matching confidence and the text highlighting algorithm for key basis; Based on the report generation results and the preset risk assessment standard, combined with the characteristics of the test type, a clinical trial matching report containing matching confidence, key matching basis, and risk prompts is generated, and the risk prompts refer to the clinical trial adverse event database for correlation annotation. 7.A tumor patient clinical trial matching device based on a large language model and an OCR technology, characterized in that, The device for implementing the method of claim 1 comprises: An acquisition module is configured to acquire clinical data of tumor patients and clinical trial information, pre-process based on a knowledge graph construction module, construct a "tumor type-stage-treatment line number" three-dimensional knowledge graph, integrate NCCN / CSCO guidelines and real-time literature data, and upgrade static judgment conditions to dynamic correlation rules. The processing module is used for analyzing clinical data and test information based on a large language model preprocessing module, processing unstructured text, combining a time sequence neural network and a RECIST1.1 standard time axis analysis model for accurate staging, and generating structured clinical feature data through context correlation analysis; paper and image data are processed by an OCR technology analysis module to extract key data and correct them through double verification to generate structured data supplement information; the structured clinical feature data and the supplement information are enhanced based on a multi-modal processing pipeline module to extract image quantitative indicators and analyze immunohistochemical results, and a comprehensive matching score is generated; through a rule engine and semantic similarity calculation, item-by-item comparison of patient features and entry and exit conditions is realized to generate a preliminary matching result; based on a large language model optimization module, edge case misjudgments are corrected through context-aware multi-round reasoning, and the ordering is adjusted in combination with clinical trial priority weight to generate an optimized clinical trial matching list; the structured clinical feature data, the supplement information, the preliminary matching result and the optimized clinical trial matching list are input into a clinical decision support module to visually display the matching full path and generate a clinical trial matching report containing matching confidence, key basis and risk prompts.
8. An electronic device, comprising: It comprises: a first processor; and a memory for storing executable instructions of the first processor; wherein the first processor is configured to execute the executable instructions to perform the tumor patient clinical trial matching method based on a large language model and an OCR technology according to any one of claims 1-6.
Citation Information
Patent Citations
Clinical test matching method, device, equipment, medium and product
CN119153008A
Clinical auxiliary decision-making system based on big data
CN120108696A