A method and system for quality control of clinical research implementation based on large language model
By processing clinical research data through large language models, generating an external knowledge base and building quality control reports, the inefficiency problem of traditional methods is solved, efficient data integration and diagnostic recommendations are achieved, and research quality and process optimization are improved.
Patent Information
- Application Number
- CN202510118642.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-01-24
AI Technical Summary
Traditional quality control methods for clinical research are inefficient, making it difficult to effectively integrate data from institutions in different regions and medical levels. There is a lack of efficient tools to generate accurate quality control reports and diagnostic recommendations, which affects the reliability and efficiency of research results.
A large language model is used to filter, format, segment, vectorize and classify target clinical research project information, user consultation questions and scientific research project information from different institutions to generate an external knowledge base. Quality control reports are generated by constructing text block vectors, similarity indexes and medical knowledge graphs, and diagnostic correlation analysis is performed by combining clinical research management influencing factors and user individual information influencing factors.
It improves the quality of clinical research, optimizes research processes, provides accurate quality control reports and diagnostic recommendations, and enhances data integration and management efficiency.
Smart Images

Figure CN120048408B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a method and system for quality control of clinical research implementation based on a large language model. Background Art
[0002] In the field of clinical research, traditional research implementation quality control has many shortcomings. With the development of medical technology and the increasing complexity of clinical research, the research process involves a large amount of patient data, a variety of research projects, and the comparison and integration of information from different regions and medical institutions.
[0003] On the one hand, clinical research needs to process massive amounts of patient data, including personal information, medical history, treatment records, and follow-up status in case report forms. The collection, organization, and analysis of this data are arduous and prone to errors. At the same time, research participants often have limited time for scientific research, making it difficult to conduct comprehensive and effective management of the patients under study, resulting in uneven data quality and affecting the reliability and validity of research results. On the other hand, there is a lack of effective integration and utilization mechanisms for similar clinical research data conducted by institutions in different regions and medical levels. These data contain a wealth of information, such as the differences in the effectiveness of different treatment options in different populations, and the experience in dealing with special cases. However, due to the dispersion of information, it is difficult to extract valuable references from them, and it is impossible to provide comprehensive background support and risk warnings for current research.
[0004] Furthermore, when problems arise during research, there is a lack of efficient tools to quickly generate accurate quality control reports and effective diagnostic recommendations. Traditional methods, which may rely on manual experience and cumbersome statistical analysis, are inefficient and fail to fully cover a wide range of complex situations, failing to meet the dual requirements of clinical research for quality and efficiency.
[0005] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the Invention
[0006] The purpose of this application is to provide a method and system for quality control of clinical research implementation based on a large language model, which at least to a certain extent overcomes the problems existing in the prior art, by obtaining target clinical research projects, user consultation questions and scientific research project information of different institutions. Subsequently, the target clinical project information is filtered, formatted, segmented, vectorized, classified and associated to form an external knowledge base; clinical research management impact factors and user individual information impact factors are generated for the scientific research information of the target hospital and other institutions respectively. The former analyzes the classification progress, etc., and the latter is obtained by feature extraction and clustering. Then, using the target large language model and the external knowledge base, a quality control report is generated by constructing a text block vector, a similarity index and a medical knowledge graph. Finally, the two influencing factors are combined to process the relevant information to generate accurate diagnostic correlation, which effectively helps improve the quality of clinical research and optimize the process.
[0007] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by practice of the invention.
[0008] According to one aspect of the present application, a method for quality control of clinical research implementation based on a large language model is provided, comprising: obtaining target clinical research project information, consulting questions of target users regarding clinical research, scientific research project management information of a target hospital, other clinical research project information, a preset large language model, and a training sample set, wherein the target clinical research project information is updated in real time, the scientific research project information of the target hospital is used to characterize the scientific research project classification information, management requirement information, and progress information of each scientific research project of the target hospital, and the other clinical research project information includes similar target clinical research data of institutions in different regions and at different medical levels; processing the preset large language model based on the training sample set, Generate a target large language model; process the target clinical research project information to generate an external knowledge base; process the scientific research project information of the target hospital to generate a clinical research management impact factor; process the other clinical research project information to generate a user individual information impact factor; based on the target large language model and the external knowledge base, process the target user's consulting questions regarding clinical research to generate a quality control report; based on the clinical research management impact factor and the user individual information impact factor, process the target clinical research project information and the quality control report to generate a diagnostic accuracy correlation, wherein the diagnostic accuracy correlation is used to characterize the improvement of research quality and the improvement of research process.
[0009] Another aspect of the present application is a clinical research implementation quality control device based on a large language model, characterized in that it includes: an acquisition module for acquiring target clinical research project information, target users' consulting questions for clinical research, target hospital's scientific research project management information, other clinical research project information, a preset large language model, and a training sample set, wherein the target clinical research project information is updated in real time, the target hospital's scientific research project information is used to characterize the target hospital's scientific research project classification information, management requirements information and progress information of each scientific research project, and the other clinical research project information includes similar target clinical research data from institutions in different regions and with different medical levels; a processing module for processing the preset large language model based on the training sample set. The target large language model is processed to generate a target large language model; the target clinical research project information is processed to generate an external knowledge base; the scientific research project information of the target hospital is processed to generate a clinical research management influencing factor; the other clinical research project information is processed to generate a user individual information influencing factor; based on the target large language model and the external knowledge base, the consulting questions of the target user regarding clinical research are processed to generate a quality control report; based on the clinical research management influencing factor and the user individual information influencing factor, the target clinical research project information and the quality control report are processed to generate a diagnostic accuracy correlation, wherein the diagnostic accuracy correlation is used to represent the improvement of research quality and the improvement of research process.
[0010] According to another aspect of the present application, an electronic device is characterized in that it includes: a first processor; and a memory for storing executable instructions of the first processor; wherein the first processor is configured to execute the above-mentioned large language model-based clinical research implementation quality control method by executing the executable instructions.
[0011] According to another aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a second processor, the method for quality control of clinical research implementation based on a large language model is implemented.
[0012] The present application provides a method and system for quality control of clinical research implementation based on a large language model, in which the server obtains target clinical research projects, user consultation questions, and scientific research project information from different institutions. Subsequently, the target clinical project information is filtered, formatted, segmented, vectorized, classified, and associated to form an external knowledge base; clinical research management impact factors and user individual information impact factors are generated for the target hospital and other institutional scientific research information, respectively. The former analyzes the classification progress, etc., and the latter is obtained through feature extraction and clustering. Then, using the target large language model and the external knowledge base, a quality control report is generated by constructing text block vectors, similarity indexes, and medical knowledge graphs. Finally, the two influencing factors are combined to process relevant information to generate accurate diagnostic correlation, which effectively helps improve the quality of clinical research and optimize processes.
[0013] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 A flowchart illustrating a method for quality control of clinical research implementation based on a large language model provided by an embodiment of the present application is shown;
[0015] Figure 2 A schematic diagram of the structure of a clinical research implementation quality control device based on a large language model provided in one embodiment of the present application is shown;
[0016] Figure 3 A prompt design roadmap based on retrieval enhancement generation technology provided by an embodiment of the present application is shown. DETAILED DESCRIPTION
[0017] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0018] The following combination Figure 1 To describe the quality control method for clinical research implementation based on a large language model according to an exemplary embodiment of the present application. It should be noted that the following application scenarios are only shown to facilitate understanding of the spirit and principles of the present application, and the embodiments of the present application are not limited in this respect. On the contrary, the embodiments of the present application are applicable to any applicable scenario.
[0019] In one embodiment, the present application also proposes a method and system for quality control of clinical research implementation based on a large language model. Figure 1 The following schematically shows a flow chart of a method for quality control of clinical research based on a large language model according to an embodiment of the present application. Figure 1 As shown, the method is applied to the server and includes:
[0020] S101, obtain target clinical research project information, target users' consulting questions regarding clinical research, target hospital's scientific research project management information, other clinical research project information, preset large language model, and training sample set.
[0021] In one embodiment, it is assumed that a research project on "Clinical research on new targeted therapeutic drugs for breast cancer" is underway. The project information is as follows: Project design: A randomized controlled trial design is adopted, and patients are divided into an experimental group and a control group. The experimental group receives new targeted drug treatment, and the control group receives traditional treatment. Implementation details: The dosage, frequency of use, treatment cycle, etc. of the drug are specified in detail. For example, patients in the experimental group take 100 mg of the new targeted drug orally every day for 28 consecutive days as a cycle, and a total of 4 cycles of treatment are performed. During this period, various physical examinations and laboratory tests are required to be performed regularly to evaluate the safety and effectiveness of the drug. Recruitment criteria: The patient's age range (such as 18-70 years old), pathological type of breast cancer (such as HER2-positive breast cancer), disease stage (such as stage II-III), previous treatment history (such as no similar targeted treatment) and other conditions are clearly specified to ensure the inclusion of suitable research subjects. Interventions: In addition to drug treatment, these also include dietary and exercise guidance for patients. For example, patients are advised to maintain a balanced diet during treatment, increase their protein intake appropriately, and engage in moderate aerobic exercise at least three times a week for more than 30 minutes each time to improve their body's immunity and tolerance to treatment. Outcome Assessment: The primary outcome measures are the patient's progression-free survival (PFS) and objective response rate (ORR), which are assessed through regular imaging examinations (such as breast ultrasound and chest CT every 8 weeks) and laboratory tests (such as tumor marker levels). Secondary outcome measures include the patient's quality of life score (using the Breast Cancer Quality of Life Questionnaire, assessed before treatment, every 4 weeks during treatment, and after treatment) and the incidence of adverse drug reactions.
[0022] Case Report Form (CRF): Patient Personal Information: Record basic information such as the patient's name, gender, age, contact information, and ID number for patient identification and follow-up management. Medical History: Detailed record of the patient's medical history, such as whether they have chronic conditions such as hypertension, diabetes, and heart disease, as well as previous breast cancer treatments, including surgical method, chemotherapy regimen, and radiation dose and duration. This information is important for assessing the patient's overall health and tolerance to the investigational drug. Treatment Record: Record the duration of each treatment session, drug dosage, and post-drug reactions during this clinical study. For example, whether the patient experienced adverse reactions such as nausea, vomiting, and rash after the first dose of the new targeted drug, as well as the severity and duration of the adverse reaction, so that timely adjustments to the treatment plan or appropriate symptomatic treatment measures can be made. Follow-up: Record the time of each follow-up visit, follow-up examination results (such as imaging examination reports and laboratory test data), the patient's self-perception and symptom changes. For example, if the patient reports symptoms such as fatigue and joint pain during follow-up, this information will be analyzed together with the examination results to assess the patient's disease progression and treatment effectiveness. During the project, this information will be updated in real time as patients are enrolled and treatment progresses. For example, when a new patient meets the recruitment criteria, their personal information and medical history are immediately entered into the system. After each treatment and follow-up visit, the corresponding treatment records and follow-up status are also updated promptly, allowing researchers to keep abreast of the latest progress of the project.
[0023] In addition, target users' inquiries regarding clinical research include, but are not limited to, "Can I take health supplements to boost my immunity while taking this new targeted drug?" and "If I develop a mild rash during treatment, do I need to stop taking the drug? What should I do?" Researchers also ask questions such as, "During data analysis, how can we more accurately assess differences in drug response among different patient subgroups (e.g., age groups and pathology subgroups)?" and "For patients experiencing severe adverse reactions, in addition to following established procedures, are there any special measures or indicators that require further observation?"
[0024] Next, taking a large general hospital as an example, its scientific research project information is as follows: Basic medical research projects: such as "Molecular Biology Research on the Pathogenesis of Breast Cancer", which aims to explore the key molecular targets and signaling pathways in the occurrence and development of breast cancer, and provide a theoretical basis for subsequent drug development and treatment strategy formulation. Clinical research projects: In addition to the above-mentioned "Clinical Research on New Targeted Therapeutic Drugs for Breast Cancer", there is also "Clinical Application Research on Early Screening Technology for Breast Cancer", which mainly studies the accuracy and effectiveness of different screening methods (such as mammography, breast ultrasound, magnetic resonance imaging, etc.) in the early diagnosis of breast cancer, and how to optimize the screening process and improve the early diagnosis rate. Translational medicine research projects: such as "From Laboratory to Clinic: Validation and Application Research of New Biomarkers for Breast Cancer", which focuses on clinical verification of potential biomarkers discovered in basic research, and explores their application value in breast cancer diagnosis, treatment monitoring and prognosis evaluation, to promote the transformation of basic research results into clinical practice.
[0025] The progress information of each scientific research project is as follows: "Molecular Biology Research on the Pathogenesis of Breast Cancer": The cell experiment and animal model construction stages have been completed, and large-scale sample data analysis is currently underway. It is estimated that it will take another 6 months to complete the data analysis and paper writing. "Clinical Research on New Targeted Therapeutic Drugs for Breast Cancer": As mentioned above, the patient recruitment and treatment stages are in progress. 50 patients have been recruited and the first treatment cycle of 20 patients has been completed. According to the plan, the recruitment of all patients will be completed within the next 3 months, and treatment and follow-up observations will continue. "Clinical Application Research on Early Screening Technology for Breast Cancer": A prospective screening study of 1,000 women has been completed, and the collected data are being collated and analyzed. Preliminary results show that combined screening of mammography and ultrasound has high sensitivity and specificity in the early diagnosis of breast cancer, but the screening thresholds and processes still need to be further optimized. The project is expected to complete the final report and the formulation of the results transformation plan within 1 year.
[0026] The following are examples of similar targeted clinical study data collected from institutions across different regions and at varying levels of medical expertise. Data from a similar clinical study conducted at an internationally renowned cancer research center: This center conducted a similar clinical study of a targeted breast cancer therapy using a different dose-escalation strategy and follow-up timelines than the target hospital. This center employed a more aggressive dose-escalation strategy, with a higher starting dose but a slower dose ramp-up to assess patient tolerance and efficacy of the higher doses. Follow-up was conducted at weeks 4, 8, and 12 after treatment, as opposed to the target hospital's primary assessment every 8 weeks. Analysis of this center's data can inform target hospitals' studies of different trial designs and potential risks and benefits. The included patient population was more ethnically diverse, with a significant number of African patients in addition to Caucasians and Asians. The study revealed differences in pharmacokinetic and adverse reaction rates among patients of different ethnicities, raising new considerations for target hospitals when analyzing patient data. Further research is needed to examine the impact of ethnicity on drug efficacy and safety, and to adjust inclusion criteria or conduct stratified analyses in subsequent studies.
[0027] Data from similar clinical studies at primary healthcare facilities in China: A small-scale clinical study of a targeted breast cancer therapy conducted by a primary healthcare hospital, while limited in sample size, exhibited unique characteristics in patient management and data collection. The hospital utilized telemedicine technology for patient follow-up. Patients uploaded relevant symptom and sign information from their homes via a mobile app, and physicians provided online assessment and guidance. This approach significantly improved patient compliance and the convenience of follow-up. The target hospital could leverage this telemedicine follow-up model to optimize its own patient management processes, particularly for patients living in remote areas or with limited mobility, thereby increasing their willingness and feasibility to participate in clinical research. Regarding data analysis, the primary healthcare hospital employed simple, accessible statistical methods to analyze preliminary data. While the analysis was limited in depth, it was able to quickly provide some basic efficacy and safety information. The target hospital could leverage this data analysis approach by implementing simple statistical methods to conduct preliminary data screening and analysis early in the project, identifying potential issues and trends and providing guidance for subsequent in-depth analysis.
[0028] The default large language model is the Llama model, a large language model based on the Transformer architecture. The following is a description of its relevant features: The Llama model is an autoregressive language model that can predict the next word or character based on a given previous text, thereby generating a coherent text sequence. It performs well in natural language processing tasks such as text generation, question-answering systems, and machine translation. It has a multi-layer neural network structure. For example, the Llama model contains dozens or even hundreds of Transformer layers, each of which continuously extracts and transforms the semantic and grammatical information of the input text. As the number of layers increases, the model can learn more complex language patterns and semantic relationships. Its core structure is the Transformer architecture, which includes components such as the Multi-Head Attention mechanism, the Feed-Forward Neural Network, and the Layer Normalization method. The multi-head attention mechanism allows the model to focus on different parts of the input text simultaneously, capturing the semantic associations between different locations. The feedforward neural network performs further nonlinear transformations on the information processed by the attention mechanism, enhancing the model's expressive power. Layer normalization helps stabilize the model's training process, improving training efficiency and the model's generalization capabilities. In this clinical research implementation quality control system, the Llama model's powerful language processing capabilities are utilized to analyze and process collected clinical research project information, user consultation questions, and so on. For example, when generating quality control reports, by processing text block information and constructing a knowledge graph, the Llama model can combine its learned medical knowledge and language patterns to provide users with accurate and professional answers, and assist researchers in improving quality control and research processes.
[0029] The training sample set includes a large amount of textual data, including medical literature, clinical research reports, case studies, and medical experts' summaries. For example: Medical Literature: This collection covers articles published in reputable domestic and international medical journals on basic research, clinical treatment, and drug development related to breast cancer, including research papers, reviews, case reports, and other types of data. Clinical Research Reports: This collection includes detailed reports of completed and ongoing breast cancer clinical studies worldwide, including trial design, patient recruitment, treatment plans, and results analysis. These reports provide a wealth of real-world examples and data support for the model, helping it learn to predict outcomes and solve problems under diverse experimental conditions. Case Studies: This collection includes detailed case studies of breast cancer patients from major hospitals, covering their clinical presentations, diagnostic processes, treatment challenges and solutions, and prognosis. This provides the model with a deep understanding of various real-world clinical scenarios and strategies. Medical Experts' Summaries: This collection includes renowned experts in the field of breast cancer, sharing their experiences and insights from clinical practice and research, including their unique understanding of the disease, treatment techniques, and research approaches. These summaries infuse the model with professional clinical wisdom and research insights, improving its accuracy and reliability in addressing complex medical problems. By training on these rich and diverse training sample sets, the preset large language model can learn professional knowledge and language expression patterns in the medical field, so that in subsequent applications it can better process target clinical research project information, answer target users' consulting questions, and assist in generating quality control reports and analyzing and improving research quality.
[0030] S102: Process the preset large language model based on the training sample set to generate a target large language model.
[0031] In one embodiment, any number of data features in the training sample set are obtained. An indefinite number of data features are randomly selected from the training sample set. These data features cover a wealth of clinical research-related information, including descriptions of disease symptoms, details of treatment methods, drug response records, patient demographic information, clinical research process steps, and other aspects. For example, in the training samples of breast cancer clinical research, data features involve imaging features of breast cancer of different pathological types, changing trends of blood test indicators of patients at different treatment stages, and quality of life assessment indicators of patients under specific treatment plans. By widely acquiring these different types of data features, the foundation for subsequent model training is laid, ensuring that the model can learn various knowledge and laws in the field of clinical research.
[0032] Based on the number of each data feature in the training sample set, a sampling ratio is generated to ensure data diversity while reasonably selecting representative data for model training. If a particular type of data feature is present in a large number in the sample set, its proportion in the sampling process will be reduced accordingly to avoid overfitting the model to this type of data; conversely, if certain data features are relatively scarce but crucial to clinical research, such as the special symptoms of certain rare diseases or early application cases of new treatments, their weight will be appropriately increased in the sampling ratio to ensure that the model can fully learn this key information. For example, in the training samples for a study of a rare cancer, although the number of cases is limited, the unique genetic variation information and special treatment response data contained in these cases will be given priority consideration when determining the sampling ratio to ensure that the model can accurately capture these key features.
[0033] The training sample set is processed based on the sampling ratio to generate a preset number of sampling features. This process is actually a targeted screening and extraction of the original training samples, so that the sampling features ultimately used for model training can not only reflect the overall feature distribution of the training sample set, but also highlight key points and critical information. For example, in a training sample set containing a large number of clinical research data on different diseases, if the number of sampling features is set to 1000, then according to the previously determined sampling ratio, 300 sampling features will be selected from the numerous cardiovascular disease test data, 400 sampling features will be selected from the tumor disease test data, and 300 sampling features will be selected from the neurological disease test data. In addition, the sampling features in each category are carefully selected to represent the key aspects of the clinical research of that type of disease, such as comparative data on the effects of different drug treatments in cardiovascular diseases, data on the correlation between tumor marker changes and treatment plans in tumor diseases, and data on the relationship between neurological function assessment indicators and treatment progress in neurological diseases.
[0034] Based on any data feature and each sampling feature, multiple data groups are generated, wherein each data group contains a preset number of data samples, and at least one data sample includes identification information. In each data group, a preset number of data samples are included, and at least one data sample has identification information. These identification information are usually used to mark whether the data samples are related to risk factors that affect clinical research projects. For example, in a data group on a clinical study of diabetes drugs, the data samples include the patient's blood glucose monitoring records, drug dosage and time, whether adverse reactions such as hypoglycemia or hyperglycemia occur, and other information. Among them, the data samples of patients with serious adverse reactions will have specific identification information, indicating that they are risk factors that affect clinical research projects. By constructing the data group in this way, the model can learn the association between different data features and their relationship with risk factors during the training process, thereby improving the model's ability to identify and judge potential problems in clinical research.
[0035] A pre-set large language model is trained based on data samples from multiple datasets to generate a trained large language model. During training, the model continuously adjusts its parameters to learn the language patterns and semantic relationships in the input data and attempts to predict various information related to clinical research, such as disease trends, evaluation of treatment effectiveness, and emerging risks and issues. For example, the model learns how to generate reasonable diagnostic recommendations based on a patient's symptom description and medical history, or predict research results and potential risk points based on the progress of clinical research. Through training with a large number of data samples, the model gradually improves its language processing and knowledge application capabilities in the field of clinical research.
[0036] The trained large language model is processed based on the validation set to generate validation results. If the data samples containing identifying information in the validation results are confirmed to represent risk factors affecting the clinical research project, the trained large language model is selected as the target large language model. The validation set also contains rich clinical research data, but differs from the training set in data distribution and specific cases. It is used to test the model's generalization and accuracy on unseen data. During the validation process, the model analyzes and predicts data from the validation set and compares the results with known ground truth. If the data samples containing identifying information in the validation results are confirmed to represent risk factors affecting the clinical research project, it indicates that the model accurately identifies these key information. The trained large language model is then selected as the target large language model and used for subsequent clinical research quality control tasks, such as generating quality control reports and analyzing research project risks and issues. Conversely, if the model's accuracy in identifying risk factors during validation is low, further adjustments and optimization are required, such as re-adjusting training parameters, increasing the amount of training data, or improving data preprocessing methods. The model is then trained and validated again until it meets the expected performance standards.
[0037] S103: Process the target clinical research project information to generate an external knowledge base.
[0038] In one embodiment, the target clinical research project information is filtered and formatted to generate clinical research project information in a target format. The target clinical research project information comes from a wide range of sources and in various formats, including electronic documents, database records, clinical reports, etc. First, this information needs to be filtered and formatted. For example, in a project information containing a large number of patient medical records and research records, irrelevant medical equipment operation logs, non-critical administrative records and other information are filtered out, and only data related to the core content of the clinical research is retained, such as the patient's basic condition description, treatment plan, test results, etc. At the same time, data in different formats are uniformly converted into a standard format, such as unifying the date format to "YYYY-MM-DD" and the text encoding to UTF-8, etc., to ensure data consistency and readability, thereby generating clinical research project information in the target format.
[0039] The target format clinical research project information is subjected to text segmentation and vectorization to generate text block information. After formatting, the target format clinical research project information is processed using a specific text segmentation tool. The text can be segmented according to logical pauses in natural language, such as line breaks, periods, question marks, exclamation marks, etc., or it can be divided into manageable blocks based on a pre-set chunk size. For example, in a detailed clinical research report, the text is segmented into paragraphs or into blocks of 500 characters. These text blocks are then converted into corresponding vector representations (embeddings) using open-source text vectorization tools or models such as Word2Vec and BERT (Llama-3.1-Nemotron-70B-Instruct is used as an example in this article). These vectors can capture the semantic information of the text, enabling computers to better understand and process the text content, thereby generating text block information.
[0040] Text blocks are categorized and processed to generate real-time progress and research challenges of clinical research projects. Categorization can include factors such as disease type, research stage, and treatment approach. For example, in a comprehensive clinical research project, text blocks related to disease diagnosis are grouped into one category, while those related to the drug treatment process are grouped into another. This categorization approach clearly identifies different aspects of a clinical research project, generating real-time progress and research challenges. For example, in a clinical study of a cancer drug, improved efficacy in a specific patient population is identified as real-time progress, while side effect management and countermeasures are identified as research challenges. The real-time progress and research challenges of clinical research projects are processed to generate text feature correlation information. The relationship between the real-time progress and research challenges of clinical research projects is analyzed to mine the text feature correlation information. For example, improvements in a treatment approach may be closely linked to breakthroughs in specific research challenges, or certain real-time achievements may be attributed to changes in specific factors. Through statistical analysis, semantic association analysis and other methods, the intrinsic connections between this information are found, providing a basis for the subsequent construction of the knowledge base.
[0041] Generate an external knowledge base based on text feature correlation information. Based on the text feature correlation information generated above, relevant text blocks, classification information, progress of results and research difficulties are integrated together to build a structured external knowledge base. This knowledge base can be stored and managed in the form of databases, knowledge graphs, etc., so that it can be quickly retrieved and utilized in the subsequent research process. For example, in an external knowledge base constructed in the form of a knowledge graph, diseases, treatment methods, research institutions, etc. are used as nodes, and the relationships between them are used as edges. Various information in clinical research projects are organically organized to provide clinical researchers with comprehensive and accurate knowledge support, assisting them in research decision-making, problem solving and quality control. It can effectively transform the original target clinical research project information into an external knowledge base with high usability and value, providing an important knowledge foundation for the entire clinical research implementation quality control system.
[0042] S104: Process the scientific research project information of the target hospital to generate a clinical research management impact factor.
[0043] In one embodiment, the scientific research project information of the target hospital is processed to generate the classification information of the scientific research projects of the target hospital and the progress information of each scientific research project. The scientific research project information of the target hospital covers numerous projects in multiple fields such as basic research, clinical research, and translational medicine research. First, by analyzing the research topics, research methods, and application directions of these projects, they are classified into different categories. For example, in medical research, they can be divided into categories such as cardiovascular disease research, tumor disease research, and nervous system disease research. At the same time, for each scientific research project, its progress information is recorded in detail, including whether the project is in the start-up stage, data collection stage, data analysis stage, or results summary stage, and the key time nodes and completion status of each stage are clarified. For example, for a clinical research project on cardiovascular disease, its progress information shows that patient recruitment has been completed, mid-term data collection and analysis are in progress, and the preliminary results evaluation is expected to be completed within the next 6 months.
[0044] The progress information of scientific research projects of the same category and scientific research projects of corresponding categories are processed separately to generate abnormal text features of different scientific research projects. For scientific research projects of the same category, their progress information and process details are compared to find out the differences from the normal process or expected results, and the relevant text descriptions are extracted as abnormal text features. For example, in multiple clinical research projects for drugs of tumor diseases, if it is found that some projects have an adjustment frequency or adjustment range that exceeds the normal range during the drug dosage adjustment stage, then these special descriptions of dosage adjustments will be extracted as abnormal text features. Similarly, for scientific research projects of different categories but with certain correlations, such as drug development projects for cardiovascular diseases and interventional treatment research projects for cardiovascular diseases, similar comparative analysis is also carried out to dig out abnormal text features in terms of research paths, changes in key indicators, etc.
[0045] In one embodiment, an example of an abnormal text feature is as follows: Project progress: In a "clinical study of a new targeted therapy for breast cancer," it was planned to conduct a comprehensive imaging review of patients in the 12th week after the start of treatment to assess tumor shrinkage. However, the review of 70% of patients was actually completed in the 14th week. The relevant text record was "The imaging review progress of the clinical study of targeted therapy for breast cancer has been delayed, and only 70% of patients have been reviewed, which is 2 weeks behind schedule." Research methods: In a "research on technology for early screening of breast cancer" project, it was originally stipulated that a standardized breast ultrasound examination process and image analysis method should be used. However, in actual operation, some operators did not set the frequency of the ultrasound probe in accordance with the standard, and did not strictly judge the image according to the established BI-RADS classification standards. The text description was "The breast cancer screening ultrasound examination process is not standardized, the probe frequency is set incorrectly, and the image analysis does not follow the BI-RADS standards." Data quality: In a study titled "Relationship between Breast Cancer Gene Expression and Prognosis," it was discovered that some patients' gene chip test data contained duplicate records and data entry errors. For example, some gene expression values were identical at different time points and clearly did not conform to biological logic. The recorded text read, "Breast cancer gene expression data contains duplicate records and entry errors, affecting data accuracy and the reliability of research results."
[0046] In another embodiment, examples of unusual text features are as follows: Regarding project progress: In a diabetes drug clinical study, an interim data analysis was originally planned for the sixth month after 200 patients had been enrolled. However, the actual progress showed that only 150 patients had been enrolled by the eighth month, and the interim analysis had not yet begun. A related text description, such as "Patient recruitment progress is lagging, affecting the timing of the interim analysis," is an unusual text feature related to project progress. Regarding research methods: In an interventional treatment study for cardiovascular disease, a specific angiographic technique was required to assess treatment efficacy according to the established protocol. However, an alternative technique was used in some cases, and no reasonable explanation was provided in the study records. The text description, "The use of angiographic techniques during the study was inconsistent with the protocol, with no reasonable explanation," is an unusual text feature related to research methods. Regarding data quality: In a tumor genetic testing research project, the gene sequencing data of some samples contained a large number of missing values, and these missing values were not properly handled in the data statistics and analysis. The corresponding text description, "The gene sequencing data contains a large number of missing values and has not been handled, affecting data reliability," is an unusual text feature related to data quality.
[0047] The anomalous text features of different scientific research projects are processed to generate information on the anomalous common features of the scientific research projects and the frequency of their occurrence. After collecting and organizing the anomalous text features of each scientific research project, data analysis and text mining techniques are used to identify common features. For example, irregularities in the sample selection process or similar errors in the application of data analysis methods found in multiple different scientific research projects constitute anomalous common features. At the same time, the number of times each anomalous common feature appears in different projects is counted to assess its prevalence and severity. For example, the anomalous common feature of irregular sample selection appeared six times in ten related scientific research projects, indicating that it is an issue that requires special attention.
[0048] In one embodiment, examples of abnormal common features are as follows: Sample management: In multiple breast cancer-related research projects, including "Pathological Studies of Different Breast Cancer Subtypes" and "Studies Predicting the Risk of Postoperative Recurrence of Breast Cancer," insufficient sample collection was observed. For example, while the research plan stipulated that each research group should collect at least 100 samples, some research groups in multiple projects only collected 70-80 samples. This was described in the text as "Breast cancer-related research sample collection fell short of the planned number, impacting the statistical power and reliability of the study conclusions." Statistical analysis: In some breast cancer treatment efficacy evaluation projects, such as the "Comparative Study of the Efficacy of Chemotherapy Combined with Radiotherapy and Chemotherapy Alone for Breast Cancer" and the "Analysis of the Effects of Different Endocrine Therapy Regimens for Breast Cancer," important confounding factors such as patient age and underlying medical conditions were mistakenly ignored when analyzing patient survival data. This was documented as "Survival analysis in breast cancer treatment efficacy evaluations failed to consider confounding factors such as age and underlying medical conditions, leading to biased results." Ethical review: In clinical research projects involving breast cancer patients, such as the "Clinical Study of Neoadjuvant Therapy for Breast Cancer" and the "Breast Cancer Rehabilitation Therapy Intervention Trial," unclear measures were taken to protect patient privacy during the ethical review process. For example, during the patient information registration and data storage process, there is no detailed description of how to encrypt and restrict access to patients' sensitive information. The text describes it as "a lack of patient privacy protection measures in the ethical review of breast cancer clinical research, and information security risks."
[0049] In another embodiment, examples of abnormal common characteristics are as follows: Sample management issues: In clinical research projects for multiple different diseases, sample storage conditions have failed to meet standard requirements. For example, in some infectious disease and genetic disease research projects, unstable temperature control during sample transportation and storage has been found, resulting in sample damage. A text description such as "Poor sample storage temperature control affects sample quality" is an example of an abnormal common characteristic in sample management. Statistical analysis errors: In several drug development projects, inappropriate statistical test methods were incorrectly used when conducting comparative drug efficacy analyses. For example, in some antibiotic and anticancer drug development projects, a simple t-test was used instead of a multi-group analysis of variance, resulting in biased results. The text description, "Incorrect selection of statistical test method in comparative drug efficacy analysis," constitutes an abnormal common characteristic in statistical analysis. Ethical review loopholes: In some scientific research projects involving human trials, whether clinical research on neurological diseases or orthopedic diseases, there are issues with incomplete ethical review documents or non-standard approval processes. For example, the informed consent forms of some projects lack key information or the signing date is unclear, and the text describes it as "there are defects in the ethical review documents, the informed consent form information is incomplete or the signing is not standardized", which has become an abnormal common feature in ethical review.
[0050] The abnormal common characteristics and the number of occurrences of abnormal common characteristics of scientific research projects are processed to generate clinical feasibility information corresponding to different scientific research project categories in the target hospital. Based on the abnormal common characteristics of scientific research projects and the number of occurrences, combined with medical professional knowledge and clinical practice experience, the feasibility of each scientific research project category in actual clinical application is evaluated. For project categories with more serious abnormal common characteristics, their clinical feasibility is lower; while project categories with fewer and relatively mild abnormalities have higher clinical feasibility. For example, in some basic medical research projects, if it is found that there are many defects in the experimental design and these problems are common in multiple projects, then the feasibility of such projects being converted into clinical applications will be questioned; on the contrary, some clinical research projects that have been strictly designed and implemented in accordance with regulations have relatively high clinical feasibility if there are only a few solvable abnormalities.
[0051] The clinical feasibility information corresponding to different scientific research project categories of the target hospital is processed to generate a clinical research management impact factor. The impact factor is a quantitative indicator that can reflect the potential impact of different scientific research project categories on clinical outcomes. For example, project categories with high clinical feasibility will be assigned a higher impact factor weight, indicating that they have a greater driving effect on improving clinical treatment effects and improving diagnostic methods; while project categories with low clinical feasibility will receive a lower impact factor, suggesting that these projects need further improvement and optimization to increase their contribution to clinical outcomes. Through the clinical research management impact factor, hospital managers and researchers can more intuitively understand the value and potential risks of each scientific research project, thereby rationally allocating resources, giving priority to supporting projects that have a positive impact on clinical outcomes, promoting the close integration of hospital scientific research and clinical practice, and improving the overall medical level and research quality.
[0052] S105: Process the other clinical research project information to generate user individual information impact factors.
[0053] In one embodiment, feature extraction processing is performed on other clinical research project information to generate target features that match the user's physiological information, wherein the target features that match the user's physiological information are used to characterize the changing trends of physiological indicators at different treatment stages, the difference information of physiological characteristics of different user groups, and the correlation information between physiological indicators. From other clinical research project information, a large number of physiological indicator data of breast cancer patients at different treatment stages (such as before surgery, after surgery, during chemotherapy, during radiotherapy, during endocrine therapy, etc.) are collected, including but not limited to the levels of tumor markers (such as CA15-3, CEA, etc.) in the blood, hormone levels (such as estrogen, progesterone, androgen, etc.), blood routine indicators (such as white blood cell count, red blood cell count, platelet count, etc.), liver and kidney function indicators (such as alanine aminotransferase, aspartate aminotransferase, creatinine, urea nitrogen, etc.) and physical symptoms (such as pain level, fatigue, sleep quality, etc.). Through data analysis and statistical methods, the changing trends of these physiological indicators with the treatment stage are extracted. For example, in the early stage of chemotherapy, the CA15-3 indicator will rise briefly and then gradually decline. This is a target feature of a changing trend. At the same time, the physiological characteristics of user groups with different ages, genders, races, and basic physical conditions (such as whether they suffer from other chronic diseases) are compared. For example, the hormone level fluctuations of young female breast cancer patients during treatment are significantly different from those of older female patients. This reflects the difference in the physiological characteristics of the user groups. In addition, the correlation between different physiological indicators is analyzed. For example, it is found that the decrease in white blood cell count of some breast cancer patients has a certain positive correlation with the dosage of chemotherapy drugs. This is the correlation information between physiological indicators. These change trends, difference information, and correlation information together constitute the target features that match the user's physiological information.
[0054] Target features matching user physiological information are processed to generate user groups with similar physiological response patterns and common influencing features within each user group. Using algorithms such as cluster analysis, based on the extracted target features, breast cancer patients with similar physiological indicator change trends, similar physiological characteristic differences, and correlations with related physiological indicators are grouped together. For example, patients experiencing similar levels of gastrointestinal reactions such as nausea and vomiting during chemotherapy and experiencing similar rates of decline in blood tumor markers are grouped together. For each user group, common influencing features are further analyzed. These include specific treatment regimen combinations (such as the combination of a chemotherapy drug and an endocrine therapy drug), lifestyle factors (such as regular exercise habits and dietary preferences), and genetic factors (such as the presence of specific breast cancer-related gene mutations). For example, if the majority of patients in a user group received the AC-T chemotherapy regimen and had similar BRCA1 gene mutation status, these would be the common influencing features of that group. The common influencing features of each user group are processed to generate feature analysis results for the target user group. For each group, an in-depth analysis is conducted to determine the comprehensive impact of these common influencing features on treatment outcomes and patient physiological status. Through statistical analysis and clinical research experience, we evaluate the relationship between these characteristics and patients' disease progression, survival, quality of life, and other aspects. For example, for the user group mentioned above with a specific chemotherapy regimen and gene mutation, the analysis found that this group of patients had a relatively low five-year survival rate after treatment and a higher probability of experiencing serious adverse reactions during treatment. This is the characteristic analysis result of this target user group.
[0055] Based on the characteristic analysis results of the target user groups, user individual information impact factors are generated. According to the characteristic analysis results of each target user group, a corresponding weight or score is assigned to each group to quantify its impact on the overall user physiological information, thereby generating user individual information impact factors. For example, for those user groups that have a greater impact on the treatment effect and patient physiological status, such as those with high-risk gene mutations and poor treatment response, their impact factors will be higher; while for groups with relatively small impact, such as those with milder symptoms and better treatment effects, the impact factors will be lower. These impact factors can be used to evaluate the characteristics and risks of different patient groups in subsequent clinical research and medical decision-making, providing an important reference basis for the formulation of personalized treatment plans and the adjustment of research directions, so as to improve the quality and efficiency of breast cancer clinical research and better serve patients' treatment and rehabilitation.
[0056] The user individual information impact factor comprehensively characterizes the degree and pattern of how physiological information of different user groups is affected by multiple factors in a specific clinical research scenario. It reflects the potential driving effect of different treatment stages, user group characteristics, and physiological indicator associations mined based on extensive clinical research data on changes in individual physiological states. Specifically, a high impact factor indicates that certain common characteristics of a specific user group or specific change trends in the treatment stage have a strong impact on physiological information. These factors need to be focused on in clinical practice. For example, when formulating treatment plans or predicting disease progression, the physiological characteristics and change patterns represented by these high impact factors should be fully considered; a low impact factor indicates that the corresponding factors have a relatively weak impact on physiological information, and their priority can be appropriately lowered when allocating resources and determining research priorities, thereby providing a quantitative reference basis for clinical decision-making and research direction adjustments, and assisting in the realization of personalized medicine and precision medicine research.
[0057] S106 , processing the target user's consulting questions regarding clinical research based on the target large language model and the external knowledge base, and generating a quality control report.
[0058] In one embodiment, the text block information is processed based on the target large language model to generate a text block vector. Figure 3 As shown, the target large language model (such as the Llama model) first receives processed text blocks from breast cancer clinical research projects. These blocks cover a wide range of topics, including breast cancer disease characteristics, treatment details, patient case data, and research progress. For example, one block describes, "In a study of endocrine therapy for breast cancer, patients in a specific age group were treated with tamoxifen for five years, during which estrogen levels and breast ultrasound results were regularly monitored. Some patients experienced side effects such as endometrial thickening." The model analyzes and encodes these blocks, converting them into vector representations with semantic and grammatical information, known as block vectors. In this process, the model leverages the language patterns and knowledge learned from extensive training on medical text to identify key concepts and semantic relationships within the text and map them into a high-dimensional vector space. This ensures that blocks with similar semantics are close together in the vector space, providing a foundation for subsequent processing.
[0059] The text block vectors and text feature correlation information are processed to generate a similarity index. This text feature correlation information contains the correlations between different elements in breast cancer clinical research, such as the association between different treatments and disease prognosis, and the association between specific gene expression and drug response. Using specific algorithms and model mechanisms, the similarity between the text block vectors and this correlation information is calculated to generate a similarity index. For example, if a text block vector involves the relationship between targeted breast cancer therapy and a specific gene mutation, the similarity index will focus on other research text blocks related to this gene mutation, as well as text blocks evaluating the effectiveness of targeted therapy. The degree of correlation between these blocks will be quantified and indexed to facilitate rapid retrieval and utilization of this relevant information. The similarity index is then processed to generate a medical knowledge graph. In the breast cancer knowledge graph, different aspects of the disease (such as pathological type, stage, and metastasis), treatment options (surgery, chemotherapy, radiotherapy, endocrine therapy, targeted therapy), and patient characteristics (age, gender, and genetic background) serve as nodes, and the relationships between them (such as whether a certain treatment is suitable for breast cancer at a specific stage, or whether certain genetic characteristics are associated with the efficacy of a specific treatment) serve as edges or labels. For example, there is a strong correlation between HER2-positive breast cancer and trastuzumab targeted therapy, while there are correlations between older breast cancer patients and certain conservative treatments. Integrating scattered breast cancer-related knowledge into a structured network comprehensively presents the knowledge system in the field of breast cancer clinical research, providing more comprehensive and in-depth background knowledge support for answering user inquiries.
[0060] Based on the medical knowledge graph, the system processes user inquiries regarding clinical research and generates quality control reports. When a user asks a question about breast cancer clinical research, such as "How can the risk of cardiotoxicity be assessed in breast cancer patients receiving combined chemotherapy and radiotherapy?" the system processes the question based on the constructed medical knowledge graph. First, the system locates nodes and edges related to chemotherapy, radiotherapy, breast cancer, and cardiotoxicity in the knowledge graph, extracting relevant knowledge and information. Then, leveraging the generative capabilities of the target large language model, the system combines this knowledge with the specific requirements of the question to generate a quality control report. The report details commonly used cardiotoxicity assessment indicators (such as myocardial enzyme spectrum testing and specific indicators from cardiac ultrasound examinations), risk factors identified in various studies (such as the patient's underlying heart disease and cumulative chemotherapy drug dose), and corresponding prevention and monitoring recommendations (such as the interval for regular cardiac function tests and measures to take when specific symptoms occur). This report provides users with accurate, professional, and targeted answers, helping clinical researchers and patients better understand and address relevant issues in breast cancer clinical research, ensuring quality control and scientific management of the research process.
[0061] S107 , processing the target clinical research project information and the quality control report based on the clinical research management influencing factor and the user individual information influencing factor to generate a diagnosis accuracy correlation.
[0062] In one embodiment, the target clinical research project information is feature-screened and matched based on the clinical research management impact factor. This generates research projects within the target hospital that are similar in type or related fields to the target clinical research project, along with progress information for each project. For the target clinical research project, "Clinical Study of a New Targeted Therapy for Breast Cancer," the clinical research management impact factor is used to screen the numerous research projects within the target hospital. For similar projects, such as studies of other targeted breast cancer drugs, similarities in study design, drug mechanism of action, and patient inclusion criteria are examined. For projects in related fields, such as studies of the pathophysiology of breast cancer and imaging assessment of breast cancer treatment, their relevance to current targeted therapy trials in terms of knowledge and research processes is analyzed. For example, a hospital has a "Clinical Study of Another New Targeted Drug Combined with Radiotherapy for Breast Cancer," which is comparable to the target project in terms of drug combination and treatment duration. Another project, "Basic Research on the Impact of the Tumor Microenvironment on Targeted Therapy in Breast Cancer," is fundamental research but holds important implications for understanding the effectiveness of targeted therapies. Through these screening and matching, we obtained the progress information of these related projects, such as the "Clinical study of another new targeted drug combined with radiotherapy for breast cancer" has completed patient recruitment and is undergoing mid-term efficacy evaluation; the "Basic research on the impact of breast cancer tumor microenvironment on targeted therapy" is in the data collection and preliminary analysis stage.
[0063] Research projects and progress information for each research project within the target hospital that are of the same type or related to the target clinical research project are processed to generate target abnormal text features. Detailed analysis is performed on the selected projects and progress information. For example, in the "Clinical Study of Another New Targeted Drug Combined with Radiotherapy for Breast Cancer," some patients developed severe radiation dermatitis during radiotherapy, exceeding the expected adverse reaction rate. This was recorded as, "In a trial of a new targeted drug combined with radiotherapy for breast cancer, some patients experienced severe radiation dermatitis, with a higher-than-expected incidence, which may affect patient compliance and treatment efficacy evaluation." In the "Basic Research on the Impact of the Tumor Microenvironment on Targeted Therapy for Breast Cancer," inconsistent processing times for tumor tissue samples were identified during sample collection. This was described as, "The significant time span between sample collection and tissue processing in breast cancer tumor microenvironment research may lead to data bias and affect the reliability of research conclusions." These are the generated target abnormal text features. Quality control reports are processed based on these target abnormal text features to generate warnings for research project abnormalities. When a quality control report involves safety and efficacy evaluation of targeted breast cancer therapies, the system will issue a warning based on these abnormal text features. For example, it suggests that "when evaluating the effectiveness of current targeted therapies for breast cancer, it is necessary to consider the potential interference of abnormal radiation dermatitis in combined radiotherapy trials on patients' overall treatment experience and outcomes, as well as the impact of data unreliability caused by sample processing problems in tumor microenvironment studies on basic theoretical support." This is an abnormal warning information for scientific research projects, reminding researchers to be cautious about these potential problems when analyzing and applying relevant research results.
[0064] Based on the user's individual information influencing factors, the target clinical research project information is feature-screened and matched, generating indicators of changes in the target user's various physiological indicators at different treatment stages and indicators of changes in the physiological indicators of other users in the same clinical research project as the target user at different treatment stages. Assume that the target user is a 50-year-old, estrogen receptor-positive patient undergoing endocrine therapy for breast cancer. From the target clinical research project information, based on the user's individual information influencing factors, the changes in their physiological indicators at different treatment stages (such as the initial, mid-term, and long-term maintenance stages of endocrine therapy) are extracted, including changes in estrogen levels, tumor markers (such as CA15-3, CEA), and bone density. At the same time, data on changes in various physiological indicators at corresponding treatment stages are collected for other similar users in the same clinical research project as the target user (such as patients aged 45-55, estrogen receptor-positive, and using the same endocrine therapy drugs). For example, it was found that the target users' estrogen levels dropped rapidly in the early stages of endocrine treatment, but the CA15-3 index rose slightly in the short term and then stabilized; some patients among other similar users also had similar estrogen level change trends, but there were differences in CA15-3 fluctuations.
[0065] Based on the change indicators of various physiological indicators of other users in the same clinical research project as the target user at different treatment stages, the change indicators of various physiological indicators of the target user at different treatment stages are processed to generate target physiological influencing factors. By performing statistical analysis on the physiological indicator change data of the target user and other similar users, such as using regression analysis and other methods, the relationship between factors such as age, treatment time, and basic health status and changes in physiological indicators is explored, and the target physiological influencing factors are quantified. For example, it was found that age is positively correlated with the rate of decline of estrogen levels in the early stage of endocrine treatment, and this correlation is more significant at specific drug doses. This relationship is converted into a target physiological influencing factor, which reflects the degree and direction of the influence of this factor on the changes in the patient's physiological state in numerical form.
[0066] The quality control report is processed based on the target physiological influencing factors to generate abnormal warning information of the target user's physiological information. During the quality control report generation process, if it involves the patient's treatment effect evaluation or treatment plan adjustment suggestions, the system will give a warning based on the target physiological influencing factors. For example, it prompts "The patient's estrogen level dropped too quickly in the early stage of endocrine treatment. Considering his age, he needs to closely monitor changes in bone density to prevent an increased risk of osteoporosis. At the same time, he needs to continue to observe fluctuations in the CA15-3 indicator. If an abnormal upward trend occurs, further examination should be carried out in time to adjust the treatment plan to ensure efficacy." This is the abnormal warning information of the target user's physiological information, which provides clinicians with personalized focus and decision-making reference for the patient.
[0067] Abnormal warning information from scientific research projects and abnormal warning information from target users' physiological information are processed to generate diagnostic accuracy correlations. For example, if similar trends in abnormal changes in physiological indicators caused by adverse drug reactions are found in multiple research projects related to endocrine therapy for breast cancer, and these factors are closely linked to research quality and processes, such as sample processing issues affecting the accuracy of research results and abnormal changes in patient physiological indicators affecting the reliability of treatment effect assessments, then the diagnostic accuracy correlation will be high. This indicates that a comprehensive review and optimization and improvement of the current research projects' drug dose adjustment strategies, sample collection and processing standards, and patient monitoring indicators and frequency are needed. For example, measures such as optimizing the dose titration scheme of endocrine therapy drugs, standardizing the processing process after sample collection, and increasing the frequency and depth of monitoring of specific physiological indicators can improve the quality and efficiency of breast cancer clinical research, ensure the scientific nature and clinical application value of research results, and achieve effective improvement and enhancement of research processes.
[0068] This application obtains various information from the server, such as target clinical research projects, user consultation questions, target hospitals and other institutions' scientific research project information. The target clinical research project information is processed to generate an external knowledge base, which goes through filtering and formatting, text segmentation and vectorization, classification, and generation of text feature correlation information. At the same time, the target hospital scientific research project information is processed to generate clinical research management influencing factors, and its classification, progress, abnormal text features, etc. are analyzed; other clinical research project information is processed to generate user individual information influencing factors, including feature extraction and cluster analysis steps. Based on the target large language model and the external knowledge base, user consultation questions are processed to generate a quality control report, which involves the construction of text block vectors, similarity indexes and medical knowledge graphs. Finally, the two influencing factors are combined to process the project information and quality control report to generate diagnostic accuracy correlation, aiming to improve research quality, improve research processes, and provide strong support for the effective conduct and quality assurance of clinical research.
[0069] In one embodiment, Figure 2 As shown, the present application also provides a clinical research implementation quality control device based on a large language model, comprising:
[0070] Acquisition module 201 is used to acquire target clinical research project information, target user consultation questions regarding clinical research, target hospital scientific research project management information, other clinical research project information, a preset large language model, and a training sample set. The target clinical research project information is updated in real time. The target hospital scientific research project information is used to represent the target hospital's scientific research project classification information, management requirements information, and progress information of each scientific research project. The other clinical research project information includes similar target clinical research data from different regions and institutions with different medical levels.
[0071] The processing module 202 is used to process the preset large language model based on the training sample set to generate a target large language model; process the target clinical research project information to generate an external knowledge base; process the scientific research project information of the target hospital to generate a clinical research management influencing factor; process the other clinical research project information to generate a user individual information influencing factor; process the target user's consulting questions regarding clinical research based on the target large language model and the external knowledge base to generate a quality control report; process the target clinical research project information and the quality control report based on the clinical research management influencing factor and the user individual information influencing factor to generate a diagnostic accuracy correlation, wherein the diagnostic accuracy correlation is used to characterize the improvement of research quality and the improvement of research process.
[0072] The computer-readable storage medium provided by the above-mentioned embodiment of the present application and the clinical research implementation quality control method based on the large language model provided by the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run or implemented by the application program stored therein. The various embodiments in this application are described in a related manner, and the same and similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. For related parts, please refer to the partial description of the embodiment of the clinical research implementation quality control method based on the large language model.
Claims
1. A method for quality control of clinical research implementation based on a large language model, characterized in that: include: Obtain target clinical research project information, target users' consulting questions regarding clinical research, target hospital's scientific research project management information, other clinical research project information, a preset large language model, and a training sample set, wherein the target clinical research project information is updated in real time, the target hospital's scientific research project information is used to represent the target hospital's scientific research project classification information, management requirements information, and progress information of each scientific research project, and the other clinical research project information includes similar target clinical research data from institutions in different regions and at different medical levels; Processing the preset large language model based on the training sample set to generate a target large language model; Processing the target clinical research project information to generate an external knowledge base; Processing the scientific research project information of the target hospital to generate clinical research management impact factors; Processing the other clinical research project information to generate user individual information impact factors; Processing the target user's consulting questions regarding clinical research based on the target large language model and the external knowledge base to generate a quality control report; The target clinical research project information and the quality control report are processed based on the clinical research management influencing factor and the user individual information influencing factor to generate diagnostic accuracy correlation, including: feature screening and matching processing of the target clinical research project information based on the clinical research management influencing factor to generate scientific research projects of the same type as the target clinical research project information and progress information of each scientific research project in the target hospital; processing the scientific research projects of the same type as the target clinical research project information and progress information of each scientific research project in the target hospital to generate target abnormal text features; processing the quality control report based on the target abnormal text features to generate scientific research project abnormality warning information; feature screening of the target clinical research project information based on the user individual information influencing factor and matching processing to generate change indicators of various physiological indicators of the target user at different treatment stages and change indicators of various physiological indicators of other users of the same clinical research project as the target user at different treatment stages; based on the change indicators of various physiological indicators of other users of the same clinical research project as the target user at different treatment stages, the change indicators of various physiological indicators of the target user at different treatment stages are processed to generate target physiological influencing factors; based on the target physiological influencing factors, the quality control report is processed to generate abnormal warning information of physiological information of the target user; the abnormal warning information of the scientific research project and the abnormal warning information of the physiological information of the target user are processed to generate diagnostic accuracy correlation, wherein the diagnostic accuracy correlation is used to characterize the improvement of research quality and the improvement of research process.
2. The method according to claim 1, wherein The preset large language model is processed based on the training sample set to generate a target large language model, including: Get any number of data features in the training sample set; Generate a sampling ratio based on the number of each data feature in the training sample set; Processing the training sample set based on the sampling ratio to generate a preset number of sampling features; Processing any data feature and each sampling feature to generate multiple data groups, wherein each data group includes a preset number of data samples, and at least one data sample includes identification information; Training the preset large language model based on data samples in the multiple data groups to generate a trained large language model; Processing the trained large language model based on the verification sample set to generate a verification result; If the data sample containing identification information in the verification result represents a risk factor affecting the clinical research project, the trained large language model is used as the target large language model.
3. The method according to claim 1, wherein Process the target clinical research project information to generate an external knowledge base, including: Filtering and formatting the target clinical research project information to generate clinical research project information in a target format; Performing text segmentation and vectorization processing on the clinical research project information in the target format to generate text block information; Classify and process the text block information to generate real-time progress of clinical research projects and information on research difficulties of clinical research projects; Processing the real-time progress of the clinical research project and information on the research difficulties of the clinical research project to generate text feature correlation information; An external knowledge base is generated based on the text feature correlation information.
4. The method according to claim 1, wherein Process the scientific research project information of the target hospital to generate clinical research management impact factors, including: Processing the scientific research project information of the target hospital to generate classification information of the scientific research projects of the target hospital and progress information of each scientific research project; The progress information of scientific research projects of the same category and corresponding categories is processed separately to generate abnormal text features of different scientific research projects; Process the abnormal text features of different scientific research projects to generate abnormal common features of scientific research projects and the number of occurrences of abnormal common features; Processing the abnormal common features of the scientific research projects and the occurrence frequency information of the abnormal common features to generate clinical feasibility information corresponding to different scientific research project categories in the target hospital; The clinical feasibility information corresponding to different scientific research project categories of the target hospitals is processed to generate clinical research management impact factors.
5. The method according to claim 1, wherein Processing the other clinical research project information to generate user individual information impact factors includes: Performing feature extraction processing on the other clinical research project information to generate target features that match the user's physiological information, wherein the target features that match the user's physiological information are used to characterize the changing trends of physiological indicators at different treatment stages, the difference information of physiological characteristics of different user groups, and the correlation information between physiological indicators; Processing target features that match user physiological information to generate user groups with similar physiological response patterns and common impact features for each user group; Process the common influencing features of each user group to generate feature analysis results for the target user group; A user individual information influencing factor is generated based on the characteristic analysis result of the target user group.
6. The method according to claim 3, wherein Based on the target large language model and the external knowledge base, the target user's consulting questions regarding clinical research are processed to generate a quality control report, including: Processing the text block information based on the target large language model to generate a text block vector; Processing the text block vector and the text feature correlation information to generate a similarity index; Processing the similarity index to generate a medical knowledge graph; Based on the medical knowledge graph, the target user's consulting questions regarding clinical research are processed to generate a quality control report.
7. A clinical research quality control device based on a large language model, characterized in that: For implementing the method of claim 1, the apparatus comprises: An acquisition module is used to obtain target clinical research project information, target users' consulting questions regarding clinical research, target hospital's scientific research project management information, other clinical research project information, a preset large language model, and a training sample set. The target clinical research project information is updated in real time. The target hospital's scientific research project information is used to represent the target hospital's scientific research project classification information, management requirements information, and progress information of each scientific research project. The other clinical research project information includes similar target clinical research data from institutions in different regions and at different medical levels. A processing module is used to process a preset large language model based on a training sample set to generate a target large language model; process the target clinical research project information to generate an external knowledge base; process the scientific research project information of the target hospital to generate a clinical research management impact factor; process the other clinical research project information to generate a user individual information impact factor; process the target user's consulting questions regarding clinical research based on the target large language model and the external knowledge base to generate a quality control report; process the target clinical research project information and the quality control report based on the clinical research management impact factor and the user individual information impact factor to generate a diagnostic accuracy correlation, wherein the diagnostic accuracy correlation is used to characterize the improvement of research quality and the improvement of research process.
8. An electronic device, characterized in that: include: a first processor; and a memory for storing executable instructions of the first processor; The first processor is configured to execute the clinical research implementation quality control method based on a large language model according to any one of claims 1 to 6 by executing the executable instructions.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the second processor, the method for quality control of clinical research implementation based on a large language model according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Automatic text dialogue method and device for clinical test, equipment and medium
CN117390145A
Intelligent subject evaluation method and device based on multi-dimensional data
CN119274797A