Clinical research implementation quality control method and system based on large language model

Through a method based on a large language model, an external knowledge base and influencing factor is generated, and a quality control report and accurate diagnostic correlation is generated by combining a large language model and knowledge graph, the problems of instability in data quality and difficulty in data integration in clinical research are solved, and efficient quality control report generation and research process optimization are achieved.

CN120048408AActive Publication Date: 2025-05-27PEKING UNIVERSITY THIRD HOSPITAL (THE THIRD CLINICAL MEDICAL SCHOOL OF PEKING UNIVERSITY)

Patent Information

Application Number
CN202510118642.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

There are problems in clinical research that data quality is unstable, data integration is difficult in different institutions, and lack of efficient quality control report generation tools, which affect the reliability and efficiency of the research results.

Method used

Using a method based on a large language model, we generate external knowledge bases, clinical research management impact factors and user individual information impact factors by obtaining and processing clinical research project information, user consultation questions and scientific research project information of different institutions, and combine large language models and knowledge graphs to generate quality control reports and diagnostic accuracy correlation.

Benefits of technology

It has improved the data quality and integration efficiency of clinical research, generated accurate quality control reports and diagnostic suggestions, and improved research quality and process optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048408A_ABST
    Figure CN120048408A_ABST
Patent Text Reader

Abstract

The invention provides a clinical research implementation quality control method and system based on a large language model, and is applied to the technical field of data processing. The method comprises the following steps: processing target clinical research project information to generate an external knowledge base; processing the scientific research project management information of the target hospital to generate clinical research management influence factors; other clinical research project information is processed, and user individual information influence factors are generated; based on the target large language model and the external knowledge base, processing the consultation problem of the target user for the clinical research, and generating a quality control report; the target clinical research project information and the quality control report are processed on the basis of the clinical research management influence factors and the user individual information influence factors, diagnosis accuracy relevance is generated, and the diagnosis accuracy relevance is used for representing improvement of research quality and improvement of a clinical research process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly relates to a method and system for quality control of clinical research implementation based on a large language model. Background Art

[0002] In the field of clinical research, there are many deficiencies in traditional research implementation quality control. With the development of medical technology and the increasing complexity of clinical research, the research process involves a large amount of patient data, diverse research projects, and information comparison and integration among different regions and medical institutions with different levels.

[0003] On the one hand, clinical research needs to process a vast amount of patient data, including personal information, medical history, treatment records, and follow-up situations in case report forms. The collection, collation, and analysis of these data are heavy and error-prone. At the same time, research participants often have limited research time and are difficult to comprehensively and effectively manage the patients in the study, resulting in uneven data quality and affecting the reliability and effectiveness of research results. On the other hand, there is a lack of an effective integration and utilization mechanism for data from similar clinical studies carried out by different regions and medical institutions with different levels. These data contain rich information, such as the differences in the effects of different treatment regimens in different populations and the coping experiences of special cases. However, due to the dispersion of information, it is difficult to extract valuable references from them, and it is impossible to provide comprehensive background support and risk warning for the current research.

[0004] In addition, in the face of problems that occur during the research process, there is a lack of an efficient tool to quickly generate accurate quality control reports and effective diagnostic suggestions. Traditional methods may rely on manual experience and cumbersome statistical analysis, with low efficiency and difficulty in comprehensively covering various complex situations, and cannot meet the dual requirements of quality and efficiency in clinical research.

[0005] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0006] The purpose of this application is to provide a quality control method and system for the implementation of clinical research based on large language models, which can at least overcome the problems existing in the prior art to a certain extent. By obtaining target clinical research project information, user consultation questions, and scientific research project information of different institutions, etc. Subsequently, the target clinical project information is sequentially filtered and formatted, segmented and vectorized, classified, and associated to form an external knowledge base; for the scientific research information of the target hospital and other institutions, clinical research management impact factors and user individual information impact factors are respectively generated. The former analyzes the classification progress, etc., and the latter is obtained through feature extraction and clustering. Then, using the target large language model and the external knowledge base, a quality control report is generated by constructing text block vectors, similarity indexes, and medical knowledge graphs. Finally, by combining the two impact factors to process relevant information, a diagnostic accurate correlation is generated, which effectively helps to improve the quality of clinical research and optimize the process.

[0007] Other features and advantages of this application will become apparent from the following detailed description, or be learned in part through the practice of the present invention.

[0008] According to one aspect of this application, a quality control method for the implementation of clinical research based on large language models is provided, including: obtaining target clinical research project information, consultation questions of target users regarding clinical research, scientific research project management information of target hospitals, other clinical research project information, a preset large language model, and a training sample set. Among them, the target clinical research project information is updated in real time, and the scientific research project information of the target hospital is used to represent the classification information, management requirement information, and progress information of each scientific research project in the target hospital. The other clinical research project information includes similar target clinical research data of institutions with different regions and different medical levels; processing the preset large language model based on the training sample set to generate a target large language model; processing the target clinical research project information to generate an external knowledge base; processing the scientific research project information of the target hospital to generate clinical research management impact factors; processing the other clinical research project information to generate user individual information impact factors; processing the consultation questions of target users regarding clinical research based on the target large language model and the external knowledge base to generate a quality control report; processing the target clinical research project information and the quality control report based on the clinical research management impact factors and the user individual information impact factors to generate a diagnostic accurate correlation, where the diagnostic accurate correlation is used to represent improving research quality and optimizing research processes.

[0009] Another aspect of the present application is a quality control device for clinical research implementation based on a large language model, which is characterized by including: an acquisition module for acquiring target clinical research project information, consultation questions of target users regarding clinical research, scientific research project management information of the target hospital, other clinical research project information, a preset large language model, and a training sample set. Among them, the target clinical research project information is updated in real time, and the scientific research project information of the target hospital is used to represent the classification information, management requirement information, and progress information of each scientific research project in the target hospital. The other clinical research project information includes similar target clinical research data of institutions with different regions and different medical levels; a processing module for processing the preset large language model based on the training sample set to generate a target large language model; processing the target clinical research project information to generate an external knowledge base; processing the scientific research project information of the target hospital to generate a clinical research management influence factor; processing the other clinical research project information to generate a user individual information influence factor; processing the consultation questions of the target user regarding clinical research based on the target large language model and the external knowledge base to generate a quality control report; processing the target clinical research project information and the quality control report based on the clinical research management influence factor and the user individual information influence factor to generate a diagnostic accuracy correlation, where the diagnostic accuracy correlation is used to represent improving research quality and improving the research process.

[0010] According to yet another aspect of the present application, an electronic device is characterized by including: a first processor; and a memory for storing executable instructions of the first processor; wherein the first processor is configured to execute the above-mentioned quality control method for clinical research implementation based on a large language model by executing the executable instructions.

[0011] According to still another aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a second processor, the above-mentioned quality control method for clinical research implementation based on a large language model is realized.

[0012] A method and system for quality control of clinical research implementation based on large language models provided by this application. The server obtains target clinical research projects, user consultation questions, and scientific research project information of different institutions, etc. Subsequently, the target clinical project information is filtered, formatted, segmented, vectorized, classified, and associated in sequence to form an external knowledge base. For the scientific research information of the target hospital and other institutions, clinical research management impact factors and user individual information impact factors are generated respectively. The former analyzes the classification progress, etc., and the latter is obtained through feature extraction and clustering. Then, using the target large language model and the external knowledge base, a quality control report is generated by constructing text block vectors, similarity indexes, and medical knowledge graphs. Finally, relevant information is processed in combination with the two impact factors to generate a diagnostic accurate correlation, effectively assisting in improving the quality of clinical research and optimizing the process.

[0013] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. Brief Description of the Drawings

[0014] Figure 1 The flowchart shows a method for quality control of clinical research implementation based on large language models provided by an embodiment of this application;

[0015] Figure 2 The structural schematic diagram shows a device for quality control of clinical research implementation based on large language models provided by an embodiment of this application;

[0016] Figure 3 The roadmap shows a prompt design based on retrieval augmented generation technology provided by an embodiment of this application. Detailed Embodiments

[0017] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only for the purpose of illustration and explanation of the present invention, and are not used to limit the present invention.

[0018] The following is combined with Figure 1 to describe a method for quality control of clinical research implementation based on large language models according to an exemplary embodiment of this application. It should be noted that the following application scenarios are only shown for the convenience of understanding the spirit and principle of this application, and the embodiments of this application are not limited in this regard. On the contrary, the embodiments of this application are applicable to any applicable scenario.

[0019] In one embodiment, this application also proposes a method and system for quality control of clinical research implementation based on large language models. Figure 1 Schematically shows a schematic flowchart of a method for quality control of clinical research implementation based on large language models according to an embodiment of this application. As Figure 1 shown, this method is applied to a server and includes:

[0020] S101, obtain the target clinical research project information, the consulting questions of the target users regarding clinical research, the scientific research project management information of the target hospital, other clinical research project information, a preset large language model, and a training sample set.

[0021] In one implementation, assume that a research project on "Clinical Research of a New Targeted Therapy Drug for Breast Cancer" is underway. The project information is as follows: Project design: A randomized controlled trial design is adopted, and patients are divided into an experimental group and a control group. The experimental group receives the new targeted drug treatment, and the control group receives the traditional treatment plan. Implementation details: The dosage, frequency of use, treatment cycle, etc. of the drug are specified in detail. For example, patients in the experimental group take the new targeted drug orally at a dose of 100 mg per day, and 28 consecutive days is a cycle, with a total of 4 cycle treatments. During this period, various physical examinations and laboratory tests need to be carried out regularly to evaluate the safety and effectiveness of the drug. Recruitment criteria: Clearly define the age range of patients (such as 18 - 70 years old), the pathological type of breast cancer (such as HER2-positive breast cancer), the disease stage (such as stage II - III), previous treatment history (such as not having received the same type of targeted treatment), etc., to ensure that suitable research subjects are included. Intervention measures: In addition to drug treatment, it also includes dietary and exercise guidance for patients. For example, it is recommended that patients maintain a balanced diet during treatment, appropriately increase protein intake, and perform at least 3 moderate aerobic exercises per week, with each exercise lasting more than 30 minutes, to improve the body's immunity and tolerance to treatment. Outcome evaluation: The main outcome indicators are set as the progression-free survival (PFS) and objective response rate (ORR) of patients, which are evaluated through regular imaging examinations (such as breast ultrasound and chest CT examinations every 8 weeks) and laboratory tests (such as detecting the levels of tumor markers); the secondary outcome indicators include the quality of life score of patients (using a breast cancer quality of life questionnaire, evaluated before treatment, every 4 weeks during treatment, and after treatment), the incidence of drug adverse reactions, etc.

[0022] Case Report Form (CRF): Patient's personal information: Record basic information such as the patient's name, gender, age, contact information, ID number, etc. to identify the patient and manage follow-up. Medical history: Record in detail the patient's past medical history, such as whether they have chronic diseases like hypertension, diabetes, heart disease, etc., and the treatment process of previous breast cancer, including information such as the surgical method, chemotherapy regimen, radiotherapy dose and time. This information is of great significance for evaluating the patient's overall health status and tolerance to the test drug in this trial. Treatment records: Record information such as the time of each treatment, drug dosage, and physical reactions after drug use in this clinical study. For example, whether the patient has adverse reactions such as nausea, vomiting, rash, etc. after taking the new targeted drug for the first time, as well as the severity and duration of the adverse reactions, in order to adjust the treatment plan in a timely manner or take corresponding symptomatic treatment measures. Follow-up situation: Record information such as the time of each follow-up, examination results at the time of follow-up (such as imaging examination reports, laboratory test data), the patient's self-perception and symptom changes. For example, if the patient reports symptoms such as fatigue and joint pain during the follow-up, this information will be comprehensively analyzed together with the examination results to judge the disease progression and treatment effect of the patient. During the implementation of the project, as patients are enrolled and treatment progresses, this information will be updated in real time. For example, when a new patient meets the recruitment criteria and is enrolled, their personal information, medical history, etc. will be immediately entered into the system; after each treatment and follow-up, the corresponding treatment records and follow-up situations will also be updated in a timely manner so that researchers can keep track of the latest progress of the project at any time.

[0023] In addition, the consulting questions of target users regarding clinical research include but are not limited to: "Can I take health supplements to enhance immunity while taking this new targeted drug?"; "If I have a mild rash during the treatment process, do I need to stop taking the drug? How should I deal with it?". And the questions asked by researchers include but are not limited to: "In the process of data analysis, how can we more accurately evaluate the differences in drug responses among different subgroups of patients (such as different age and different pathological type subgroups)?", "For patients with severe adverse reactions, in addition to following the established treatment process, are there any other special countermeasures or indicators that need further observation?".

[0024] Next, taking a large general hospital as an example, its scientific research project information is as follows: Basic medical research projects: such as "Molecular Biology Research on the Pathogenesis of Breast Cancer", which aims to explore the key molecular targets and signaling pathways in the occurrence and development of breast cancer, and provide a theoretical basis for subsequent drug development and treatment strategy formulation. Clinical research projects: In addition to the above-mentioned "Clinical Research on New Targeted Therapeutic Drugs for Breast Cancer", there is also "Clinical Application Research on Early Screening Technology for Breast Cancer", which mainly studies the accuracy and effectiveness of different screening methods (such as mammography, breast ultrasound, magnetic resonance imaging, etc.) in the early diagnosis of breast cancer, and how to optimize the screening process and improve the early diagnosis rate. Translational medicine research projects: such as "From laboratory to clinic: Validation and application research of new biomarkers for breast cancer", focus on clinical verification of potential biomarkers discovered in basic research, and explore their application value in breast cancer diagnosis, treatment monitoring and prognosis evaluation, and promote the transformation of basic research results into clinical practice.

[0025] The progress information of each scientific research project is as follows: "Molecular Biology Research on the Pathogenesis of Breast Cancer": The cell experiment and animal model construction phases have been completed, and large-scale sample data analysis is currently underway. It is expected that it will take another 6 months to complete the data analysis and paper writing. "Clinical Research on New Targeted Therapeutic Drugs for Breast Cancer": As mentioned above, the patient recruitment and treatment phase is in progress. 50 patients have been recruited and the first treatment cycle of 20 patients has been completed. According to the plan, the recruitment of all patients will be completed within the next 3 months, and treatment and follow-up observations will continue. "Clinical Application Research on Early Screening Technology for Breast Cancer": A prospective screening study of 1,000 women has been completed, and the collected data are being sorted and analyzed. Preliminary results show that combined screening of mammography and ultrasound has high sensitivity and specificity in the early diagnosis of breast cancer, but the screening threshold and process still need to be further optimized. The project is expected to complete the formulation of the final report and results transformation plan within 1 year.

[0026] Examples of similar target clinical research data collected from different regions and institutions with different medical levels are as follows: Similar clinical research data from an internationally renowned cancer research center: A similar clinical study on breast cancer targeted therapy drugs conducted by this center adopted a different drug dose escalation scheme and follow-up time point settings compared to the target hospital. In terms of drug dose escalation, a more aggressive strategy was adopted, with a higher starting dose but a slower dose escalation rate to observe the tolerance and efficacy of patients to high-dose drugs. The follow-up time points were set for comprehensive evaluations at 4 weeks, 8 weeks, 12 weeks, etc. after treatment, rather than a major evaluation every 8 weeks as in the target hospital. Through the analysis of the data from this center, it can provide different trial design ideas and references for potential risks and benefits for the research in the target hospital. The patient population included was more diverse in terms of ethnic composition. In addition to Caucasians and Asians, there were also a relatively large number of African American patients. It was found that there were certain differences in pharmacokinetics and the incidence of adverse reactions among patients of different ethnic groups, which provided new considerations for the target hospital when analyzing patient data. It is necessary to further study the impact of ethnic factors on drug efficacy and safety and appropriately adjust the inclusion criteria or conduct stratified analysis in subsequent studies.

[0027] Similar clinical research data from domestic primary medical units: A relatively small-scale clinical study on breast cancer targeted therapy drugs conducted by a primary hospital, although with a limited sample size, had certain characteristics in patient management and data collection. This hospital used telemedicine technology to conduct follow-up management of patients. Patients uploaded relevant symptom and sign information through a mobile application at home, and doctors conducted online evaluations and guidance. This method improved patient compliance and the convenience of follow-up to a certain extent. The target hospital can draw on this telemedicine follow-up model to optimize its own patient management process, especially for patients living in remote areas or with limited mobility, to improve their enthusiasm and feasibility for participating in clinical research. In terms of data analysis, this primary hospital used simple and easy-to-understand statistical methods to analyze the preliminary data. Although the analysis depth was limited, it could quickly provide some basic information on efficacy and safety. The target hospital can refer to its data analysis ideas and use simple statistical methods to conduct preliminary screening and analysis of the data in the early stage of the project to timely discover potential problems and trends and provide directions for subsequent in-depth analysis.

[0028] The pre-set large language model is the Llama model, which is a large language model based on the Transformer architecture. The following is an explanation of its related features: The Llama model belongs to an autoregressive language model and can predict the next word or character based on the given previous text, thus generating a coherent text sequence. It performs well in natural language processing tasks such as text generation, question answering systems, machine translation, etc. It has a multi-layer neural network structure. For example, the Llama model contains dozens or even hundreds of Transformer layers. Each layer continuously extracts and transforms the semantic and syntactic information of the input text. As the number of layers increases, the model can learn more complex language patterns and semantic relationships. Its core structure is the Transformer architecture, including components such as the multi-head attention mechanism (Multi-Head Attention), feed-forward neural network (Feed-Forward Neural Network), and layer normalization (Layer Normalization). The multi-head attention mechanism allows the model to simultaneously focus on different parts of the input text and capture the semantic associations between different positions; the feed-forward neural network further performs non-linear transformations on the information processed by the attention mechanism to enhance the model's expressive ability; layer normalization helps to stabilize the training process of the model, improve the training efficiency and the generalization ability of the model. In this clinical research implementation quality control system, the powerful language processing ability of the Llama model is used to analyze and process the collected clinical research project information, user consultation questions, etc. For example, when generating a quality control report, through the processing of text block information and the construction of a knowledge graph, the Llama model can combine the medical knowledge and language patterns it has learned to provide accurate and professional answers for users and assist researchers in quality control and the improvement of research processes.

[0029] The training sample set includes a large amount of text data such as medical literature, clinical research reports, case analyses, and experience summaries of medical experts. For example: Medical literature: It covers articles on basic research, clinical treatment, drug development, etc. of breast cancer published in well-known domestic and international medical journals, including various types such as research papers, reviews, and case reports. Clinical research reports: Detailed reports of completed and ongoing breast cancer clinical research worldwide are collected, including trial design, patient recruitment, treatment regimens, result analysis, etc. These reports provide rich practical cases and data support for the model, helping the model learn result prediction and problem-solving methods under different trial conditions. Case analyses: Detailed case analysis materials of breast cancer patients from major hospitals, including patients' clinical manifestations, diagnosis processes, problems and solutions during treatment, prognosis, etc., enabling the model to deeply understand various situations and coping strategies in actual clinical work. Experience summaries of medical experts: Renowned experts in the field of breast cancer are invited to write their experiences and insights in clinical practice and research work, including unique understandings of the disease, treatment techniques, research ideas, etc. These experience summaries inject professional clinical wisdom and research ideas into the model, improving the accuracy and reliability of the model in dealing with complex medical problems. By training with these rich and diverse training sample sets, the pre-set large language model can learn professional knowledge and language expression patterns in the medical field, so as to better process target clinical research project information, answer target users' consultation questions, and assist in generating quality control reports and analyzing and improving research quality in subsequent applications.

[0030] S102, process the pre-set large language model based on the training sample set to generate a target large language model.

[0031] In one implementation, obtain any number of data features from the training sample set. Randomly select an indefinite number of data features from the training sample set. These data features cover rich clinical research-related information, including disease symptom descriptions, treatment method details, drug reaction records, patient demographic information, clinical research process steps, etc. For example, in the training samples of breast cancer clinical research, the data features involve imaging features of different pathological types of breast cancer, changing trends of blood test indicators of patients at different treatment stages, quality of life assessment indicators of patients under specific treatment regimens, etc. By widely obtaining these different types of data features, it lays a foundation for subsequent model training and ensures that the model can learn various knowledge and rules in the field of clinical research.

[0032] Generate a sampling ratio based on the quantity of each data feature in the training sample set. While ensuring data diversity, reasonably select representative data for model training. If a particular type of data feature has a large quantity in the sample set, its proportion in the sampling process will be correspondingly reduced to avoid the model overfitting to this type of data; conversely, if some data features are scarce but crucial for clinical research, such as the special symptom manifestations of certain rare diseases or the initial application cases of new treatment methods, then their weights will be appropriately increased in the sampling ratio to ensure that the model can fully learn this key information. For example, in the research training sample for a certain rare cancer, although the number of cases is limited, the unique gene mutation information and special treatment response data contained in these cases will be key considerations when determining the sampling ratio to ensure that the model can accurately capture these key features.

[0033] Process the training sample set based on the sampling ratio to generate a preset number of sampling features. This process is actually a targeted screening and extraction of the original training samples, so that the sampling features finally used for model training can not only reflect the overall feature distribution of the training sample set, but also highlight the key points and crucial information. For example, in a training sample set containing a large amount of clinical research data on different diseases, if the set number of sampling features is 1000, then according to the previously determined sampling ratio, 300 sampling features will be selected from the cardiovascular disease trial data, 400 sampling features will be selected from the oncology disease trial data, and 300 sampling features will be selected from the neurological disease trial data. Moreover, the sampling features in each category are carefully selected and can represent the key aspects of the clinical research of this type of disease, such as the comparison data of different drug treatment effects in cardiovascular diseases, the correlation data between tumor marker changes and treatment plans in oncology diseases, and the relationship data between neurological function evaluation indicators and treatment progress in neurological diseases.

[0034] Process any data feature with each sampling feature to generate multiple groups of data sets, where each data set contains a preset number of data samples, and at least one data sample includes identification information. In each data set, there are a preset number of data samples, and at least one data sample has identification information. These identification information are usually used to mark whether the data samples are related to the risk factors affecting the clinical research project. For example, in a data set for a clinical study of diabetes drugs, the data samples include information such as the patient's blood glucose monitoring records, drug dosage and time, and whether adverse reactions such as hypoglycemia or hyperglycemia occur. Among them, the data samples of patients with serious adverse reactions will carry specific identification information, indicating that they are risk factors affecting the clinical research project. By constructing data sets in this way, the model can learn the associations between different data features and their relationships with risk factors during the training process, thereby improving the model's ability to identify and judge potential problems in clinical research.

[0035] Train a preset large language model based on the data samples in multiple groups of data sets to generate a trained large language model. During the training process, the model continuously adjusts its own parameters to learn the language patterns and semantic relationships in the input data, and tries to predict various information related to clinical research, such as the development trend of diseases, the evaluation of treatment effects, the risks and problems that occur, etc. For example, the model will learn how to generate reasonable diagnostic suggestions based on the patient's symptom descriptions and medical history information, or predict the research results and potential risk points according to the progress of the clinical study. Through training with a large number of data samples, the model gradually improves its language processing and knowledge application abilities in the field of clinical research.

[0036] Process the trained large language model based on the validation sample set to generate validation results. If the data samples containing identification information in the validation results are risk factors characterizing the impact on clinical research projects, then the trained large language model is used as the target large language model. The validation sample set also contains rich clinical research data, but is different from the training sample set in terms of data distribution and specific cases, and is used to test the generalization ability and accuracy of the model on unseen data. During the validation process, the model analyzes and predicts the data in the validation sample set and compares the results with the known true situation. If the data samples containing identification information in the validation results are confirmed as risk factors characterizing the impact on clinical research projects, it means that the model can accurately identify these key information, and the trained large language model will be selected as the target large language model for subsequent quality control tasks in clinical research implementation, such as generating quality control reports, analyzing the risks and problems of research projects, etc.; conversely, if the model has a low accuracy in identifying risk factors during the validation process, then the model needs to be further adjusted and optimized, such as re-adjusting training parameters, increasing the amount of training data, or improving data preprocessing methods, and then training and validating again until the model meets the expected performance standards.

[0037] S103, process the target clinical research project information to generate an external knowledge base.

[0038] In one implementation, filter and format the target clinical research project information to generate clinical research project information in a target format. The source of the target clinical research project information is extensive and the formats are diverse, including forms such as electronic documents, database records, and clinical reports. First, these information need to be filtered and formatted. For example, in a project information containing a large number of patient medical records and research records, filter out irrelevant information such as medical device operation logs and non-critical administrative records, and only retain data related to the core content of clinical research, such as the basic condition description of patients, treatment plans, and examination results. At the same time, convert data in different formats into a standard format, such as unifying the date format to "YYYY-MM-DD" and the text encoding to UTF-8, etc., to ensure the consistency and readability of the data, thereby generating clinical research project information in a target format.

[0039] Perform text segmentation and vectorization on the clinical research project information in the target format to generate text block information. For the clinical research project information in the target format after formatting, use a specific text segmentation tool for processing. It can be segmented according to the logical pauses of natural language, such as line breaks, full stops, question marks, exclamation marks, etc., or the text can be divided into manageable chunks according to a pre-set chunk size. For example, in a detailed clinical research report, segment it by paragraph or divide it into chunks of 500 characters each. Then, use open-source text vectorization tools or models such as Word2Vec, BERT (taking Llama-3.1-Nemotron-70B-Instruct as an example in the text) to convert these text chunks into corresponding vector representations (embeddings). These vectors can capture the semantic information of the text, enabling the computer to better understand and process the text content, and then generating text block information.

[0040] Classify the text block information to generate the real-time achievement progress of the clinical research project and the research difficulty information of the clinical research project. The classification basis can include dimensions such as disease type, research stage, treatment method, etc. For example, in a comprehensive clinical research project, classify the text blocks related to disease diagnosis into one category and those related to the drug treatment process into another category. Through this classification method, different aspects of the clinical research project can be clearly sorted out, thus generating the real-time achievement progress of the clinical research project and the research difficulty information of the clinical research project. For example, in a clinical study of a cancer drug, identify the improvement in the efficacy of the drug in a specific patient group as the real-time achievement progress, and regard the side effect management and countermeasures of the drug as the research difficulty information. Process the real-time achievement progress of the clinical research project and the research difficulty information of the clinical research project to generate text feature correlation information. Analyze the relationship between the real-time achievement progress and the research difficulty information of the clinical research project, and mine the text feature correlation information therein. For example, find that the improvement of a certain treatment method is closely related to the breakthrough of specific research difficulties, or that certain real-time achievements are due to the change of specific factors. Through statistical analysis, semantic association analysis and other methods, find the internal connection between these information, providing a basis for the subsequent construction of the knowledge base.

[0041] Generate an external knowledge base based on the text feature correlation information. Based on the generated text feature correlation information, integrate relevant text blocks, classification information, achievement progress, research difficulties, etc. to construct a structured external knowledge base. This knowledge base can be stored and managed in the form of databases, knowledge graphs, etc., so as to be quickly retrieved and utilized in subsequent research processes. For example, in an external knowledge base constructed in the form of a knowledge graph, with diseases, treatment methods, research institutions, etc. as nodes and their association relationships as edges, various information in clinical research projects is organically organized to provide comprehensive and accurate knowledge support for clinical researchers, assisting them in research decision-making, problem-solving, quality control, and other tasks. It can effectively transform the original target clinical research project information into a highly available and valuable external knowledge base, providing an important knowledge foundation for the entire clinical research implementation quality control system.

[0042] S104, process the scientific research project information of the target hospital to generate a clinical research management impact factor.

[0043] In one implementation, process the scientific research project information of the target hospital to generate the scientific research project classification information of the target hospital and the progress information of each scientific research project. The scientific research project information of the target hospital covers many projects in multiple fields such as basic research, clinical research, and translational medicine research. First, analyze aspects such as the research topics, research methods, and application directions of these projects and classify them into different categories. For example, in medical research, it can be divided into categories such as cardiovascular disease research, tumor disease research, and nervous system disease research. At the same time, for each scientific research project, record its progress information in detail, including whether the project is in the startup phase, data collection phase, data analysis phase, or result summary phase, and clarify the key time nodes and completion status of each phase. For example, for a clinical research project on cardiovascular diseases, its progress information shows that patient recruitment has been completed, mid-term data collection and analysis are underway, and it is expected to complete the preliminary result evaluation within the next 6 months.

[0044] The progress information of scientific research projects of the same category and scientific research projects of corresponding categories are processed separately to generate abnormal text features of different scientific research projects. For scientific research projects of the same category, their progress information and process details are compared to find out the differences from the normal process or expected results, and the relevant text descriptions are extracted as abnormal text features. For example, in multiple drug clinical research projects for tumor diseases, if it is found that some projects have an adjustment frequency or adjustment range that exceeds the normal range during the drug dose adjustment stage, then these special descriptions of dose adjustment will be extracted as abnormal text features. Similarly, for scientific research projects of different categories but with certain correlations, such as drug development projects for cardiovascular diseases and interventional treatment research projects for cardiovascular diseases, similar comparative analysis is also performed to dig out abnormal text features in terms of research paths, changes in key indicators, etc.

[0045] In one embodiment, an example of an abnormal text feature is as follows: Project progress: In a "clinical study of a new targeted therapy for breast cancer", it was planned to conduct a comprehensive imaging review of patients in the 12th week after the start of treatment to evaluate tumor shrinkage, but the review of 70% of patients was actually completed in the 14th week. The relevant text record is "The imaging review progress of the clinical study of targeted treatment for breast cancer has been delayed, and only 70% of patients have been reviewed, which is 2 weeks later than planned." Research methods: In a "research on early screening technology for breast cancer" project, it was originally stipulated to use standardized breast ultrasound examination procedures and image analysis methods, but in actual operation, some operators did not set the frequency of the ultrasound probe in accordance with the standard, and did not strictly judge according to the established BI-RADS classification standards during image analysis. The text description is "The breast cancer screening ultrasound examination process is not standardized, the probe frequency is set incorrectly, and the image analysis does not follow the BI-RADS standards." Data quality: In a study on the relationship between breast cancer gene expression and prognosis, it was found that some patients' gene chip test data had duplicate records and data entry errors. For example, the records of some gene expression values ​​at different time points were exactly the same and obviously did not conform to biological logic. The record text was "Breast cancer gene expression data had duplicate records and entry errors, affecting data accuracy and the reliability of research results."

[0046] In another implementation, examples of abnormal text features are as follows: In terms of project progress: In a clinical research project of a diabetes drug, it was originally planned to conduct mid-term data analysis in the 6th month after the number of enrolled patients reached 200. However, the actual progress showed that when it was the 8th month, the number of enrolled patients had just reached 150, and the mid-term analysis had not yet started. The relevant text description was like "The progress of patient recruitment is lagging behind, affecting the time node of mid-term analysis", and this is an abnormal text feature in terms of project progress. In terms of research methods: In a research on interventional treatment of cardiovascular diseases, according to the established protocol, a specific angiography technique should be used to evaluate the treatment effect. However, in actual operation, another alternative technique was used for some cases, and no reasonable explanation was given in the research records. The text description was "The application of angiography technique in the research process does not conform to the protocol and there is no reasonable explanation", and this is an abnormal text feature in terms of research methods. In terms of data quality: In a certain research project on tumor gene detection, it was found that there were a large number of missing values in the gene sequencing data of some samples, and these missing values were not properly processed in data statistics and analysis. The corresponding text description was "There are a large number of missing values in the gene sequencing data and they are not processed, affecting the data reliability", which belongs to the abnormal text feature in terms of data quality.

[0047] Process the abnormal text features of different scientific research projects to generate the abnormal common features of scientific research projects and the information on the occurrence times of the abnormal common features. After collecting and organizing the abnormal text features of each scientific research project, use data analysis and text mining techniques to find the common parts among them. For example, in multiple different scientific research projects, it was found that there were irregularities in the sample selection process, or there were similar errors in the application of data analysis methods, and these constituted the abnormal common features. At the same time, count the number of times each abnormal common feature appears in different projects to evaluate its universality and severity. For example, the abnormal common feature of irregular sample selection appeared 6 times in 10 related scientific research projects, indicating that this is an issue that needs to be focused on.

[0048] In one implementation, examples of abnormal common features are as follows: Sample management: In multiple scientific research projects related to breast cancer, including "Pathological Research on Different Subtypes of Breast Cancer" and "Research on Prediction of Postoperative Recurrence Risk of Breast Cancer", there are problems with insufficient sample collection. For example, in the research plan, it is stipulated that each research group should collect at least 100 samples, but in some research groups of multiple projects, only 70 - 80 samples are actually collected. The text description is "The sample collection in breast cancer-related research fails to reach the planned quantity, affecting the statistical power of the research and the reliability of the conclusions". Statistical analysis: In some research projects on the evaluation of breast cancer treatment effects, such as "Comparative Study on the Efficacy of Chemotherapy Combined with Radiotherapy and Chemotherapy Alone for Breast Cancer" and "Analysis of the Effects of Different Endocrine Therapy Regimens for Breast Cancer", important confounding factors such as the age and underlying diseases of patients are wrongly ignored when analyzing the survival rate data of patients. The text record is "The age, underlying diseases and other confounding factors are not considered in the survival rate analysis during the evaluation of breast cancer treatment effects, resulting in biased results". Ethical review: In clinical research projects involving breast cancer patients, such as "Clinical Research on Neoadjuvant Therapy for Breast Cancer" and "Intervention Trial on Rehabilitation Therapy for Breast Cancer", there are unclear privacy protection measures for patients during the ethical review process. For example, during the patient information registration and data storage process, there is no detailed description of how to encrypt and restrict access to patients' sensitive information. The text description is "The privacy protection measures for patients are missing during the ethical review of breast cancer clinical research, and there are potential information security risks".

[0049] In another implementation, examples of abnormal common features are as follows: Sample management problems: In multiple clinical research projects on different diseases, there are situations where the sample storage conditions do not meet the standard requirements. For example, in some infectious disease research and genetic disease research projects, it is found that the temperature control during sample transportation and storage is unstable, resulting in sample damage. The text description such as "Poor temperature control of sample storage affects sample quality" is an abnormal common feature in terms of sample management. Statistical analysis errors: In several drug R & D projects, inappropriate statistical test methods are wrongly used when conducting drug efficacy comparison analysis. For example, in some antibiotic R & D projects and anti-cancer drug R & D projects, analysis of variance for multiple group comparisons should have been used, but a simple t-test was used instead, resulting in biased results. The text description is "The wrong statistical test method is selected in the drug efficacy comparison analysis", which constitutes an abnormal common feature in terms of statistical analysis. Ethical review loopholes: In some scientific research projects involving human trials, whether it is clinical research on neurological diseases or orthopedic diseases, there are problems with incomplete ethical review documents or non-standard approval processes. For example, the informed consent forms of some projects lack key information or the signing dates are unclear. The text description is "There are defects in the ethical review documents, and the informed consent forms are incomplete or not signed in a standard manner", which becomes an abnormal common feature in terms of ethical review.

[0050] Process the abnormal common features of scientific research projects and the information on the occurrence times of abnormal common features to generate the clinical feasibility information corresponding to different scientific research project categories for the target hospital. Based on the abnormal common features of scientific research projects and the information on their occurrence times, combined with medical professional knowledge and clinical practice experience, evaluate the feasibility of each scientific research project category in actual clinical applications. For project categories with more serious abnormal common features, their clinical feasibility is lower; while project categories with fewer and relatively minor abnormal situations have higher clinical feasibility. For example, in some basic medical research projects, if it is found that there are many defects in the experimental design and these problems are common in multiple projects, then the feasibility of transforming such projects into clinical applications will be questioned; on the contrary, for some clinical research projects with strict design and standardized implementation, if there are only a small number of resolvable abnormal situations, their clinical feasibility is relatively high.

[0051] Process the clinical feasibility information corresponding to different scientific research project categories for the target hospital to generate a clinical research management impact factor. The impact factor is a quantitative indicator that can reflect the potential impact degree of different scientific research project categories on clinical outcomes. For example, project categories with high clinical feasibility will be assigned a higher impact factor weight, indicating that they have a greater promoting effect on improving clinical treatment effects, improving diagnostic methods, etc.; while project categories with low clinical feasibility will get a lower impact factor, suggesting that these projects need to be further improved and optimized to increase their contribution to clinical outcomes. Through the clinical research management impact factor, hospital managers and researchers can more intuitively understand the value and potential risks of each scientific research project, so as to reasonably allocate resources, give priority to supporting projects that have a positive impact on clinical outcomes, promote the close combination of hospital scientific research and clinical practice, and improve the overall medical level and research quality.

[0052] S105, process the other clinical research project information to generate a user individual information impact factor.

[0053] In one implementation, feature extraction processing is performed on other clinical research project information to generate target features that match the user's physiological information. Among them, the target features that match the user's physiological information are used to characterize the change trends of physiological indicators at different treatment stages, the difference information of physiological characteristics among different user groups, and the correlation information among physiological indicators. From other clinical research project information, physiological indicator data of a large number of breast cancer patients at different treatment stages (such as before surgery, after surgery, during chemotherapy, during radiotherapy, during endocrine therapy, etc.) are collected, including but not limited to the levels of tumor markers (such as CA15-3, CEA, etc.) in the blood, hormone levels (such as estrogen, progesterone, androgen, etc.), blood routine indicators (such as white blood cell count, red blood cell count, platelet count, etc.), liver and kidney function indicators (such as alanine aminotransferase, aspartate aminotransferase, creatinine, urea nitrogen, etc.), and physical symptom manifestations (such as pain level, fatigue, sleep quality, etc.). Through data analysis and statistical methods, the change trends of these physiological indicators with the treatment stage are extracted. For example, the CA15-3 indicator will show a brief increase and then gradually decrease at the initial stage of chemotherapy, which is a target feature of a change trend. At the same time, the physiological characteristic differences among different user groups such as age, gender, race, and physical baseline conditions (such as whether suffering from other chronic diseases) are compared. For example, the hormone level fluctuations of young female breast cancer patients during treatment are significantly different from those of elderly female patients, which reflects the difference information of the physiological characteristics of user groups. In addition, the correlation between different physiological indicators is analyzed. For example, it is found that there is a certain positive correlation between the decrease in white blood cell count and the chemotherapy drug dose in some breast cancer patients, which is the correlation information among physiological indicators. These change trends, difference information, and correlation information together constitute the target features that match the user's physiological information.

[0054] Process the target features that match the user's physiological information to generate user groups with similar physiological response patterns and the common influencing features of each user group. Using algorithms such as cluster analysis, according to the extracted target features, breast cancer patients with similar trends in physiological index changes, similar manifestations of physiological feature differences, and related physiological index correlations are divided into the same user group. For example, patients who experience similar degrees of nausea, vomiting and other gastrointestinal reactions during chemotherapy and have similar rates of decline in tumor markers in the blood are grouped together. For each user group, further analyze its common influencing features, which include specific treatment plan combinations (such as the combination method of a certain chemotherapy drug and an endocrine therapy drug), lifestyle factors (such as regular exercise habits, dietary preferences, etc.), genetic factors (such as whether carrying specific breast cancer-related gene mutations), etc. For instance, in a certain user group, most patients have adopted the AC-T chemotherapy plan and have similar BRCA1 gene mutation conditions, and these are the common influencing features of this group. Process the common influencing features of each user group to generate the feature analysis results of the target user group. For each divided user group, deeply analyze the comprehensive impact of its common influencing features on the treatment effect and the patient's physiological state. Through statistical analysis and clinical research experience, evaluate the relationships between these features and aspects such as the disease progression, survival period, and quality of life of the patients. For example, for the above-mentioned user group with a specific chemotherapy plan and gene mutation, it is found through analysis that the 5-year survival rate of the patients in this group is relatively low after treatment, and the probability of experiencing severe adverse reactions during the treatment process is relatively high, and this is the feature analysis result of this target user group.

[0055] Generate the influencing factors of user individual information based on the feature analysis results of the target user group. According to the feature analysis results of each target user group, assign corresponding weights or scores to each group to quantify its impact degree on the overall user physiological information, thereby generating the influencing factors of user individual information. For example, for those user groups that have a greater impact on the treatment effect and the patient's physiological state, such as groups with high-risk gene mutations and poor treatment responses, their influencing factors will be higher; while for groups with relatively small impacts, such as groups with mild symptoms and good treatment effects, the influencing factors will be lower. These influencing factors can be used in subsequent clinical research and medical decision-making to evaluate the characteristics and risks of different patient groups, provide important reference bases for the formulation of personalized treatment plans and the adjustment of research directions, so as to improve the quality and efficiency of breast cancer clinical research and better serve the treatment and rehabilitation of patients.

[0056] The user individual information influence factor comprehensively represents the degree and pattern of the physiological information of different user groups being affected by various factors in a specific clinical research context. It reflects the potential driving effect of the associations between different treatment stages, user group characteristics, and physiological indicators, mined from extensive clinical research data, on the change of an individual's physiological state. Specifically, a high influence factor indicates that certain common characteristics of a specific user group or specific change trends in a treatment stage have a strong impact on physiological information, and these factors need to be focused on in clinical practice. For example, when formulating a treatment plan or predicting disease progression, the physiological characteristics and change laws represented by these high influence factors should be fully considered; while a low influence factor indicates that the corresponding factors have a relatively weak impact on physiological information, and the priority can be appropriately reduced when allocating resources and determining research priorities, thereby providing a quantitative reference basis for clinical decision-making and research direction adjustment, and assisting in achieving personalized medicine and precision medicine research.

[0057] S106, process the consultation questions of the target user regarding clinical research based on the target large language model and the external knowledge base to generate a quality control report.

[0058] In one implementation, process the text block information based on the target large language model to generate text block vectors. Refer to Figure 3 As shown, first, the target large language model (such as the Llama model) receives the processed text block information in a breast cancer clinical research project. This text block information covers various aspects such as the disease characteristics of breast cancer, details of treatment plans, patient case data, and research progress. For example, a text block describes that "in a breast cancer endocrine therapy study, tamoxifen was used to treat patients in a specific age group for 5 years, during which estrogen levels and breast ultrasound examination results were regularly detected, and some patients had side effects such as endometrial thickening". The model will analyze and encode these text blocks, converting them into vector representations with semantic and syntactic information, that is, text block vectors. In this process, the model uses the language patterns and knowledge learned from training on a large number of medical texts to identify key concepts and semantic relationships in the text and map them into a high-dimensional vector space, so that text blocks with similar semantics are close in the vector space, providing a basis for subsequent processing.

[0059] Process the text block vectors and text feature correlation information to generate a similarity index. The text feature correlation information includes the correlation relationships between different elements in breast cancer clinical research, such as the correlation between different treatment methods and disease prognosis, the correlation between specific gene expressions and drug responses, etc. Through specific algorithms and model mechanisms, calculate the similarity between the text block vectors and these correlation information, so as to generate a similarity index. For example, if a text block vector involves the relationship between targeted therapy for breast cancer and specific gene mutations, then when generating the similarity index, it will focus on other research text blocks related to this gene mutation and text blocks for evaluating the efficacy of targeted therapy, quantify and index the degree of association between them, so as to quickly retrieve and utilize this relevant information. Process the similarity index to generate a medical knowledge graph. In the knowledge graph of breast cancer, use different aspects of the disease (such as pathological type, stage, metastasis status, etc.), treatment methods (surgery, chemotherapy, radiotherapy, endocrine therapy, targeted therapy, etc.), patient characteristics (age, gender, gene background, etc.) as nodes, and use the relationships between them (such as a certain treatment method is applicable to breast cancer at a specific stage, certain gene characteristics are related to the efficacy of specific treatments, etc.) as edges or labels. For example, there will be a strongly associated edge between HER2-positive breast cancer and trastuzumab targeted therapy, and there is an associated edge between older breast cancer patients and certain conservative treatment methods. Integrate the scattered breast cancer-related knowledge into a structured network, comprehensively display the knowledge system in the field of breast cancer clinical research, and provide more comprehensive and in-depth background knowledge support for answering users' consultation questions.

[0060] Based on the medical knowledge graph, process the consultation questions of the target user regarding clinical research to generate a quality control report. When the target user raises a consultation question regarding breast cancer clinical research, such as "How to evaluate the risk of cardiotoxicity in breast cancer patients after receiving combined chemotherapy and radiotherapy?", the system will process it based on the constructed medical knowledge graph. First, locate the relevant nodes and edges related to chemotherapy, radiotherapy, breast cancer, and cardiotoxicity in the knowledge graph, extract the relevant knowledge and information, and then use the generation ability of the target large language model, combined with these knowledge and the specific requirements of the question, to generate a quality control report. The report will elaborate in detail on the currently commonly used cardiotoxicity evaluation indicators (such as specific indicators for myocardial enzyme spectrum detection and echocardiography), risk factors summarized in different studies (such as the patient's underlying heart disease, cumulative dose of chemotherapy drugs, etc.), and corresponding prevention and monitoring recommendations (such as the time interval for regular cardiac function examinations, measures to be taken when specific symptoms appear, etc.), providing users with accurate, professional, and targeted answers, assisting clinical researchers and patients to better understand and address relevant issues in breast cancer clinical research, and ensuring the quality control and scientific management of the research process.

[0061] S107. Process the target clinical research project information and the quality control report based on the clinical research management impact factor and the user individual information impact factor to generate a diagnostic accurate relevance.

[0062] In one implementation, feature screening and matching processing are performed on the target clinical research project information based on the clinical research management impact factor to generate scientific research projects of the same type or related fields as the target clinical research project information in the target hospital and the progress information of each scientific research project. Taking the target clinical research project "Clinical Research on a New Targeted Therapy Drug for Breast Cancer" as an example, the clinical research management impact factor will screen among numerous scientific research projects in the target hospital. For projects of the same type, such as the research on other breast cancer targeted therapy drugs, the similarities in research design, drug mechanism of action, patient inclusion criteria, etc. will be concerned; for related field projects, like the research on the pathophysiological mechanism of breast cancer and the imaging evaluation research of breast cancer treatment, the relevance with the current targeted therapy trial in the knowledge system and research process will be analyzed. For example, the hospital has a "Clinical Research on the Combination of Another New Targeted Drug and Radiotherapy for Breast Cancer", which has comparable aspects with the target project in terms of drug combination method and treatment cycle setting; there is also a "Basic Research on the Influence of the Tumor Microenvironment of Breast Cancer on Targeted Therapy", although it belongs to basic research, it is of great significance for understanding the effect of targeted therapy. Through these screenings and matches, the progress information of these related projects is obtained. For example, the "Clinical Research on the Combination of Another New Targeted Drug and Radiotherapy for Breast Cancer" has completed patient recruitment and is in the mid-term efficacy evaluation; the "Basic Research on the Influence of the Tumor Microenvironment of Breast Cancer on Targeted Therapy" is in the data collection and preliminary analysis stage.

[0063] Process the scientific research projects in the target hospital that are of the same type or in related fields as the target clinical research project information, as well as the progress information of each scientific research project, to generate target abnormal text features. Conduct a detailed analysis of the above-selected projects and progress information. In the "Clinical Research on the Combination of Another New Type of Targeted Drug and Radiotherapy for Breast Cancer", it is found that some patients have severe radiation dermatitis during radiotherapy, exceeding the expected incidence of adverse reactions, which is recorded as "In the clinical trial of the combination of a new type of targeted drug and radiotherapy for breast cancer, some patients have severe radiation dermatitis, and the incidence is higher than expected, which may affect patient compliance and the evaluation of treatment effects"; in the "Basic Research on the Influence of the Tumor Microenvironment of Breast Cancer on Targeted Therapy", it is found that there is a problem of inconsistent processing times for tumor tissue samples during the sample collection process, which is described as "After the sample collection for the study of the tumor microenvironment of breast cancer, the tissue processing time span is large, which may lead to data deviation and affect the reliability of research conclusions". These are the generated target abnormal text features. Process the quality control report based on the target abnormal text features to generate abnormal warning information for scientific research projects. When generating a quality control report involving the safety and effectiveness evaluation of breast cancer targeted therapy, the system will give a warning based on the above abnormal text features. For example, it prompts "When evaluating the current effect of breast cancer targeted therapy, it is necessary to consider the potential interference of the abnormal situation of radiation dermatitis in the combination radiotherapy trial on the overall treatment experience and results of patients, as well as the impact of data unreliability caused by sample processing problems in the tumor microenvironment study on the basic theory support", and this is the abnormal warning information for scientific research projects, reminding researchers to be cautious about these potential problems when analyzing and applying relevant research results.

[0064] Based on the user individual information influence factor, perform feature screening and matching processing on the target clinical research project information to generate the change indicators of the physiological indicators of the target user at different treatment stages and the change indicators of the physiological indicators of other users with the same clinical research project as the target user at different treatment stages. Assume that the target user is a 50-year-old patient with positive estrogen receptor and in the endocrine therapy stage of breast cancer. From the target clinical research project information, extract the changes in physiological indicators at different treatment stages (such as the initial stage, middle stage, and long-term maintenance stage of endocrine therapy) according to the user individual information influence factor, including changes in estrogen levels, tumor markers (such as CA15-3, CEA), and bone density. At the same time, collect the change data of the physiological indicators of other similar users (such as patients aged 45-55 years, with positive estrogen receptor and using the same endocrine therapy drugs) with the same clinical research project as the target user at the corresponding treatment stages. For example, it is found that the estrogen level of the target user decreases rapidly in the initial stage of endocrine therapy, but the CA15-3 index shows a small increase in the short term and then stabilizes; among other similar users, some patients have a similar trend of estrogen level change, but there are differences in the CA15-3 fluctuation.

[0065] Based on the change indicators of various physiological indicators of other users in the same clinical research project as the target user at different treatment stages, the change indicators of various physiological indicators of the target user at different treatment stages are processed to generate target physiological impact factors. By statistically analyzing the physiological index change data of the target user and other similar users, such as using regression analysis and other methods, explore the relationships between factors such as age, treatment time, and basic health status and the changes in physiological indicators, and quantify the target physiological impact factors. For example, it is found that age is positively correlated with the decline rate of estrogen levels in the early stage of endocrine therapy, and at a specific drug dose, this correlation is more significant. This relationship is transformed into a target physiological impact factor to numerically reflect the degree and direction of the impact of this factor on the changes in the patient's physiological state.

[0066] Based on the target physiological impact factor, the quality control report is processed to generate early warning information for abnormal physiological information of the target user. During the generation of the quality control report, if it involves the evaluation of the treatment effect or suggestions for treatment plan adjustment of this patient, the system will give an early warning according to the target physiological impact factor. For example, it is prompted that "the estrogen level of this patient drops too fast in the early stage of endocrine therapy. Considering its age factor, it is necessary to closely monitor the change of bone density to prevent the increased risk of osteoporosis. At the same time, continuously observe the fluctuation of the CA15-3 index. If there is an abnormal upward trend, further examinations should be carried out in a timely manner to adjust the treatment plan to ensure the curative effect". This is the early warning information for abnormal physiological information of the target user, providing personalized attention focuses and decision-making references for clinicians for this patient.

[0067] Process the early warning information for abnormal scientific research projects and the early warning information for abnormal physiological information of the target user to generate a high diagnostic relevance. For example, if similar abnormal change trends of physiological indicators caused by drug adverse reactions are found in multiple breast cancer endocrine therapy-related research projects, and these factors are closely related to the research quality and process, such as sample processing problems affecting the accuracy of research results and abnormal changes in patient physiological indicators affecting the reliability of treatment effect evaluation, etc., then the diagnostic relevance will be relatively high. This indicates that it is necessary to comprehensively review and optimize and improve aspects such as the drug dose adjustment strategy, sample collection and processing specifications, patient monitoring indicators and frequencies of the current research project, such as optimizing the dose titration plan of endocrine therapy drugs, standardizing the processing process after sample collection, increasing the monitoring frequency and depth of specific physiological indicators, etc., to improve the quality and efficiency of breast cancer clinical research, ensure the scientific nature and clinical application value of research results, and achieve effective improvement and promotion of the research process.

[0068] This application obtains various information from the server, such as target clinical research projects, user consultation questions, scientific research project information of target hospitals and other institutions. The information of the target clinical research project is processed to generate an external knowledge base, which goes through processes such as filtering and formatting, text segmentation and vectorization, classification, and generating text feature correlation information. At the same time, the information of the scientific research projects of the target hospital is processed to generate a clinical research management impact factor, and its classification, progress, abnormal text features, etc. are analyzed; the information of other clinical research projects is processed to generate a user individual information impact factor, including steps such as feature extraction and clustering analysis. Based on the target large language model and the external knowledge base, the user consultation questions for clinical research are processed to generate a quality control report, which involves text block vectors, similarity indexes, and the construction of a medical knowledge graph. Finally, combining the two impact factors to process the project information and the quality control report to generate a diagnostic accurate correlation, aiming to improve the research quality, improve the research process, and provide strong support for the effective implementation and quality assurance of clinical research.

[0069] In one implementation, as Figure 2 shown, this application also provides a clinical research implementation quality control device based on a large language model, including:

[0070] An acquisition module 201, configured to acquire target clinical research project information, consultation questions of target users regarding clinical research, scientific research project management information of target hospitals, information of other clinical research projects, a preset large language model, and a training sample set. Among them, the target clinical research project information is updated in real time, and the scientific research project information of the target hospital is used to represent the classification information, management requirement information, and progress information of each scientific research project of the target hospital. The information of other clinical research projects includes similar target clinical research data of institutions with different regions and different medical levels;

[0071] A processing module 202, configured to process the preset large language model based on the training sample set to generate a target large language model; process the target clinical research project information to generate an external knowledge base; process the scientific research project information of the target hospital to generate a clinical research management impact factor; process the information of other clinical research projects to generate a user individual information impact factor; process the consultation questions of target users regarding clinical research based on the target large language model and the external knowledge base to generate a quality control report; process the target clinical research project information and the quality control report based on the clinical research management impact factor and the user individual information impact factor to generate a diagnostic accurate correlation, where the diagnostic accurate correlation is used to represent improving research quality and improving the research process.

[0072] The computer-readable storage medium provided by the above embodiments of the present application and the method for implementing quality control of clinical research based on large language models provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein. Each embodiment in the present application is described in a related manner. For the same or similar parts between the embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. For the related parts, reference can be made to the partial description of the embodiments of the method for implementing quality control of clinical research based on large language models above.

Claims

1. A method for quality control of clinical research implementation based on a large language model, characterized in that: include: Obtain target clinical research project information, target user's consulting questions regarding clinical research, target hospital's scientific research project management information, other clinical research project information, preset large language model, and training sample set, wherein the target clinical research project information is updated in real time, the target hospital's scientific research project information is used to characterize the target hospital's scientific research project classification information, management requirements information, and progress information of each scientific research project, and the other clinical research project information includes similar target clinical research data of institutions in different regions and with different medical levels; Processing the preset large language model based on the training sample set to generate a target large language model; Processing the target clinical research project information to generate an external knowledge base; Processing the scientific research project information of the target hospital to generate clinical research management impact factors; Processing the other clinical research project information to generate user individual information impact factors; Processing the consulting questions of the target user regarding clinical research based on the target large language model and the external knowledge base to generate a quality control report; The target clinical research project information and the quality control report are processed based on the clinical research management influencing factor and the user individual information influencing factor to generate a diagnostic accuracy correlation, wherein the diagnostic accuracy correlation is used to characterize the improvement of research quality and the improvement of research process.

2. The method according to claim 1, characterized in that The preset large language model is processed based on the training sample set to generate a target large language model, including: Get any number of data features in the training sample set; Generate a sampling ratio based on the number of each data feature in the training sample set; Processing the training sample set based on the sampling ratio to generate a preset number of sampling features; Processing based on any data feature and each sampling feature to generate multiple data groups, wherein each data group includes a preset number of data samples, and at least one data sample includes identification information; Training the preset large language model based on data samples in the multiple data groups to generate a trained large language model; Processing the trained large language model based on the verification sample set to generate a verification result; If the data sample containing identification information in the verification result represents a risk factor affecting the clinical research project, the trained large language model is used as the target large language model.

3. The method according to claim 1, characterized in that Process the target clinical research project information to generate an external knowledge base, including: Filtering and formatting the target clinical research project information to generate clinical research project information in a target format; Performing text segmentation and vectorization processing on the clinical research project information in the target format to generate text block information; Classify and process the text block information to generate real-time progress of clinical research projects and research difficulty information of clinical research projects; Processing the real-time progress of the clinical research project and the research difficulty information of the clinical research project to generate text feature correlation information; An external knowledge base is generated based on the text feature correlation information.

4. The method according to claim 1, characterized in that Process the scientific research project information of the target hospital to generate clinical research management impact factors, including: Processing the scientific research project information of the target hospital to generate classification information of the scientific research projects of the target hospital and progress information of each scientific research project; Process the progress information of scientific research projects of the same category and corresponding categories respectively to generate abnormal text features of different scientific research projects; Process the abnormal text features of different scientific research projects to generate abnormal common features of scientific research projects and the number of occurrences of abnormal common features; Processing the abnormal common characteristics of the scientific research projects and the occurrence frequency information of the abnormal common characteristics to generate clinical feasibility information corresponding to different scientific research project categories of the target hospital; The clinical feasibility information corresponding to different scientific research project categories of the target hospital is processed to generate clinical research management impact factors.

5. The method according to claim 1, characterized in that Processing the other clinical research project information to generate user individual information impact factors includes: Performing feature extraction processing on the other clinical research project information to generate target features that match the user's physiological information, wherein the target features that match the user's physiological information are used to characterize the changing trends of physiological indicators at different treatment stages, the difference information of physiological characteristics of different user groups, and the correlation information between physiological indicators; Processing target features that match the user's physiological information to generate user groups with similar physiological response patterns and common impact features for each user group; Process the common influencing features of each user group to generate feature analysis results for the target user group; The user individual information influencing factor is generated based on the characteristic analysis result of the target user group.

6. The method according to claim 3, characterized in that Based on the target large language model and the external knowledge base, the target user's consulting questions for clinical research are processed to generate a quality control report, including: Processing the text block information based on the target large language model to generate a text block vector; Processing the text block vector and the text feature correlation information to generate a similarity index; Processing the similarity index to generate a medical knowledge graph; Based on the medical knowledge graph, the target user's consulting questions regarding clinical research are processed to generate a quality control report.

7. The method according to claim 6, characterized in that The target clinical research project information and the quality control report are processed based on the clinical research management influencing factor and the user individual information influencing factor to generate a diagnostic accuracy correlation, including: Performing feature screening and matching processing on the target clinical research project information based on the clinical research management impact factor, generating scientific research projects of the same type or in related fields as the target clinical research project information and progress information of each scientific research project in the target hospital; Process the scientific research projects of the same type or in related fields as the target clinical research project information in the target hospital and the progress information of each scientific research project to generate target abnormal text features; Processing the quality control report based on the target abnormal text features to generate abnormal warning information for the scientific research project; Performing feature screening and matching processing on the target clinical research project information based on the user individual information influencing factors, generating change indicators of various physiological indicators of the target user at different treatment stages and change indicators of various physiological indicators of other users of the same clinical research project as the target user at different treatment stages; Based on the change indicators of various physiological indicators of other users of the same clinical research project as the target user at different treatment stages, the change indicators of various physiological indicators of the target user at different treatment stages are processed to generate target physiological influencing factors; Processing the quality control report based on the target physiological influencing factor to generate abnormal warning information of physiological information of the target user; The abnormal warning information of the scientific research project and the abnormal warning information of the physiological information of the target user are processed to generate a diagnosis accuracy correlation.

8. A clinical research implementation quality control device based on a large language model, characterized in that: The device comprises: An acquisition module is used to acquire target clinical research project information, consulting questions of target users regarding clinical research, scientific research project management information of target hospitals, other clinical research project information, a preset large language model, and a training sample set, wherein the target clinical research project information is updated in real time, the scientific research project information of the target hospital is used to characterize the scientific research project classification information, management requirements information, and progress information of each scientific research project of the target hospital, and the other clinical research project information includes similar target clinical research data of institutions in different regions and with different medical levels; A processing module is used to process a preset large language model based on a training sample set to generate a target large language model; process the target clinical research project information to generate an external knowledge base; process the scientific research project information of the target hospital to generate a clinical research management impact factor; process the other clinical research project information to generate a user individual information impact factor; process the target user's consulting questions regarding clinical research based on the target large language model and the external knowledge base to generate a quality control report; process the target clinical research project information and the quality control report based on the clinical research management impact factor and the user individual information impact factor to generate a diagnostic accuracy correlation, wherein the diagnostic accuracy correlation is used to characterize the improvement of research quality and the improvement of research process.

9. An electronic device, characterized in that: include: a first processor; and a memory for storing executable instructions of the first processor; Wherein, the first processor is configured to execute the clinical research implementation quality control method based on a large language model as described in any one of claims 1 to 7 by executing the executable instructions.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the second processor, the method for quality control of clinical research implementation based on a large language model as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Medical quality evaluation system based on large language model

    CN116978524A

  • Automatic text dialogue method and device for clinical test, equipment and medium

    CN117390145A

  • Knowledge graph-driven medical large model diagnosis method

    CN118280562A

  • Clinical test scheme information analysis system based on large language model

    CN118824439A

  • Intelligent subject evaluation method and device based on multi-dimensional data

    CN119274797A

Cited By

  • Image report quality control method and device based on large model, and storage medium

    CN121439071A