Over-specification medication evidence evaluation method and system, electronic equipment and storage medium

By combining a large language model with cue word engineering, we can achieve methodological quality assessment and validity grading of off-label drug use evidence, which solves the problems of low assessment efficiency and inconsistent results in existing technologies, and realizes rapid and accurate evidence assessment.

CN120878281APending Publication Date: 2025-10-31GENERAL HOSPITAL OF SOUTHERN THEATRE COMMAND OF PLA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511027371.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-07-23
Filing Date
2025-07-24
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing technologies for evaluating evidence of off-label drug use are inefficient and subject to high subjective bias, leading to inconsistent evaluation results and affecting scientific decision-making regarding off-label drug use.

Method used

By employing a large language model combined with cue word engineering, we can achieve methodological quality assessment and validity grading of evidence. By identifying the research type and obtaining cue words, we can conduct automated evaluation, including background, role, task, and quality control components, to ensure the rigor and standardization of the evaluation.

Benefits of technology

It enables rapid and efficient evaluation of evidence for off-label drug use, reducing the review time for a single article from 2-3 hours to 10 minutes, significantly improving scoring consistency, and solving the problems of low efficiency and high subjective bias in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120878281A_ABST
    Figure CN120878281A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of medical literature analysis, and discloses an out-of-specification medication evidence evaluation method, which comprises the following steps of: determining a research type of evidence, and obtaining a first prompt word corresponding to the research type; based on the first cue word, performing methodological quality evaluation on the evidence by adopting a large language model to obtain an evaluation result; when the evaluation result is medium-high-quality evidence, based on a preset second prompt word, performing evidence validity grading on the evidence by adopting a large language model to obtain a grading result; and obtaining an evaluation result according to the evaluation result and the grading result. The method can quickly and efficiently evaluate the medicine evidence exceeding the specification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical literature analysis technology, specifically relating to a method, system, electronic device, and storage medium for evaluating evidence of off-label drug use. Background Technology

[0002] In clinical practice, off-label use refers to situations where the dosage, indications, route of administration, or patient population of a drug exceeds the scope of the instructions approved by the drug regulatory authority. The rationality and safety of off-label use are highly dependent on sufficient evidence. A lack of scientific basis may increase the risk of adverse reactions in patients, lead to medical disputes, or even violate the principles of evidence-based medicine.

[0003] Currently, assessing the strength of evidence for off-label drug use primarily relies on a systematic review of medical literature. Clinical pharmacists, physicians, or relevant professionals must combine the "Expert Consensus on Off-Label Drug Use," drug instructions, and the latest clinical guidelines to evaluate the methodological quality (such as risk of bias, blinding), effectiveness (effectiveness and evidence grading according to the Thomson grading definition), and safety (such as adverse reactions) of the evidence. Ultimately, a comprehensive recommendation for evidence-based evaluation of off-label drug use (strong recommendation, weak recommendation, or no recommendation) should be formed.

[0004] The manual evaluation model suffers from significant efficiency bottlenecks, and inconsistencies in results may arise due to differences in experience among reviewers, affecting the objectivity of the evaluation. Therefore, it is necessary to improve the efficiency of off-label drug use evidence evaluation to provide efficient support for scientific decision-making regarding off-label drug use. Summary of the Invention

[0005] The purpose of this invention is to provide a method, system, electronic device, and computer-readable storage medium for evaluating evidence of off-label drug use, which can quickly and efficiently evaluate evidence of off-label drug use.

[0006] The first aspect of this invention discloses a method for evaluating evidence of off-label drug use, comprising:

[0007] Determine the research type of the evidence and obtain the first cue word corresponding to that research type;

[0008] Based on the first prompt word, a large language model is used to evaluate the methodological quality of the evidence and obtain the evaluation results.

[0009] When the evaluation result is medium to high quality evidence, the evidence is graded based on the preset second prompt word using a large language model to obtain the grading result; and the evaluation result is obtained based on the evaluation result and the grading result.

[0010] In some embodiments, determining the research type of the evidence and obtaining the first cue word corresponding to the research type includes:

[0011] Based on preset prompts, a large language model is used to determine the research type of the evidence;

[0012] The first prompt word is determined based on the scale corresponding to the research type.

[0013] In some embodiments, a large language model is used to classify the evidence validity. After obtaining the classification result, the large language model is also used to perform GRADE classification on the evidence to obtain GRADE classification result, and the GRADE classification result is added to the classification result.

[0014] In some embodiments, it also includes:

[0015] When the evaluation result is low-quality evidence, the manual evaluation result obtained by manually reviewing the evidence will be compared with the evaluation result. If the first preset condition is met, a re-evaluation will be conducted.

[0016] And / or,

[0017] The grading results are manually reviewed, and a reassessment is conducted when the review results meet the second preset condition.

[0018] In some embodiments, obtaining an evaluation result based on the evaluation result and the grading result includes:

[0019] A large language model is used to extract descriptions of security from the evidence to obtain security information;

[0020] An assessment result is obtained based on the evaluation results, the classification results, and the security information.

[0021] In some embodiments, both the first and second prompt words include a background section, a role section, a task section, and a quality control section. The background section is used to clarify the research type, the core purpose of the scale, and the research objective. The role section is used to guide the large language model to process the literature from the perspective of a professional evaluator. The task section is used to provide the large language model with an operation manual to standardize the evaluation process. The quality control section is used to set a validation framework for the output of the large language model to ensure the rigor of the evaluation.

[0022] In some embodiments, the evaluation results include item scores, evidence summaries, and quality levels; the grading results include item scores, evidence summaries, and validity grading.

[0023] A second aspect of this invention discloses an off-label drug use evidence evaluation system, comprising:

[0024] The research type module is used to determine the research type of the evidence and obtain the first prompt word corresponding to the research type.

[0025] The methodology quality assessment module is used to assess the methodology quality of the evidence based on the first prompt word using a large language model, and to obtain the assessment result.

[0026] The evidence validity grading module is used to grade the evidence validity based on the second prompt word and a large language model when the evaluation result is medium to high quality evidence, and to obtain the grading result.

[0027] The evaluation results module is used to obtain evaluation results based on the evaluation results and the grading results.

[0028] A third aspect of the present invention discloses an electronic device, including a memory storing executable program code and a processor coupled to the memory; the processor calls the executable program code stored in the memory to execute the off-label drug use evidence assessment method disclosed in the first aspect.

[0029] The fourth aspect of the present invention discloses a computer-readable storage medium storing a computer program, wherein the computer program causes a computer to perform the off-label drug use evidence evaluation method disclosed in the first aspect.

[0030] The beneficial effects of this invention are as follows: First, the research type of the evidence is determined. Then, based on the research type, a methodological quality assessment of the evidence is initiated using a cue-based large language model. If the evidence is of medium to high quality, the evidence validity is further graded using the cue-based large language model. Finally, the evaluation result is derived by combining the methodological quality assessment and the evidence grading. This allows for the rapid and efficient evaluation of off-label drug use evidence. Attached Figure Description

[0031] The accompanying drawings illustrate specific examples of the technical solutions described in this invention and, together with the detailed embodiments, form part of the specification, serving to explain the technical solutions, principles, and effects of this invention.

[0032] Unless otherwise specified or defined, the same reference numerals in different figures represent the same or similar technical features, and different reference numerals may be used to represent the same or similar technical features.

[0033] Figure 1 This is a flowchart of the off-label drug use evidence evaluation method according to an embodiment of the present invention;

[0034] Figure 2 This is a schematic diagram of the off-label drug use evidence evaluation system according to an embodiment of the present invention;

[0035] Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0036] Unless otherwise specified or defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. When combined with the technical solutions of the invention in a real-world scenario, all technical and scientific terms used herein may also have meanings corresponding to the purpose of achieving the technical solutions of the invention. The terms "first," "second," etc., used herein are merely for distinguishing names and do not represent a specific number or order. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0037] It should be noted that when a component is considered "fixed" to another component, it can be directly fixed to the other component or there can be an intervening component; when a component is considered "connected" to another component, it can be directly connected to the other component or there can be an intervening component; when a component is considered "mounted" on another component, it can be directly mounted on the other component or there can be an intervening component; when a component is considered "placed" on another component, it can be directly placed on the other component or there can be an intervening component.

[0038] Unless otherwise specified or defined, the terms "described" or "the" as used herein refer to the technical features or technical content mentioned or described prior to the relevant section, which may be the same as or similar to the technical features or technical content mentioned herein. Furthermore, the terms "comprising" and "having," and any variations thereof, as used herein, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or apparatus.

[0039] To rapidly and efficiently evaluate evidence of off-label drug use, this invention utilizes a dynamic prompting engineering-driven large language model to automate the entire process of "methodological quality assessment - effectiveness grading - safety description" for off-label drug use evidence for the first time. This reduces the review time for a single article from 2-3 hours to 10 minutes, while significantly improving scoring consistency, thus addressing the technical pain points of traditional methods such as low efficiency, incomplete coverage, and high subjective bias.

[0040] This invention discloses a method for evaluating evidence of off-label drug use, which can be implemented by computer programming. The subject executing this method can be an electronic device such as a computer, laptop, or tablet, or a control chip embedded in an electronic device. This invention does not limit the subject to this.

[0041] like Figure 1 As shown, this embodiment includes the following steps:

[0042] Step S100: Determine the research type of the evidence and obtain the first cue word corresponding to the research type;

[0043] The evidence is medical literature. In this example, the literature comes from the evidence-based medicine data provided by clinicians when submitting off-label drug use applications. The title of the literature is "Randomized Phase 2 Trial of Telitacicept in Patients With IgA Nephropathy With Persistent Proteinuria".

[0044] The types of evidence-based research include: meta-analysis, randomized controlled trials (RCTs), and non-RCTs. Different research types require different scales for methodological quality assessment. Specifically, meta-analysis corresponds to the AMSTAR scale, RCT to the Jadard scale, and non-RCT to the Minors scale.

[0045] In this embodiment, after the evidence is uploaded to the large language model, the model determines the research type of the evidence based on preset prompts. Specifically, the preset prompts are: "Please act as an expert in evidence-based medicine and help me determine which research type the evidence belongs to: meta-analysis, RCT, or non-RCT study, and provide the basis for your judgment."

[0046] Considering the significant differences in evaluation dimensions, indicator definitions, and focuses among different scales (such as the JADAD scale, MINARS scale, and AMSRAR scale), this embodiment employs a large language model based on cue engineering for methodological quality assessment, designing different cue words for different scales. Designing dedicated cue words for different scales allows the large language model to accurately pinpoint the core evaluation elements of each scale, avoiding "fuzzy judgments" caused by generic cue words, ensuring that the evaluation closely adheres to the scale logic, reducing interference from irrelevant information, and improving the accuracy of the results. Dedicated cue templates are constructed for RCT and non-RCT literature respectively (Jadad scale adapted to RCT, MINARS scale dynamically identifies non-RCT control groups), overcoming the limitation of a single model being incompatible with multiple study designs.

[0047] Therefore, after obtaining the research type of the evidence, this embodiment retrieves the first prompt word from a pre-set prompt word table or prompt word database according to the scale corresponding to the research type. Specifically, to achieve the objective, this embodiment uses the Chain of Thought (CoT) technology to force the model to execute a standardized process of "research type determination → scale clause matching → original text evidence location → structured scoring → quality grading," ensuring a strong correlation between the scoring results and the literature evidence. The designed first prompt word includes four parts: background, role, task, and quality control. The background part clarifies the research type, the core purpose of the scale, and the research objective; the role part guides the large language model to process the literature from the perspective of a professional evaluator; the task part provides an operation manual for the large language model, standardizing the evaluation process; and the quality control part sets a validation framework for the output of the large language model, ensuring the rigor of the evaluation.

[0048] For example, for the Jadad scale, the first prompt word is:

[0049]

background

[0050] Please conduct a systematic quality assessment of the target literature based on the provided Jadad scale assessment criteria. The assessment targets are RCT research papers.

[0051]

Role

[0052] As an AI assessment system with expertise in clinical research methodologies, you need to: (1) accurately analyze the various assessment dimensions of the Jadad scale; (2) systematically identify methodological features in the literature; and (3) execute standardized quality assessment procedures.

[0053]

Task

[0054] Compare each item with the Jadard scoring criteria to identify evidence fragments in the literature that meet or do not meet the criteria. Requirements: (1) Cite the original text (in Chinese and English); (2) Mark the specific chapter / section; (3) Provide the basis for the compliance judgment; (4) Generate a structured scoring table.

[0055] Methodological quality review: (1) Highlighting the advantages of the research design; (2) Identifying methodological limitations; (3) Proposing improvement suggestions.

[0056] Quality Control

[0057] Establish a dual verification mechanism: (1) Perform logical consistency verification after the initial evaluation; (2) Perform compliance verification of the scoring criteria.

[0058] Implement iterative optimization: initiate secondary literature review for questionable entries.

[0059] For the MINORS scale, the first cue word is:

[0060]

background

[0061] Please conduct a quality assessment of observational research literature using the provided MINARS scale assessment criteria. First, identify the research design type (whether there is a control group), and then follow the corresponding assessment procedure.

[0062]

Role

[0063] As an AI evaluation system for observational studies, you need to: (1) accurately distinguish the type of study design (with or without a control group); (2) dynamically adjust the evaluation dimensions (8 basic items / 12 extended items); and (3) implement a standardized quality evaluation process.

[0064]

Task

[0065] Study type identification: (1) Confirm the control group setting through methodological description; (2) Activate the corresponding evaluation path: For studies without a control group, perform evaluation of items 1-8; for studies with a control group, perform evaluation of items 1-12.

[0066] Systematic evaluation: (1) Match evidence against the scale standards item by item; (2) Mark supporting textual evidence (Chinese and English); (3) Generate a visual scoring result table.

[0067] Methodological quality assessment: (1) Identify the methodological strengths and weaknesses; (2) Propose improvement suggestions.

[0068] Quality Control

[0069] Establish a dual verification mechanism: (1) Perform logical consistency verification after the initial evaluation; (2) Perform compliance verification of the scoring criteria.

[0070] Implement iterative optimization: initiate secondary literature review for questionable entries.

[0071] As shown above, the first prompt word is designed to include a background section, a role section, a task section, and a quality control section. The background section clarifies the research type, the core purpose of the scale, and the research objective; it helps the large language model quickly locate the professional field and application scenario of the task, avoiding evaluation direction deviations caused by misunderstandings of basic information such as "the scope of application of the Jadad scale" and "the differences between RCTs and other research types," ensuring the model's analysis is based on the correct context from the outset. The role section assigns the model a clear professional role, guiding it to handle the task with a mindset consistent with domain consensus. This avoids the model interpreting literature from the perspective of general text analysis rather than evidence-based medicine evaluation, ensuring that judgments on professional issues such as "whether randomization is appropriate" and "whether blinding hides group assignments" conform to standard understanding within the field, enhancing the professionalism and authority of the evaluation. The task section breaks down the core requirements in detail, essentially providing the model with an operation manual. This ensures the model does not omit key steps, avoiding information gaps due to task ambiguity; and it standardizes the structure of the evaluation output, facilitating subsequent manual verification or data integration. The quality control section clarifies the dual verification mechanism and iterative optimization process, essentially setting a validation framework for the model's output. The model is guided to proactively avoid logical contradictions when generating evaluation results, reducing basic errors at the output end. Through iterative optimization, the model tends to mark questionable content as requiring secondary verification rather than making arbitrary judgments, ensuring the rigor of the evaluation.

[0072] Step S200: Based on the first prompt word, use a large language model to evaluate the methodological quality of the evidence and obtain the evaluation results;

[0073] Based on the initial prompt word designed in the model, input the prompt word into DeepSeek-R1, which automatically performs a methodological quality assessment of the evidence and obtains the assessment results. It is important to emphasize that large language models are not limited to DeepSeek-R1.

[0074] Methodological quality assessment, in this context, involves the role of an expert in evidence-based medicine quality evaluation, using the AMSTAR / Jadad / MINORS scales to assess the quality of medical literature. The input to the large language model consists of the full text of the literature and the scale details. The output evaluation results include item scores, evidence summaries, and quality grades. Setting the output of the large language model as a combination of item scores, evidence summaries, and quality grades has clear and specific significance, fully replicating the thought process of human evaluation and ensuring that the output of the large language model conforms to the professional standards of methodological evaluation. Specifically, the output item scores ensure that the evaluation strictly follows the requirements of standardized tools, avoiding vague overall judgments and guaranteeing the granularity and operability of the evaluation. Evidence summaries address the issue of "traceability," deriving from the objective content of the literature, facilitating human review, and providing clear traceability for subsequent disputes. The quality grade is a general judgment of the research methodology quality.

[0075] The designed first cue word can guide the step-by-step reasoning of the large language model through the CoT method: research type determination → scale clause matching → evidence extraction → clause scoring → quality grading.

[0076] In this embodiment, the clinical evidence submitted was a Phase II clinical trial, classified as an RCT, and evaluated using the Jadad scale. The evaluation results showed that the DeepSeek-R1 score was 5–6, with all scores indicating high quality. Overall, the DeepSeek-R1 results were consistent with the human assessment.

[0077]

[0078] Step S300: When the evaluation result is medium to high quality evidence, the evidence is graded based on the preset second prompt word using a large language model to obtain the grading result; the evaluation result is obtained based on the evaluation result and the grading result.

[0079] Evidence will only be graded for validity if the evaluation result is medium-to-high quality evidence (i.e., medium-quality or high-quality evidence). Low-quality evidence will be subject to manual review; if the review still classifies it as low-quality evidence, it will not be accepted.

[0080] In this embodiment, the large language model is used as an expert in evaluating the Thomson grading of evidence validity in evidence-based medicine, assessing the validity of evidence based on the Thomson grading scale. The input includes the full text of the literature and the detailed rules of the Thomson grading scale. The output grading results include item scores, evidence summaries, and validity ratings. The role of each part is explained in the analysis of the large language model's output results during methodological quality assessment.

[0081] To guide the large language model using the CoT method, the reasoning process follows a step-by-step procedure: research type determination → scale item matching → evidence extraction → item scoring → evidence validity grading. The designed second cue words also include background, role, task, and quality control sections. Specifically:

[0082]

background

[0083] Based on the provided Thomson grading scale assessment criteria, a systematic quality assessment of the target literature was conducted.

[0084]

Role

[0085] As an AI assessment system with Thomson grading expertise, you need to: (1) accurately identify the research type of the target literature; (2) accurately analyze the various assessment dimensions of the Thomson grading scale; and (3) execute a standardized Thomson grading assessment process.

[0086]

Task

[0087] The effectiveness assessment requires identifying the following categories by comparing each item with the Thomson grading scale: whether the treatment is effective, whether the evidence supports its effectiveness, whether the effectiveness is controversial, and whether the treatment is ineffective. The evidence assessment requires classifying the treatment into categories A, B, C, and no evidence, according to the grading scale requirements. Requirements: (1) Cite the original text (Chinese-English bilingual version); (2) Mark the specific chapter / section; (3) Provide the basis for the compliance judgment; (4) Generate a structured rating table.

[0088] A review of the Thomson rating scale: (1) highlighting the strengths of the study design; (2) identifying the limitations that led to a lower rating; and (3) proposing improvements.

[0089] Quality Control

[0090] Establish a dual verification mechanism: (1) Perform logical consistency verification after the initial evaluation; (2) Perform compliance verification of the scoring criteria.

[0091] Implement iterative optimization: initiate secondary literature review for questionable entries.

[0092] Similar to the four parts in the first prompt, the background part in the second prompt directly anchors the core premise of "based on the Thomson grading scale," avoiding confusion between different grading systems in the large language model. Meanwhile, the statement "to conduct a systematic quality assessment of the target literature" clarifies the application scenario of the task, ensuring that the model focuses on the specific logic of the Thomson scale from the outset, laying a unified standard foundation for subsequent analysis. The role part accurately identifies the research type, analyzes the scale assessment dimensions, and executes standardized procedures, guiding the model to process literature from the perspective of a professional evaluator, rather than simply text matching. This is especially true for judging ambiguous items in the Thomson scale, which aligns more closely with the scale's professional connotations, reducing bias caused by a lack of understanding of the scale. The task part breaks down validity and evidence levels into specific judgment targets, avoiding model confusion between the two. Through the design of dual assessment dimensions and four output requirements, the assessment process is standardized. The quality control part forces the model to trace back to the details of the literature for questionable items, reducing misjudgments caused by incomplete information extraction.

[0093] The evidence validity rating result of this embodiment is:

[0094]

[0095] The results of the evidence validity rating show that the large language model based on prompting engineering can accurately assess the validity and level of evidence, and the basis for the judgment is clear and sufficient.

[0096] After obtaining the evaluation results and the grading results, the two results are integrated according to the preset rules or templates to obtain the assessment result. The specific form of the assessment result is not limited. For example, the assessment result obtained according to the preset rules is: recommended or not recommended.

[0097] In some implementations, a large language model is used to classify the evidence's validity. After obtaining the classification result, the large language model is also used to perform a GRADE classification of the evidence, obtaining a GRADE classification result, which is then used as another type of classification result. That is, the resulting classification results include both Thomson classification results and GRADE classification results. The process of performing GRADE classification is the same as that of Thomson classification; only the prompt words need to be modified accordingly based on the differences between the two classifications, which will not be elaborated further here.

[0098] In some implementations, when the evaluation result is low-quality evidence, manual review is mandatory. The obtained manual evaluation result is compared with the evaluation result output by the large language model. When the first preset condition is met (e.g., the difference between the two is greater than 8%), a re-evaluation is conducted and a second verification is performed, which improves efficiency while ensuring the reliability of the conclusion.

[0099] The grading results are also manually reviewed. When the review results meet the second preset condition (for example, when the grading results are determined to be inaccurate), a reassessment and a second verification are conducted.

[0100] In some implementations, a large language model is also used to extract descriptions of security from the evidence to obtain security information; then, the quality assessment results, validity evidence ratings, and security description results are combined to obtain an assessment result, such as recommendation or non-recommendation.

[0101] Specifically, the large language model reads the off-label filing evidence and extracts descriptions related to safety. The prompt used is: "Please act as an evidence-based medicine expert and help me extract the safety-related descriptions from this document. The content should be concise and do not omit important information."

[0102] The security information obtained in this embodiment is:

[0103]

[0104]

[0105] Key safety conclusions:

[0106] 1. Items with significant differences (* marked):

[0107] o Injection site reaction (treatment group >60% vs control group 0%)

[0108] o Decreased immunoglobulin levels (IgG / IgM treatment group >28% vs control group ≤7%)

[0109] 2. Undifferentiated core terms:

[0110] o Risk of serious infection (upper respiratory tract infection rate comparable to the control group)

[0111] No deaths or kidney failure events occurred (none occurred in any of the three groups).

[0112] 3. Special findings:

[0113] Despite a significant decrease in immunoglobulin levels, the risk of infection was not increased (P>0.05), which may be related to the drug's mechanism of action.

[0114] In summary, this embodiment first determines the research type of the evidence, then, based on the research type, initiates a methodological quality assessment of the evidence using a cue-based large language model. If the evidence is of medium to high quality, the cue-based large language model is used to classify the validity of the evidence. Finally, the methodological quality assessment and evidence classification are combined to arrive at an evaluation conclusion: recommended or not recommended. Furthermore, descriptions related to safety are extracted from the evidence to obtain a safety assessment.

[0115] like Figure 2 As shown, based on the above-mentioned off-label drug use evidence assessment method, this embodiment of the invention discloses an off-label drug use evidence assessment system, including:

[0116] Research type module 600 is used to determine the research type of the evidence and obtain the first prompt word corresponding to the research type;

[0117] The methodology quality assessment module 610 is used to perform a methodology quality assessment on the evidence based on the first prompt word using a large language model, and obtain the assessment result.

[0118] The evidence validity grading module 620 is used to grade the evidence validity based on the second prompt word and a large language model when the evaluation result is medium to high quality evidence, and obtain the grading result.

[0119] The evaluation result module 630 is used to obtain the evaluation result based on the evaluation result and the grading result.

[0120] like Figure 3 As shown, an embodiment of the present invention discloses an electronic device, including a memory 401 storing executable program code and a processor 402 coupled to the memory 401;

[0121] Specifically, the processor 402 calls the executable program code stored in the memory 401 to execute the off-label drug use evidence evaluation method described in the above embodiments.

[0122] This invention also discloses a computer-readable storage medium storing a computer program that causes a computer to execute the off-label drug use evidence assessment methods described in the above embodiments.

[0123] The purpose of the above embodiments is to reproduce and derive the technical solution of the present invention by way of example, and to fully describe the technical solution, purpose and effect of the present invention. The purpose is to enable the public to have a more thorough and comprehensive understanding of the disclosure of the present invention, and not to limit the scope of protection of the present invention.

[0124] The above embodiments are not an exhaustive list based on the present invention, and there may be many other embodiments not listed. Any substitutions and improvements made without departing from the concept of the present invention are within the protection scope of the present invention.

Claims

1. A method for evaluating evidence of off-label drug use, characterized in that, include: Determine the research type of the evidence and obtain the first cue word corresponding to that research type; Based on the first prompt word, a large language model is used to evaluate the methodological quality of the evidence and obtain the evaluation results. When the evaluation result is medium to high quality evidence, the evidence is graded based on the preset second prompt word using a large language model to obtain the grading result. The evaluation results are obtained based on the evaluation results and the grading results.

2. The method for evaluating evidence of off-label drug use as described in claim 1, characterized in that, The process of determining the research type of the evidence and obtaining the first cue word corresponding to the research type includes: Based on preset prompts, a large language model is used to determine the research type of the evidence; The first prompt word is determined based on the scale corresponding to the research type.

3. The method for evaluating evidence of off-label drug use as described in claim 1, characterized in that, The evidence is graded for validity using a large language model. After obtaining the grading results, the evidence is also graded using a large language model to obtain GRADE grading results. The GRADE grading results are then added to the grading results.

4. The method for evaluating evidence of off-label drug use as described in claim 3, characterized in that, Also includes: When the evaluation result is low-quality evidence, the manual evaluation result obtained by manually reviewing the evidence will be compared with the evaluation result. If the first preset condition is met, a re-evaluation will be conducted. And / or, The grading results are manually reviewed, and a reassessment is conducted when the review results meet the second preset condition.

5. The method for evaluating evidence of off-label drug use as described in claim 1, characterized in that, Based on the evaluation results and the grading results, an assessment result is obtained, including: A large language model is used to extract descriptions of security from the evidence to obtain security information; An assessment result is obtained based on the evaluation results, the classification results, and the security information.

6. The method for evaluating evidence of off-label drug use as described in claim 1, characterized in that, Both the first and second prompt words include a background section, a role section, a task section, and a quality control section. The background section is used to clarify the research type, the core purpose of the scale, and the research objective. The role section is used to guide the large language model to process the literature from the perspective of a professional evaluator. The task section is used to provide the large language model with an operation manual to standardize the evaluation process. The quality control section is used to set up a validation framework for the output of the large language model to ensure the rigor of the evaluation.

7. The method for evaluating evidence of off-label drug use as described in claim 1, characterized in that, The evaluation results include item scores, evidence summaries, and quality levels; the grading results include item scores, evidence summaries, and effectiveness grading.

8. A system for evaluating evidence of off-label drug use, characterized in that, include: The research type module is used to determine the research type of the evidence and obtain the first prompt word corresponding to the research type. The methodology quality assessment module is used to assess the methodology quality of the evidence based on the first prompt word using a large language model, and to obtain the assessment result. The evidence validity grading module is used to grade the evidence validity based on the second prompt word and a large language model when the evaluation result is medium to high quality evidence, and to obtain the grading result. The evaluation results module is used to obtain evaluation results based on the evaluation results and the grading results.

9. An electronic device, characterized in that, It includes a memory storing executable program code and a processor coupled to the memory; the processor calls the executable program code stored in the memory to execute the off-label drug use evidence assessment method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program causes a computer to perform the off-label drug use evidence assessment method according to any one of claims 1-7.