Random control test literature screening method based on LLM large model
Through a hybrid model screening method based on large language models, combining inference and general LLM big models and expert manual judgment, the prompt word scheme is optimized, which solves the shortcomings in accuracy, efficiency and applicability of the existing RCT literature screening methods, and achieves higher efficiency and higher accuracy literature screening.
Patent Information
- Application Number
- CN202510274403.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
There is a lot of room for improvement in the accuracy, efficiency and applicability of existing randomized controlled trials (RCT) literature screening methods, especially due to the limited generalization ability of traditional machine learning models and the reliance on manual judgment on the annotation process, resulting in increased costs and difficulty.
The hybrid model screening method based on the large language model (LLM) is adopted, and through pre-labeling training and formal mixed labeling steps, the inference and general LLM big models are used, combined with expert manual judgment, and the initial prompt word scheme is optimized to improve screening efficiency and accuracy.
Achieve higher efficiency and higher accuracy RCT literature screening, reduce manual annotation costs, and improve the model's adaptability and generalization ability in different fields.
Smart Images

Figure CN120216618A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a randomized controlled trial literature screening method based on an LLM large model, and belongs to the technical field of RCT literature screening. Background Art
[0002] At present, the screening of randomized controlled trial (RCT) literature mainly relies on traditional methods, such as manual screening, rule-based screening tools, and machine learning models. Among them, although manual screening has high accuracy, it is time-consuming and labor-intensive. Rule-based methods usually rely on preset keywords or screening conditions, which cannot fully cover and screen out RCT literature. There are many cases of missed or false detections, and it is difficult to improve the accuracy. In recent years, machine learning models, especially text classification methods based on deep learning, have been applied to RCT literature screening. However, machine learning models usually require a large amount of high-quality manually annotated data for training in the early stage, and the annotation process still relies on manual judgment and annotation by experts, which is time-consuming and labor-intensive. In addition, the generalization ability of traditional machine learning models is limited. When facing RCT literature in different fields or different types, it is often necessary to retrain or fine-tune the model, which further increases the cost and difficulty of screening. Therefore, there is still much room for improvement in the accuracy, efficiency, and applicability of existing RCT literature screening methods.
[0003] The Large Language Model (LLM) has powerful natural language understanding capabilities, can efficiently process and screen a large number of RCT documents, and reduce manual annotation work. Through pre-training and few-sample learning, LLM can quickly adapt to RCT documents in different fields, improve screening efficiency and accuracy, optimize the screening process, and provide researchers with a more convenient document screening tool.
[0004] However, simply calling the LLM model directly to screen RCT literature cannot achieve the desired effect. This is because some general LLM models have fast answer speeds, but significant answer volatility, resulting in accuracy that needs to be improved. Some reasoning LLM models are conducive to solving complex logical problems, but the risk of fabricating data is significant. When generating complex reasoning content, the hallucination rate is high, and it is especially easy to fabricate false data or fictitious logical chains. Therefore, a randomized controlled trial literature screening method based on the LLM model was designed, using a mixed model screening method to provide a more efficient and accurate RCT literature screening method. Summary of the invention
[0005] The technical problem to be solved by the present invention is to provide a method for screening randomized controlled trial literature based on the LLM large model, which solves the problem of improving the efficiency and accuracy of screening randomized controlled trial literature through the LLM large model.
[0006] The technical problems to be solved by the present invention are achieved by the following technical solutions:
[0007] A method for screening randomized controlled trial literature based on a large LLM model, comprising the following steps:
[0008] S1. Convert the literature to be annotated into a unified format;
[0009] S2. Pre-annotation training:
[0010] S21. Randomly select a number of literature to be annotated as pre-annotation training literature;
[0011] S22. Manually determine and annotate whether the pre-annotation training literature is RCT literature by experts, and generate the first manual annotation result: RCT literature is annotated as RCT_0, and non-RCT literature is annotated as nonRCT_0;
[0012] S23. Design an initial prompt scheme promptA that can call the large LLM model to judge whether the literature to be annotated is RCT literature;
[0013] S24. Based on the initial prompt scheme promptA, call the inference large LLM model to automatically determine and annotate whether the pre-annotation training literature is RCT literature, and generate an inference annotation result: RCT literature is annotated as RCT_1, non-RCT literature is annotated as nonRCT_1, literature that cannot be determined is annotated as FALSE_1, and generate the chain of thought CoT for each result;
[0014] S25. Compare the inference annotation result with the first manual annotation result to obtain cases where the result is inconsistent with the expected result badcase, obtain the chain of thought CoT of the bad case, and calculate the correct rate of the judgment of the inference large LLM model;
[0015] S26. Analyze the reason for the inconsistency based on the chain of thought CoT of the bad case, so as to optimize and iterate the initial prompt scheme promptA until the correct rate of the inference annotation result reaches the preset value, and obtain the formal prompt scheme promptB;
[0016] S3. Formal mixed annotation:
[0017] S31. Call the general LLM based on the formal prompt scheme promptB to automatically determine and label whether the literature to be labeled is an RCT literature, and generate a general labeling result: label RCT literature as RCT_2, non-RCT literature as nonRCT_2, and literature that cannot be determined as FALSE_2. Merge the literature to be labeled with a general labeling result of RCT_2 into the final labeled RCT literature list, and cycle and determine the literature to be labeled with a general labeling result of FALSE_2 multiple times;
[0018] S32. Based on the professional RCT literature screening small model, perform a secondary automatic determination and labeling on whether the literature to be labeled with a general labeling result of nonRCT_2 is an RCT literature, and generate a professional labeling result: label RCT literature as Maybe, non-RCT literature as nonRCT_3, and merge the literature to be labeled with a professional labeling result of nonRCT_3 into the final labeled nonRCT literature list;
[0019] S33. Manually determine and label the literature to be labeled with a professional labeling result of Maybe by experts, and generate a second manual labeling result: label RCT literature as RCT_4, non-RCT literature as nonRCT_4, merge the literature to be labeled with a second manual labeling result of RCT_4 into the final labeled RCT literature list, and merge the literature to be labeled with a second manual labeling result of nonRCT_4 into the final labeled nonRCT literature list.
[0020] Preferably, the LLM includes an inference LLM and a general LLM. Among them, the inference LLM uses Deepseek-R1, and the general LLM uses Deepseek-V3.
[0021] Preferably, the professional RCT literature screening small model uses the Robotsearch RCT literature screening tool.
[0022] Preferably, for the literature to be labeled with an inference labeling result of FALSE_1 and a general labeling result of FALSE_2 in steps S24 and S31, if it still cannot be determined after 10 cycles of determination, it is labeled as the final FALSE_1 and FALSE_2. Among them, FALSE_1 is directly determined by experts in the pre-labeling training step and used to improve promptA.
[0023] Preferably, it further includes a secondary pre-annotation training step: supplement the to-be-annotated documents that are still FALSE after being repeatedly determined for multiple times from the general annotation results of FALSE_2 into step S21 as pre-annotation training documents, replace the initial prompt scheme promptA in step S23 with the formal prompt scheme promptB used in the formal mixed annotation in this step S3, and re-execute steps S22-S26 to obtain the formal prompt scheme promptC optimized by the secondary pre-annotation training.
[0024] Preferably, at least 200 to-be-annotated documents are randomly selected as pre-annotation training documents in step S21.
[0025] Preferably, in step S22, at least two experts are used to manually determine and annotate whether the pre-annotation training documents are RCT documents.
[0026] Preferably, the preset value in step S26 is not less than 95%.
[0027] The beneficial effects of the present invention are as follows:
[0028] (1) During pre-annotation training, relying on experts for a small amount of annotation can effectively reduce the hallucination rate and factual risk existing in the inference-type LLM large model, and improve the reliability of the inference-type LLM large model in the professional field; only a small amount of pre-annotation training documents are required, and the thought chain CoT demonstrated by the inference-type LLM large model can be used to assist experts in optimizing the prompt scheme, thereby reducing the labor and time costs of pre-annotation training.
[0029] (2) During formal mixed annotation, using the promptB scheme with a correct rate exceeding 97% obtained from pre-annotation training can overcome the problem of large fluctuations in the answers of the general-purpose LLM large model, and can quickly screen a large number of to-be-annotated documents. Compared with the manually designed promptA or the existing small model Robotsearch, the promptB scheme obtained through pre-annotation training can screen out RCT documents more comprehensively and accurately.
[0030] (3) During formal mixed annotation, among the non-RCT documents identified by the general-purpose LLM large model, a professional RCT document screening small model is used again for screening. The small model can pick out the RCT documents misclassified as non-RCT documents by the large model, further improving the accuracy of RCT document screening. For the documents determined by the small model to be RCT, they are re-annotated by experts, and the experts can correctly annotate the non-RCT documents misclassified as RCT documents by the small model, further improving the correct rate of comprehensive annotation.
[0031] (4) Capture the documents to be annotated that always show FALSE, conduct secondary pre-annotation training, and make full use of the Chain of Thought CoT of the inference-based LLM large model for abnormal data during the inference process, improving the efficiency of experts in optimizing the prompt scheme. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a schematic diagram of the three major steps of the present invention;
[0033] Figure 2 It is a schematic diagram of the detailed steps of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] In order to make it easy to understand the technical means, creative features, achieved purposes and effects of the present invention, the present invention will be further described below with reference to specific drawings.
[0035] The overall scheme design of the method for screening randomized controlled trial documents based on the LLM large model is an image of a docker container on the linux operating system, which is convenient for rapid deployment on the linux server side.
[0036] As Figure 1 , Figure 2 shown, a method for screening randomized controlled trial documents based on the LLM large model includes the following steps:
[0037] To run the following steps, a front-end web page is constructed, which includes the following content:
[0038] (1) Format conversion: The user imports their own document form (which must include two columns: title and abstract) into this web page, and the web page converts it into a.csv format file that can be recognized by Label Studio;
[0039] (2) Editing of the prompt: Used to update and adjust the prompt content;
[0040] (3) Different large language model API interface columns provide adjustable interfaces for LLMs including Deepseek, ChatGPT, Claude, Gemini, Grok, etc. In this embodiment, the content of this column is the python format call code for the official and third-party APIs of Deepseek, and it includes both R1 and V3 parts;
[0041] (4) The operation of the automatic annotation switch includes two parts: LLM and Robotsearch, and the start of automatic annotation of the LabelStudio form is achieved through a button.
[0042] S1. Convert the documents to be annotated into a unified format.
[0043] Retrieve and import literature data as the literature to be annotated, and convert the data format.
[0044] S2. Pre-annotation training:
[0045] S21. Randomly select no less than 200 pieces of literature to be annotated as pre-annotation training literature;
[0046] S22. Manually determine and annotate whether the pre-annotation training literature is RCT literature by experts, and generate the first manual annotation result: RCT literature is annotated as RCT_0, and non-RCT literature is annotated as nonRCT_0.
[0047] Specifically, experts enter the Label Studio page to perform data annotation on the corresponding form content, annotating 200 pieces of typical data. The form content of the literature to be annotated includes the title, abstract, and other content.
[0048] S23. Design an initial prompt scheme promptA that can call the LLM large model to determine whether the literature to be annotated is RCT literature.
[0049] Design the initial prompt scheme promptA through experts in this field.
[0050] In one embodiment, promptA includes the following content:
[0051] <Role>
[0052] You are a highly professional, capable, and specialized research assistant. You have read millions of papers and deeply understand the expressions in papers and literature.
[0053] < / Role>
[0054] <Purpose>
[0055] Our goal is to screen out papers that may use RCT (randomized controlled trials). More specifically, imagine that the RCT methodology we are looking for is formed by combining these words: (clinical trials, grouping the population to form a contrast relationship so that controlled things can be studied).
[0056] < / Purpose>
[0057] <Format>
[0058] Output strictly in the valid json format, otherwise it will cause a json loading error.
[0059] "{{n"
[0060] "\"is_rct\":\"RCT\" or \"Non-RCT\"\n"
[0061] "}} \n"
[0062] < / Format>
[0063] <General Guidelines>
[0064] -Make a judgment based on the provided text. -Read the full text carefully and consider the overall context and details in the text.
[0065] -Construct a scenario in your mind and pay attention to the criteria, cues, and warning signals that are directly passed.
[0066] -Judge whether the literature behind the provided text clearly uses RCT; if the possibility of RCT cannot be clearly excluded, further reading of the full text is required for confirmation; or determine that the literature is not an RCT paper.
[0067] -For papers that are clearly RCT or may be RCT, output "Yes".
[0068] -For papers that are not RCT, output "No".
[0069] < / General Guidelines>
[0070] <Criteria for direct pass: Very, very likely to be an RCT paper. If these are seen, it can be directly marked as "Yes">
[0071] -The phrase "safety and effectiveness" appears in the title
[0072] -Causal inference is used to draw conclusions because journals usually only publish such papers when accompanied by RCT
[0073] <Cues: RCT papers usually use the following words>
[0074] -Comparison, allocation, double-blind, random allocation, trial, effectiveness
[0075] <Warning signals: Non-RCT papers usually use the following words>
[0076] - "We searched", "questionnaire survey", "survey", "database", "test tube / in vitro experiment", "cross-sectional study", "retrospective study", "animal", "case-control"
[0077] -It is obvious from the overall context that this is not our target scenario: (clinical trial, grouping the population to form a contrast relationship so that controlled things can be studied)
[0078] <Rules that must be followed>
[0079] -Don't fabricate anything out of thin air! Make a judgment based on the provided text.
[0080] - Do not create any warning signals, tips, or pass-through criteria that are not listed
[0081] - Usually, clinical trial papers will mention a certain grouping, but do not explicitly mention "randomized" or "controlled". As long as there is a mention of a certain grouping, such papers are still worth reading the full text further to confirm the possibility of an RCT and mark it as "yes".
[0082] - The possibility of an RCT cannot be excluded just because of the lack of cue words
[0083] - When considering warning signals, tips, and pass-through criteria, the context must be combined. When they are contradictory, the strict priority order is: 1. Pass-through criteria, 2. Warning signals, 3. Tips
[0084] - Only output valid json data without adding any extra text
[0085] If the literature to be labeled is in English, all the prompt schemes in this article can be translated into English, and then the LLM large model can be called. If the literature is in other languages, it is translated into the corresponding language. Using the prompt in the same language as the literature will reduce possible translation errors and improve the accuracy
[0086] S24. Call Deepseek-R1 based on the initial prompt scheme promptA to automatically determine and label whether the pre-labeled training literature is an RCT literature, and generate an inferential labeling result: RCT literature is labeled as RCT_1, non-RCT literature is labeled as nonRCT_1, literature that cannot be determined is labeled as FALSE_1, and generate the chain of thought CoT for each result
[0087] Specifically, through the execution of the button on the above front-end web page, call the API of Deepseek-R1 to label 200 pieces of data by running a python script in the server background, and record the chain of thought CoT of Deepseek-R1 in the form. The value of the data label is set to three items
[0088] (1) RCT_1, literature determined by Deepseek-R1 to be an RCT
[0089] (2) nonRCT_1, literature determined by Deepseek-R1 to be a non-RCT
[0090] (3) FALSE_1, cannot be determined by the background set time threshold (reference value is set to 10 seconds), such as the API of Deepseek-R1
[0091] Parameter settings for calling Deepseek-R1
[0092] (1) Set "Temperature" to 0 (ensure that only "yes", "no", and "False" are returned, where "False" indicates that the LLM API does not respond).
[0093] (2) Set "Top-P" to 1, which allows the model to consider a wider range of feature combinations (such as identifying low-frequency but important descriptions like "blinded allocation"), improving the coverage of atypical description documents.
[0094] (3) Set "context" to 0 (ensure that the annotation result of the previous data does not affect the next one).
[0095] (4) Ensure that the prompt is the only variable and control randomness as much as possible.
[0096] By setting the above parameters, it is possible to reduce the volatility of the control answer.
[0097] For literature samples with an annotation value of "FALSE", the API must be called repeatedly multiple times to determine the result. If it cannot be determined (FALSE) after 10 consecutive determinations, then mark it as "FALSE_1" and break out of the loop. The reasons for the annotation value of "FALSE" include:
[0098] (1) The official or third-party API called is unstable and the speed is reduced.
[0099] (2) The literature being analyzed is difficult to judge, and the model keeps outputting thoughts and is in a state of entanglement.
[0100] (3) Other unknown reasons.
[0101] S25. Compare the inference-based annotation results with the first manual annotation results to obtain bad cases where the results are inconsistent with the expected results, obtain the thought chain CoT of the bad cases, and calculate the accuracy rate of the inference-based LLM large model's judgment.
[0102] Specifically, compare the results where the expert annotation and Deepseek-R1 annotation are inconsistent (for example, if the expert annotates an RCT literature and Deepseek-R1 annotates it as a non-RCT literature or "FALSE", both are considered inconsistent). Identify the inconsistent results as bad cases and adjust the initial prompt scheme promptA according to the CoT corresponding to the bad cases.
[0103] S26. Analyze the reasons for the inconsistency based on the thought chain CoT of the bad cases, and then optimize and iterate the initial prompt scheme promptA until the accuracy rate of the inference-based annotation results reaches 97%, obtaining the formal prompt scheme promptB.
[0104] Specifically, the prompt adjustment method is as follows (taking English literature as an example):
[0105] (1) Role identification: Require the LLM to act as an RCT literature screening expert and summarize the annotation experience of the expert during the adjustment process;
[0106] (2) Content structuring;
[0107] (3) Emphasize the use of causal reasoning to draw conclusions;
[0108] (4) Given keywords for judging RCT literature: random, control, group, "control group", comparison, assign, doubly, double blind, allocated, trial, efficency, efficacy, effect in the title, safety and efficacy, etc.;
[0109] (5) Given keywords for judging non - RCT literature: cell, mice, rat, mouse, etc.;
[0110] (6) Few - shot, give examples of typical RCT literature and non - RCT literature in the prompt part;
[0111] (7) Make special notes or adjust prompt elements according to bad cases;
[0112] (8) Emphasize that it cannot be determined as non - RCT just because the keywords for judging RCT do not appear in the title and abstract;
[0113] (9) Specific details and techniques for RCT, make changes and adjustments according to the misjudged CoT. Emphasize "possible RCT" in the prompt, that is, as long as a literature has the potential to be an RCT initially judged, it is determined as an RCT literature. For those with clinical or psychological, background (population) and / or background (general intervention comparison) or similar expressions in the title and abstract, they are regarded as literatures with the potential to be an RCT.
[0114] (10) When considering the dedicated direct pass / red flag / hint, etc. in the prompt, the context must be considered. When they are contradictory, they are strictly arranged in the order of priority: 1. direct pass, 2. red flag, 3. hint. If a word is both a direct pass and a hint, its priority is set as a direct pass.
[0115] (11) Unless otherwise specified, the subject in the provided text is defaulted to human.
[0116] The accuracy rate of the model annotation result refers to the comparison between the model annotation result and the expert annotation result. When they are consistent, it is correct. In this embodiment, based on the pre-annotation training accuracy rate reaching 97%, the server-side python script is run. For other medical literature data that needs to be classified and annotated, the API of Deepseek-V3 is called to perform automatic annotation of the data. On the basis of completing the data annotation using Deepseek-V3, the literature marked as nonRCT is extracted and for it, the small model kernel of Robotsearch deployed on the server side is called for secondary annotation, and the annotation values are set to: Maybe and nonRCT. The experts judge the literature marked as Maybe above. In this way, the false negative data (literature that is actually an RCT but is misjudged as nonRCT by the LLM) is reduced as much as possible. Through the above steps of operation, the RCT literature in the medical literature is effectively screened and its false exclusion probability is reduced.
[0117] S3. Formal mixed annotation:
[0118] S31. Call Deepseek-V3 based on the formal prompt scheme promptB to automatically determine and annotate whether the literature to be annotated is an RCT literature, and generate a general annotation result: RCT literature is annotated as RCT_2, non-RCT literature is annotated as nonRCT_2, and the literature that cannot be determined is annotated as FALSE_2. The literature to be annotated with the general annotation result of RCT_2 is merged into the final annotated RCT literature list, and the literature to be annotated with the general annotation result of FALSE_2 is judged repeatedly in a loop;
[0119] S32. Based on the Robotsearch small model, perform secondary automatic determination and annotation on whether the literature to be annotated with the general annotation result of nonRCT_2 is an RCT literature, and generate a professional annotation result: RCT literature is annotated as Maybe, non-RCT literature is annotated as nonRCT_3, and the literature to be annotated with the professional annotation result of nonRCT_3 is merged into the final annotated nonRCT literature list;
[0120] S33. Have the experts perform manual determination and annotation on the literature to be annotated with the professional annotation result of Maybe, and generate a second manual annotation result: RCT literature is annotated as RCT_4, non-RCT literature is annotated as nonRCT_4. The literature to be annotated with the second manual annotation result of RCT_4 is merged into the final annotated RCT literature list, and the literature to be annotated with the second manual annotation result of nonRCT_4 is merged into the final annotated nonRCT literature list.
[0121] RobotSearch is a machine learning-based literature screening tool that focuses on quickly identifying and screening randomized controlled trial (RCT) literature. Its core function is to achieve automated literature screening and data extraction in scenarios such as medical research and systematic reviews through a trained high-sensitivity model, which can reduce the manual screening volume by 10%-90% and significantly shorten the research cycle.
[0122] Its relevant technical details are as follows: relying on a machine learning model for RCT feature recognition, supporting dynamic adjustment of sensitivity and specificity parameters to achieve mode switching between high recall rate and high precision. Through intelligent analysis of titles / abstracts, standardized data extraction, and bias evaluation modules, literature classification and structured processing are completed.
[0123] In the preliminary test, the patent applicant compared the effects of large language models such as Deepseek and Robotsearch on screening and annotating prior literature data and found that the annotation results of the two models for the same dataset had different biases. Specifically, the literature misjudged by the large language model differed significantly from that of Robotsearch. Combining the annotation of RCT literature by the two models can further reduce the number of false negatives (misclassifying RCT literature as non-RCT literature).
[0124] The two have different focuses: the proportion of RCT literature mixed in the non-RCT literature identified by the general-purpose LLM large model is higher than that of the Robotsearch small model; while the proportion of non-RCT literature mixed in the RCT literature annotated by the Robotsearch small model is higher than that of the general-purpose LLM large model.
[0125] Therefore, during the formal mixed annotation, among the non-RCT literature identified by the general-purpose LLM large model, the professional small model for screening RCT literature is used again for screening. The small model can pick out the RCT literature that the large model misclassifies as non-RCT literature, further improving the accuracy of RCT literature screening. For the literature annotated as RCT by the small model, it is annotated again by experts. The experts can correctly annotate the non-RCT literature that the small model misclassifies as RCT literature, further improving the accuracy rate of the comprehensive annotation.
[0126] Using the above method, 500 prior RCT literature datasets and 500 prior non-RCT literature datasets are annotated.
[0127] (1) After importing 1,000 documents into the system's pre - front page, the format is first automatically changed to.csv, and then uploaded to Label Studio. At least two experts pre - annotate 200 RCT and non - RCT documents in the system. The experts described in this article are those with an extremely low error rate in this field, and the error rate is approximately equal to 0. Therefore, this article defaults that the accuracy rate of the documents pre - annotated by experts is 100%.
[0128] (2) Write the initial promptA, the content of which is mainly to judge whether a document belongs to an RCT or non - RCT document based on the title and abstract content of the document.
[0129] (3) Call the API of Deepseek - R1 to automatically annotate these pre - annotated documents according to the content of promptA. The annotation results are: RCT documents are annotated as RCT_1, non - RCT documents are annotated as nonRCT_1, documents that cannot be determined are annotated as FALSE_1, and the thought chain CoT for each result is generated.
[0130] (4) During the automatic annotation process, due to 1. reasons within the server itself being called, 2. the LLM being in a state of entanglement in judging the documents, resulting in no response for a relatively long time, and other unknown reasons, the returned result is FALSE at this time. It is necessary to repeatedly call Deepseek - R1 to judge the documents with the returned FALSE result until it is still FALSE after looping 10 times, and then annotate it as FALSE_1 and jump out of the loop.
[0131] (5) Compare the final results of the expert annotation and Deepseek - R1, extract the misjudged documents by Deepseek - R1 and the corresponding CoT, and optimize and improve the prompt content.
[0132] (6) According to the improved prompt, call the API of Deepseek - R1 to annotate 200 pre - annotated documents again, compare the results of the expert annotation and the LLM annotation, and repeatedly iterate and optimize the prompt according to the results until the false negative rate is reduced to less than 2% and the accuracy rate reaches 97%.
[0133] (7) According to the analysis of the thought chain CoT of the bad cases, it is found that the following restrictive indicators of promptA have been optimized:
[0134] <Direct - pass criteria: Very, very likely to be an RCT paper. When seeing these, it can be directly marked as "yes">
[0135] <Hint: RCT papers usually use the following words>
[0136] <Warning signal: Non - RCT papers usually use the following words>
[0137] <Rules to be followed>
[0138] Thus, the optimized promptB is generated as follows:
[0139] <Role>
[0140] You are a highly professional, capable and specialized research assistant. You have read millions of papers and deeply understand the expressions in papers and literature.
[0141] - If you feel confused or contradictory, it is very likely that the author of the paper title or abstract has made a mistake, which you have seen many times in low-quality papers. In this case, you should follow the direct-pass standard as much as possible to handle it.
[0142] < / Role>
[0143] <Purpose>
[0144] Our goal is to screen out RCT (randomized controlled trial) or potential RCT papers. More specifically, imagine the methodology we are looking for is formed by combining these words: (clinical or psychology, context involving population grouping and / or context comparing intervention measures, or expressions with the same meaning).
[0145] < / Purpose>
[0146] <Format>
[0147] Output strictly in the valid json format, otherwise it will cause json loading errors.
[0148] "{{n"
[0149] "\"is_rct\":\"RCT\" or \"Non - RCT\"\n"
[0150] "}}\n"
[0151] "Only output json data, no extra text, no md wrapping, and the output starts with a left curly brace."
[0152] < / Format>
[0153] <General guidelines>
[0154] - Please make a judgment based on the provided text.
[0155] - Read the full text carefully and consider the overall background and details in the text carefully. - Construct a scenario in your mind and pay attention to the direct - pass standard, tips and warning signals.
[0156] - Determine whether the literature behind the provided text clearly uses RCT; if the possibility of RCT cannot be clearly excluded, further read the full text for confirmation; or determine that the literature is not an RCT paper.
[0157] - For papers that are clearly RCT or may be RCT, output RCT.
[0158] - For non-RCT papers, output Non-RCT.
[0159] < / General Guidelines>
[0160] <Direct Passing Criteria: Very likely to be an RCT paper, can be directly marked as RCT when seeing these>
[0161] - "Safety and efficacy" appears in the title
[0162] - "Comparison" or expressions with the same meaning appear in the title or abstract
[0163] - "Control" or "control group" or expressions with the same meaning appear in the title or abstract
[0164] - "Group" or expressions with the same meaning appear in the title or abstract
[0165] - Causal inference is used to draw conclusions because journals usually only publish such papers when accompanied by RCT
[0166] <Hint: RCT papers usually use the following words>
[0167] - Allocation, double-blind, random allocation, trial, efficacy
[0168] <Warning Signals: Non-RCT papers usually use the following words>
[0169] - "We searched", "questionnaire survey", "survey", "database", "test tube / in vitro experiment", "cross-sectional study", "retrospective study", "animal", "case-control"
[0170] - It can be clearly seen from the overall background that this is not our target scenario: (clinical or psychological, background involving population grouping and / or background comparing intervention measures, or expressions with the same meaning)
[0171] <Rules to be Followed>
[0172] - Don't fabricate anything out of thin air! Make judgments based on the provided text.
[0173] - Don't create any guidelines, warning signals or hints that are not listed. These are based on existing standards and don't go beyond these.
[0174] - Generally, clinical or psychological papers will construct the background of a certain grouping, but do not explicitly mention "random" or "control". As long as the word "group" is mentioned, such papers are marked as RCTs.
[0175] - The possibility of an RCT cannot be excluded solely due to the lack of cue words.
[0176] - When considering the warning signals, cues, and direct-pass criteria, the context must be combined. When they conflict, the strict priority order is: 1. Direct-pass criteria; 2. Warning signals; 3. Cues. If a word is both a direct-pass criterion and a cue, it is preferentially regarded as a direct-pass criterion.
[0177] - The research subjects in the provided text are defaulted to humans unless otherwise specified.
[0178] - Only output valid json data, do not add any extra text, and start the output with a left curly brace.
[0179] (8) Use the trained promptB with 97% accuracy to call the Deepseek-V3 API to judge the types of RCT / nonRCT for 1000 documents. Repeatedly call the Deepseek-V3 API to judge the documents with a return value of FALSE among them, and loop 10 times. If it is still FALSE, mark it as FALSE and break out of the loop. Finally, mark the RCT documents as RCT_2, the non-RCT documents as nonRCT_2, and the undecidable documents as FALSE_2.
[0180] (9) For the documents marked as non-RCT, call the Robotsearch API for analysis and judgment. Marking results: Mark the RCT documents as Maybe, and the non-RCT documents as nonRCT_3. Merge the documents to be marked with a professional marking result of nonRCT_3 into the final marked non-RCT document list.
[0181] (10) For the results of Robotsearch marked as RCT, have experts evaluate and re-mark them (the default expert marking accuracy is 100%). Manually judge and mark the documents to be marked with a professional marking result of Maybe by experts, and generate the second manual marking results: Mark the RCT documents as RCT_4, and the non-RCT documents as nonRCT_4. Merge the documents to be marked with a second manual marking result of RCT_4 into the final marked RCT document list, and merge the documents to be marked with a second manual marking result of nonRCT_4 into the final marked non-RCT document list.
[0182] (11)After the above mixed annotation, the finally annotated RCT / nonRCT dataset is obtained, namely the RCT literature list and the nonRCT literature list.
[0183] According to the above promptB and 1000 randomly selected documents, including 500 prior RCT literature datasets and 500 prior nonRCT literature datasets, the following four schemes are verified:
[0184] Scheme 1: Deepseek-V3_promptB;
[0185] Scheme 2: Robotsearch;
[0186] Scheme 3: Deepseek-V3_promptB + Robotsearch (removing the expert annotation of this scheme);
[0187] This scheme: Deepseek-V3_promptB + Robotsearch + expert (complete version scheme).
[0188] The results of the four schemes for annotation are shown in the following table:
[0189]
[0190] 1. From the comparison between Scheme 1 and Scheme 2, it can be seen that: when judging RCT literature, the accuracy of Robotsearch in judging RCT literature is slightly higher than that of Deepseek-V3_promptB, and the false negative rate is low; when judging nonRCT literature, the accuracy of Deepseek-V3_promptB in judging nonRCT literature is higher than that of Robotsearch, and the false positive rate is low. This scheme is designed for these problems. The goal of this scheme is to minimize the false negative results and avoid missing as few RCT literatures as possible on the premise of controlling the excessive increase of false positives.
[0191] 2. From the comparison between Scheme 1, Scheme 2 and Scheme 3, it can be seen that: Scheme 3 without experts has better results than Deepseek-V3_promptB and Robotsearch alone when judging RCT literature. However, when judging nonRCT literature, there is an error of mixing some nonRCT literatures into RCT literatures.
[0192] 3. To address the problem of degradation when Solution 2's Robotsearch is used in combination with Deepseek-V3_promptB and non-RCT documents are being judged, experts are employed to conduct a secondary judgment on the RCT documents judged by Robotsearch, and it is assumed that the experts' judgment is 100% accurate. As a result, when Robotsearch is judging non-RCT documents, a portion of the non-RCT documents mixed in with the RCT documents will be picked out by the experts and classified as non-RCT documents.
[0193] 4. From the comparison of the annotation results of Solution 1, Solution 2, Solution 3, and this solution: Considering the final annotation accuracy rate of the above 1000 items, the accuracy rate of the full version of this solution is 98.0%, which is the best.
[0194] In some embodiments, for the documents to be annotated with an inference-based annotation result of FALSE_1 and a general-purpose annotation result of FALSE_2 in steps S24 and S31, if they still cannot be determined after 10 rounds of cyclic determination, they are annotated as the final FALSE_1 and FALSE_2. Among them, FALSE_1 is directly determined by the experts in the pre-annotation training step, and the promptA is improved based on the reasons for non-determination obtained by the experts.
[0195] In some embodiments, it further includes a secondary pre-annotation training step: The documents to be annotated that remain FALSE after multiple rounds of cyclic determination with a general-purpose annotation result of FALSE_2 are supplemented into step S21 as pre-annotation training documents. The formal prompt word scheme promptB used in the official mixed annotation in step S3 this time replaces the initial prompt word scheme promptA in step S23, and steps S22 - S26 are executed again. At this time, the preset value is greater than the preset value of the first pre-annotation training. For example, the first preset value is 97%, and the preset value this time is 98%. Until a certain amount of FALSE is accumulated and meets the standard to generate a new promptC. Specifically, if Deepseek-R1 also finally outputs FALSE, it is directly determined by the experts in the pre-annotation training step, and promptB is improved based on the reasons for non-determination obtained by the experts to generate a more optimized promptC.
[0196] Through secondary pre-training annotation, assuming the experts' annotation accuracy rate is 100%, the FALSE in the above table will be classified into the correct table, further improving the accuracy rate of subsequent versions.
[0197] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.
Claims
1. A method for screening randomized controlled trial literature based on the LLM model, characterized in that: The following steps are involved: S1. Convert the documents to be annotated into a unified format; S2. Pre-labeling training: S21. Randomly select a number of documents to be annotated as pre-annotated training documents; S22. Experts manually determine whether the pre-annotated training documents are RCT documents and generate a first manual annotation result: RCT documents are annotated as RCT_0, and non-RCT documents are annotated as nonRCT_0; S23. Design an initial prompt word scheme promptA that can call the LLM model to determine whether the document to be annotated is an RCT document; S24. Based on the initial prompt word scheme promptA, the inference-type LLM large model is called to automatically determine whether the pre-annotated training document is an RCT document, and generate inference-type annotation results: RCT documents are marked as RCT_1, non-RCT documents are marked as nonRCT_1, and documents that cannot be determined are marked as FALSE_1, and a chain of thought CoT for each result is generated; S25. Compare the inference-based annotation result with the first manual annotation result, obtain the bad case that is inconsistent with the expected result, obtain the thinking chain CoT of the bad case, and calculate the accuracy of the judgment of the inference-based LLM large model; S26. Based on the analysis of the thinking chain CoT of the bad case, the cause of inconsistency is obtained, and the initial prompt word solution promptA is optimized and iterated until the accuracy of the reasoning-type annotation result reaches the preset value, and the formal prompt word solution promptB is obtained; S3. Formal mixed annotation: S31. Based on the formal prompt word scheme promptB, the general LLM model is called to automatically determine whether the document to be annotated is an RCT document, and a general annotation result is generated: RCT documents are annotated as RCT_2, non-RCT documents are annotated as nonRCT_2, and documents that cannot be determined are annotated as FALSE_2. The documents to be annotated with the general annotation result of RCT_2 are merged into the final annotated RCT document list; S32. Based on the professional RCT literature screening model, the literature to be annotated with the general annotation result of nonRCT_2 is automatically judged and annotated for a second time to see whether it is an RCT literature, and the professional annotation results are generated: RCT literature is annotated as Maybe, and non-RCT literature is annotated as nonRCT_3. The literature to be annotated with the professional annotation result of nonRCT_3 is merged into the final annotated nonRCT literature list; S33. Experts manually judge and annotate the documents to be annotated whose professional annotation results are Maybe, and generate the second manual annotation results: RCT documents are annotated as RCT_4, and non-RCT documents are annotated as nonRCT_4. The documents to be annotated whose second manual annotation results are RCT_4 are merged into the final annotated RCT document list, and the documents to be annotated whose second manual annotation results are nonRCT_4 are merged into the final annotated nonRCT document list.
2. A method for screening randomized controlled trial literature based on the LLM large model according to claim 1, characterized in that: The LLM big model includes an inference-type LLM big model and a general-purpose LLM big model, wherein the inference-type LLM big model adopts Deepseek-R1, and the general-purpose LLM big model adopts Deepseek-V3.
3. The method for screening randomized controlled trial literature based on the LLM large model according to claim 1, characterized in that: The professional RCT literature screening model adopts the Robotsearch RCT literature screening tool.
4. The method for screening randomized controlled trial literature based on the LLM large model according to claim 1, characterized in that: If the document to be annotated whose inference-type annotation result is FALSE_1 and general-type annotation result is FALSE_2 in steps S24 and S31 cannot be determined after 10 cycles of determination, it will be marked as the final FALSE_1 and FALSE_2.
5. The method for screening randomized controlled trial literature based on the LLM large model according to claim 1, characterized in that: The method further includes a secondary pre-labeling training step: adding the undetermined documents whose general labeling result is FALSE_2 and whose to-be-labeled documents are still FALSE after repeated cyclical determination into step S21 as pre-labeling training documents, replacing the initial prompt word scheme promptA in step S23 with the formal prompt word scheme promptB used in the formal mixed labeling in step S3, and re-executing steps S22-S26 to obtain the formal prompt word scheme promptC after the secondary pre-labeling training optimization.
6. The method for screening randomized controlled trial literature based on the LLM model according to claim 1, characterized in that: In step S21, at least 200 documents to be annotated are randomly selected as pre-annotated training documents.
7. The method for screening randomized controlled trial literature based on the LLM large model according to claim 1, characterized in that: In step S22, at least two experts manually determine whether the pre-labeled training document is an RCT document.
8. The method for screening randomized controlled trial literature based on the LLM model according to claim 1, characterized in that: The preset value in step S26 is not less than 95%.
Citation Information
Cited By
RCT literature labeling method and system based on model consensus
CN120429756A