RCT literature classification model training method and system based on adversarial training
By constructing an RCT literature classification model based on adversarial training, using divergent literature to build an adversarial model and training a binary classification model, the problem of insufficient accuracy of RCT literature classification in the existing technology is solved, and higher classification accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202510925964.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-07-07
AI Technical Summary
The existing RCT literature classification model has problems such as data labeling defects, labeling noise, sample imbalance, and excessive recall and false positive rates, resulting in insufficient classification accuracy.
Build an RCT literature classification model based on adversarial training. By obtaining the designated medical literature data set with labels, building an adversarial model using divergent literature, generating a sample library, and training a binary classification model based on the sample library to ensure sample balance, and use semantic understanding model and fully connected prediction layer for feature extraction and classification.
It improves the accuracy of RCT literature classification, reduces the false positive rate and false negative rate, improves the recall rate and accuracy rate, and optimizes the sample utilization efficiency.
Smart Images

Figure CN120429646A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of document classification, and particularly relates to a method and system for training an RCT document classification model based on adversarial training. Background Art
[0002] With the development of medical concepts, the current medical model has shifted from past experience-based medicine to evidence-based medicine (EBM). Evidence-based medicine, which adheres to the principle that "all clinical decisions should be based on clinical evidence", can provide the most powerful evidence support and rigorous guidance for clinical research design in medical clinical work, and is of great guiding significance for clinical practice and research. The main evidence carrier of evidence-based medicine is systematic reviews, and their writing requirements are extremely strict. Researchers need to conduct a systematic search and literature screening for a specific clinical problem to find the current best clinical evidence, and then evaluate the risk of bias and integrate the results of these evidences. Its steps involve systematic retrieval, literature screening, information extraction, risk of bias assessment, and data synthesis, etc. To control the risk of bias in the included literature itself, the current best clinical evidence that systematic review writers need to find is generally the randomized controlled clinical trial (RCT) with the most rigorous research design. And to ensure the systematicness of the retrieval, reviewers often adopt a retrieval strategy with extremely high recall rate (sensitivity) but very low precision rate (accuracy), which makes the literature screening process often face thousands of bibliographies composed of title abstracts, and the vast majority of the bibliographies are not RCTs. This will cost researchers a lot of time and energy, and at the same time, the writing and updating speed of systematic reviews far cannot meet the needs of clinical decisions. A study shows that on average, 1,781 literatures need to be screened to evaluate a systematic review, and the average screening rate of irrelevant literatures reaches 97.1%, and it takes an average of 64.3 weeks to publish a systematic review.
[0003] Although the literature screening work itself has strict requirements, complex processes, and large workloads, the currently developed intelligent classifiers still have the following defects: 1. Data annotation defects; 2. Annotation noise: The Kappa coefficient of traditional manual annotation is only 0.65 (the ideal value > 0.9); 3. Sample imbalance: The diversity of non-RCT literatures is insufficient (the existing data set only covers 42 research types); 4. The recall rate still needs to be further improved and optimized; 5. Although the existing recall scheme for missed-screened RCT literatures can further reduce the number of missed-screened RCT literatures, the number of non-RCT literatures mislabeled as RCTs is too large (false positive rate is too high). Summary of the Invention
[0004] The embodiments of the present application provide a method and system for training an RCT literature classification model based on adversarial training, which can solve the problem of insufficient accuracy in RCT literature classification.
[0005] In a first aspect, the embodiments of the present application provide a method for training an RCT literature classification model based on adversarial training, including the following steps: Obtain a specified medical literature dataset with annotations; wherein, the annotations include RCT literature, non-RCT literature, and divergent literature; the divergent literature refers to literature for which multiple annotators have different judgments on whether the literature belongs to RCT literature. Construct an adversarial model based on the literature labeled as divergent literature in the dataset, and train the adversarial model based on the literature labeled as RCT literature and non-RCT literature in the dataset. Run the adversarial model to generate a sample library, and train a binary classification model based on the sample library; wherein, the binary classification model is a machine learning model for identifying whether a literature belongs to RCT literature, and the ratio of RCT literature and non-RCT literature in the sample library satisfies a preset sample balance condition.
[0006] In a possible implementation manner of the first aspect, the binary classification model includes a semantic understanding model and a fully connected prediction layer; wherein, the semantic understanding model is used to extract features; the fully connected prediction layer is used to predict the probability that a literature belongs to RCT literature based on the features extracted by the semantic understanding model, and output a binary classification result according to the probability.
[0007] In a possible implementation manner of the first aspect, the semantic understanding model includes: A first feature extraction layer for extracting the research framework features of the literature.
[0008] In a possible implementation manner of the first aspect, the semantic understanding model includes: A second feature extraction layer for extracting the hidden exclusion features of the literature.
[0009] In a possible implementation manner of the first aspect, the outputting the binary classification result according to the probability includes: Determine that the probability is greater than the threshold, then output the binary classification result that the literature belongs to RCT literature, otherwise output the binary classification result that the literature belongs to non-RCT literature. In a possible implementation manner of the first aspect, the threshold is a dynamic threshold adjusted according to the publication year of the literature.
[0010] The advantages of the embodiments of the present application compared with the prior art are: 1. Construct an adversarial model using the literature with disagreements in manual annotation, so as to obtain a model for sample augmentation with better performance, and be able to obtain more effective sample data based on this model to improve the training effect of the binary classification model.
[0011] 2. Augment samples using the adversarial model and selectively train the binary classification model using a balanced sample ratio, so as to obtain a binary classification model with higher classification accuracy (i.e., better accuracy, precision, and recall).
[0012] In a second aspect, an embodiment of the present application provides an RCT literature classification model training system based on adversarial training, including: A dataset module for obtaining a specified medical literature dataset with annotations; wherein, the annotations include RCT literature, non-RCT literature, and discrepant literature; the discrepant literature refers to the literature for which multiple annotators have different judgments on whether the literature belongs to RCT literature. An adversarial module for constructing an adversarial model based on the literature annotated as discrepant literature in the dataset, and training the adversarial model based on the literature annotated as RCT literature and non-RCT literature in the dataset. A training module for running the adversarial model, generating a sample library, and training a binary classification model based on the sample library; wherein, the binary classification model is a machine learning model for identifying whether a literature belongs to RCT literature, and the ratio of RCT literature and non-RCT literature in the sample library satisfies a preset sample balance condition.
[0013] In a third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for training an RCT literature classification model based on adversarial training according to any one of the above first aspects is implemented.
[0014] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the method for training an RCT literature classification model based on adversarial training according to any one of the above first aspects is implemented.
[0015] In a fifth aspect, an embodiment of the present application provides a computer program product, and when the computer program product runs on a terminal device, the terminal device is enabled to execute the method for training an RCT literature classification model based on adversarial training according to any one of the above first aspects.
[0016] It can be understood that the beneficial effects of the above second to fifth aspects can be referred to the relevant descriptions in the above first aspect, and will not be elaborated here. Description of the Drawings
[0017] Figure 1 It is a comparison chart of the literature classification results between the binary classification model Diting Model and the similar small model Robotsearch of this embodiment; Figure 2 It is a framework diagram of the RCT literature classification model training system based on adversarial training; Figure 3 It is a framework diagram of the terminal device; Figure 4 It is a specific flow chart of the actual operation of step 102. Detailed Implementation Manner
[0018] It should be understood that when used in the specification and appended claims of this application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0019] It should also be understood that the term "and / or" used in the specification and appended claims of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0020] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when", "once", "in response to determining" or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]" according to the context.
[0021] In addition, in the description of the specification and appended claims of this application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0022] Reference to "one embodiment" or "some embodiments" etc. described in the specification of this application means that a specific feature, structure or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear at different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0023] The training method and system of the RCT literature classification model based on adversarial training of the present invention will be described in more detail below with reference to the schematic diagrams, in which the preferred embodiments of the present invention are shown. It should be understood that those skilled in the art can modify the present invention described herein while still achieving the advantageous effects of the present invention.
[0024] An embodiment of this application provides a training method for an RCT literature classification model based on adversarial training, including the following steps: Step 102, obtain a specified medical literature dataset with annotations; wherein, the annotations include RCT literatures, non-RCT literatures, and divergent literatures; the divergent literatures refer to literatures for which multiple annotators have different judgments on whether the literature belongs to an RCT literature. Step 104, construct an adversarial model based on the literatures marked as divergent literatures in the dataset, and train the adversarial model based on the literatures marked as RCT literatures and non-RCT literatures in the data. Step 106, run the adversarial model to generate a sample library, and train a binary classification model based on the sample library; wherein, the binary classification model is a machine learning model for identifying whether a literature belongs to an RCT literature, and the ratio of RCT literatures and non-RCT literatures in the sample library satisfies a preset sample balance condition.
[0025] Regarding step 102: In an optional implementation manner, the medical literature dataset is a dataset obtained by manual annotation by at least 2 annotators.
[0026] In a preferred implementation manner, the medical literature dataset is obtained based on the annotation of experts in the relevant field. Traditional literature annotation work is usually responsible by specialized annotators, and the understanding of these annotators in a specific field is uneven, which may lead to a certain proportion of errors in the annotation results.
[0027] However, this annotation method is still widely used because the data requirements involved in the annotation work are usually large; while the number of experts in a specific field - such as medical experts in the medical field involved in this application - is limited and the output per unit time is lower, and the cost of using experts for annotation is too high. Once this cost consideration is superimposed on the problem of the aforementioned large data requirements, the cost problem will become extremely prominent.
[0028] At the same time, the construction of the adversarial model in the subsequent solutions of this embodiment depends on the conflict of the annotation results. This idea can effectively improve the performance of the adversarial model, provided that the quality of the annotation results is high enough. That is to say, the solution based on the idea of this embodiment requires annotators with as much experience as possible - ideally experts in the medical field.
[0029] The following will provide a specific process example of the actual operation of the applicant in step 102, as Figure 4 shown: Based on the associated research project currently being implemented by the applicant, 3 experts constructed an initial seed library based on the retraction literature library of the Retraction Watch Database, annotated a total of 12,472 literatures, and performed binary classification on whether the literature is a randomized controlled trial (RCT) through a consensus method, dividing it into a literature inclusion group (Include, that is, annotated as RCT literature), an exclusion group (Exclude, that is, annotated as non-RCT literature), and marking the MAYBE group with large differences (the views of 1 person are opposite to those of the other 2 people or cannot make a judgment). Finally, 1,330 inclusion group literatures and 11,142 exclusion literatures were marked.
[0030] In the above example, the annotation results with differences are still finally marked as RCT or non-RCT. In some other optional implementation manners, the subsequent annotation of the difference group may not be performed either.
[0031] In addition, the number of the expert group participating in the annotation work is preferably 3. This preferred solution can reduce the labor cost while ensuring the simplicity of the subsequent annotation of the literature in the difference group (that is, in the case where 2 people think it is the first result and 1 person thinks it is the second result, the common view of the 2 people can be referred to for annotation), so as to provide higher-quality samples for the subsequent training of the adversarial model.
[0032] Step 104 is used to expand the training samples of the binary classification model. The reason for using an adversarial network to expand the samples in this embodiment is that the characteristics of typical RCT literatures and typical non-RCT literatures are usually relatively clear. According to the experimental results of the applicant, typical RCT literatures and typical non-RCT literatures rarely become the sources of errors in the binary classification model. Therefore, the identification of the difference samples (MAYBE group) has become the key factor restricting the performance of the binary classification model.
[0033] The original intention of sample expansion is to overcome the problem of sample imbalance, that is, to ensure that the ratio of RCT documents and non-RCT documents in the sample library meets the sample balance condition. Since divergent samples play a more critical role in adjusting the performance of the binary classification model, step 104 of this embodiment also takes into account the performance adjustment effect of the expanded samples on the subsequent binary classification model - that is, trying to generate more expanded samples similar to the divergent samples.
[0034] This beneficial effect is achieved through the idea of building an adversarial model, that is, building an adversarial model based on divergent literature.
[0035] Divergent documents can be thought of as "blueprints of ambiguity." These documents are not simply noisy data, but rather collections of examples identified as ambiguous by human annotators. This ambiguity is not random, but rather stems from specific textual features. An adversarial model "built" on such documents can learn to replicate or target these specific ambiguity patterns, rather than simply applying random perturbations. This is a more sophisticated form of adversarial attack / generation. If the adversarial model is a generator, it can be trained to generate new examples that mimic this blueprint; if it is a model that learns ambiguity patterns, it can learn to identify the characteristics of this blueprint.
[0036] Non-traditional definition of adversarial model: The adversarial model of this invention is fundamentally different from traditional generative adversarial networks (GANs). Its purpose is not to generate fake samples to deceive the discriminator, but to extract key ambiguous features that cause annotators to disagree by deeply analyzing the divergent literature (such as semantic ambiguity in the description of the research method and incompleteness of the randomization logic). These features are then encoded into the core discriminant rules of the adversarial model. Targeted feature mining of MAYBE literature: 1. Leveraging NLP semantic parsing technology, we structured the MAYBE literature to identify frequently ambiguous sentence patterns (e.g., the non-randomness implied by "allocation according to patient preference") and contradictory text structures (e.g., the mixed prospective design and retrospective analysis terms in the abstract). 2. When designing the adversarial model's loss function, we introduced annotator disagreement metrics (e.g., Kappa coefficient variance) to force the model to focus on the characteristics of the MAYBE literature during training, enabling it to quantify the uncertainty areas at the classification boundaries.
[0037] The following provides an example of the specific process that applicants actually follow in Step 104: Due to the obvious text features of typical included and excluded group literature, it is difficult to improve the recall rate and discrimination accuracy of small models. Therefore, the MAYBE group, that is, the literature entries with semantic ambiguity leading to expert disagreement, is used to construct a text adversarial system. The core lies in leveraging the "ambiguity" and "deceptiveness" features inherent in these discrepant literatures to guide the adversarial network (especially the generator) to learn and generate higher-quality and more challenging adversarial samples. The results of high-confidence annotation are dichotomized into Include and Exclude to construct an initial literature seed bank for training; For the text adversarial generator in the adversarial model: For RCT literature: Randomly insert non-randomized description statements (such as "subjectively assigned by the researcher"); For non-RCT literature: Generate false positive paragraphs that conform to the CONSORT specification.
[0038] When constructing the adversarial model, by analyzing these discrepant literatures, it can learn: Which words, phrase combinations, or sentence structures are likely to cause ambiguity? Which ambiguities in the description of research methods will lead to the blurred boundary between RCT and non-RCT? How does non-RCT literature achieve the effect of being indistinguishable from the real by imitating the expression of RCT? For step 106, in a preferred embodiment, the sample balance condition is that the number of RCT literature and non-RCT literature is equivalent, and the "equivalent" here allows a certain proportion / number of errors.
[0039] According to the above embodiments, in another embodiment: The binary classification model includes a semantic understanding model and a fully connected prediction layer; among them, the semantic understanding model is used to extract features; the fully connected prediction layer is used to predict the probability that the literature belongs to RCT literature based on the features extracted by the semantic understanding model, and output a binary classification result according to the probability.
[0040] In a preferred embodiment, the semantic understanding model is a machine learning model based on a BERT-like model.
[0041] BERT (Bidirectional Encoder Representations from Transformers), as a pre-trained language model, its bidirectional attention mechanism can capture long-range context dependencies in text, which is particularly important for the deep semantic parsing of medical literature. For example, when judging whether a certain literature belongs to a randomized controlled trial (RCT), the model needs to simultaneously understand the relevance of multiple parts such as research design, methodology description, and result analysis. Traditional unidirectional models such as LSTM often have problems of information attenuation when dealing with such cross-paragraph semantic associations, while BERT can dynamically adjust the weights of word tokens at different positions through the multi-head attention layer in the self-attention mechanism.
[0042] In addition to BERT, there are also various alternative model architectures available for semantic understanding, each with its own unique advantages and limitations. For example, RoBERTa, as an improved version of BERT, outperforms the original BERT in many natural language understanding tasks through a dynamic masking mechanism and larger-scale training data. For the RCT literature classification task, RoBERTa may be better at capturing subtle semantic differences.
[0043] Another model worthy of attention is Longformer, which can handle long documents of up to 4096 tokens by introducing a combination of local attention mechanism and global attention mechanism, which has unique value for scenarios that require a complete analysis of the entire medical literature.
[0044] In terms of lightweight models, ALBERT significantly reduces the number of model parameters through parameter sharing and embedding decomposition techniques, and is suitable for deployment in resource-constrained environments. For example, in an online service that needs to process a large number of literature streams in real time, ALBERT can achieve fast inference while maintaining reasonable accuracy. However, its performance ceiling is usually lower than the full version of BERT, especially when dealing with complex semantic reasoning tasks.
[0045] The ELECTRA model uses a pre-training task of replacement detection. Compared with BERT's masked language modeling, its training efficiency is higher and it makes more efficient use of computing resources. This has practical value for scenarios that require frequent model updates (such as tracking the evolution of medical research methods), because the fine-tuning process for new data can be completed faster. However, ELECTRA may be slightly inferior to the traditional BERT architecture in capturing deep semantic relationships, which may have a slight impact on tasks that require a fine understanding of research design.
[0046] All in all, the core task of semantic understanding models is to extract features, and quite a number of natural language understanding models can complete this work. However, as mentioned above, BERT-like models usually have certain advantages when facing the tasks of this application.
[0047] Of course, some BERT - like models also have deficiencies. For example, the limitation of input tokens makes it impossible to fully understand the content of the literature, resulting in incomplete and comprehensive features being extracted.
[0048] To overcome these problems, an optional approach is to combine models. For example, use a BERT - like model to extract features of key paragraphs of the literature (such as titles, abstracts, etc.), use Longformer to extract features of the full text of the literature, and set the weights of the two.
[0049] Thus, this combined model can extract different features of the same literature.
[0050] Continuing this idea, this embodiment provides an optional implementation method as follows.
[0051] The semantic understanding model includes: The first feature extraction layer is used to extract the research framework features of the literature.
[0052] The second feature extraction layer is used to extract the hidden exclusion features of the literature.
[0053] Among them, the research framework feature refers to the logical chain of research ideas in the literature, such as the logical chain of "random assignment, double - blind, multi - center".
[0054] The hidden exclusion feature refers to some keywords in the text of some seemingly RCT literatures that can individually negate their RCT features. The existence of such keywords can, to a certain extent, identify the literature as a non - RCT literature, such as "retrospective cohort" mixed in the abstract.
[0055] This implementation method can theoretically more effectively complete the identification of RCT literatures. However, since the binary classification model has very poor interpretability for humans, it is difficult for us to find problems and make fine - tuning during the training process. Therefore, based on the above two feature extraction layers, we provide a preferred implementation method as follows: Construct the semantic understanding model as a two - branch model; Among them, the first branch is as described above, and is used to extract the research framework features and / or hidden exclusion features of the literature.
[0056] The second branch is used to output the explanatory text of the research framework features and / or hidden exclusion features.
[0057] That is to say, based on the output parameters of the original semantic understanding model, these parameters are "translated" into natural language, so that during the training process, researchers can count the "logical chain" in the research framework features and / or the "keywords" in the hidden exclusion features to estimate the training effect, and then fine-tune the training process to achieve better model performance.
[0058] According to the above embodiments, in another embodiment: The outputting of the binary classification result according to the probability includes: Determine that the probability is greater than the threshold, then output the binary classification result that the literature belongs to RCT literature, otherwise output the binary classification result that the literature belongs to non-RCT literature; Preferably, the threshold is a dynamic threshold adjusted according to the publication year of the literature.
[0059] It is particularly worth noting that the embodiments of the present application involve an adversarial model, a binary classification model, and the annotation / expansion of samples. The reason is that these embodiment solutions are used to illustrate the construction and training process of the binary classification model. When the training work of the binary classification model is completed, the model used to perform the RCT literature classification task is the trained binary classification model itself and does not involve others. That is to say: In the training stage: T1. It is necessary to obtain a labeled dataset (step 102), and the annotation work of these datasets is completed manually - especially by experts in related fields - which can effectively improve the reliability of the data (this constitutes a preferred pre-step of step 102), and then improve the training effect of the model to obtain a binary classification model with better performance (so as to have better performance when performing the RCT literature classification task).
[0060] T2. It is necessary to construct an adversarial model based on the discrepant literature labeled as MAYBE in T1. Here, the construction can also be understood as a kind of training for the adversarial model. After the adversarial model is trained, it is necessary to use the trained adversarial model to expand the samples.
[0061] T3. The foregoing samples are used to train the binary classification model. The sample library can include a part of the samples output by the trained adversarial model, a part of the samples in the dataset obtained in step 102 whose labels are not MAYBE, and a part of the samples that are not obtained in step 102 (for example, in step 102, a preferred solution is adopted to have professional personnel in the medical field annotate a higher-quality dataset to optimize the performance of the adversarial model. When the samples generated by the adversarial model are excellent enough, including some samples annotated by ordinary annotators in the training samples for the binary classification model can also achieve a good training effect for the binary classification model).
[0062] More specifically, in an optional embodiment, the samples in the sample library (i.e., the sample set for training the binary classification model) are all labeled (i.e., the true values indicating whether the sample belongs to RCT or non-RCT), so as to enable supervised training for the binary classification model.
[0063] In addition, considering cost issues: the marginal cost of samples generated based on the trained adversarial model is the lowest (i.e., after excluding the data annotation cost for training the adversarial model, the cost of samples generated by the adversarial model can be ignored), the marginal cost of samples labeled by ordinary annotators is medium (the annotation of each piece of data requires labor hours, but the unit price of labor hours is relatively low), and the marginal cost of samples labeled by experts in related fields (the preferred embodiment of step 102) is the highest. However, due to the training requirements of the adversarial model, this part of high-cost samples is essential.
[0064] Therefore, the preferred solution considering cost is to have experts label sufficient data for training the adversarial model. Among them, the MAYBE group is used for training the adversarial model, and the RCT and non-RCT groups are used for training the binary classification model.
[0065] Subsequently, since the number of data in the RCT and non-RCT groups in the expert-labeled data cannot be estimated and determined (even when considering the data labeled by ordinary annotators, the same problem exists), it may lead to the problem of unbalanced samples for training the binary classification model - for example, the ratio of positive samples (data labeled as RCT) and negative samples (data labeled as non-RCT) is extremely different; at the same time, when such labeled data is used as samples, there may be a problem that the binary classification features are too clear (refer to the previous description), which is not conducive to optimizing the training effect of the binary classification model.
[0066] Therefore, thanks to the fuzzy characteristics and almost zero marginal cost of samples generated by the adversarial model, the above two problems can be better solved - that is, by generating a sufficient number and balanced samples through the trained adversarial model to make the samples in the sample library meet the requirements (preset sample balance conditions).
[0067] T4. Based on the sample library of T3, train the binary classification model. After obtaining the trained binary classification model, the training phase is completed, and the RCT literature classification task can be carried out subsequently.
[0068] Different from the above training phase, the steps involved in the execution phase of the RCT literature classification task are quite simple. Specifically: E1. Obtain the literature to be classified.
[0069] Please note that the literature obtained in step E1 is the original literature without any annotation. Of course, the original literature can be digitalized through some preprocessing steps such as encoding, partial extraction, and standardization, but these preprocessing do not involve annotators or annotation labels and true values.
[0070] E2. Input the literature to be classified into the trained binary classification model.
[0071] E3. Obtain the literature annotation result output by the binary classification model, and the RCT literature classification task is completed.
[0072] Referring to the foregoing description, the following will provide relevant data based on the actual implementation process of the applicant to reflect the improvement effect of the present application.
[0073] The data of the relevant process is as follows.
[0074] Training data volume: 1000 initial seeds (i.e., the data volume annotated by experts in the field), which is expanded to 50,000 through adversarial training (the data volume of the sample library for training the binary classification model including the samples generated by the trained adversarial model).
[0075] After training the binary classification model based on these 50,000 data, the performance comparison between the trained binary classification model performing the RCT literature classification task and the traditional model performing the RCT literature classification task is as described in Table 1.
[0076] Table 1 Test set Traditional model (F1) This solution (F1) Retracted RCT 0.79 0.93 Children's RCT 0.68 0.89 Cohort 0.52 0.81 It can be seen that the model based on the above solution has been able to achieve better performance compared with the traditional model.
[0077] However, there is still room for optimization in the aforementioned cost problem.
[0078] Among the literatures annotated by experts, a considerable part of the literatures that do not belong to the MAYBE group "waste" the ability of experts in a sense (although these literatures that do not belong to the MAYBE group can be used as training samples for the binary classification model, the output of the adversarial model can be replaced at a lower cost). If there are fewer proportions of RCT literatures / non-RCT literatures with clear features and more proportions of MAYBE group literatures with fuzzy features in the data annotated by experts, and the professional knowledge of experts is more fully utilized to eliminate the RCT literatures / non-RCT literatures with clear features in these data, it will enable the expert group to complete the screening of more MAYBE group literatures at a shorter time cost.
[0079] Then, the optimization goal of the cost problem is transformed into how to make there be fewer proportions of RCT literatures / non-RCT literatures with clear features and more proportions of MAYBE group literatures with fuzzy features in the data set for expert annotation.
[0080] Considering the preferred embodiment, the (trained) binary classification model with better performance itself includes a semantic understanding module. That is to say, the above-mentioned solution can still achieve considerable performance improvement on the basis of at least partially relying on the model's understanding of semantics. This to some extent proves that a machine learning model with semantic understanding ability can provide positive value in the identification of RCT literature.
[0081] Therefore, the cost optimization problem is further specified as how to use a machine learning model to identify RCT literature at a low cost.
[0082] Naturally, some open-source and pre-trained large language models (such as DeepSeek) are considered for performing the above tasks (the pre-step of step 102, providing a dataset with a smaller proportion of RCT literature / non-RCT literature with clear features and a larger proportion of MAYBE group literature with fuzzy features for annotation by experts).
[0083] Logically, we continue this idea. Before the trained binary classification model performs the RCT literature identification task, a rough screening step based on an open-source and pre-trained large language model (such as DeepSeek) is added. Different from the final classification result, the preliminary classification result of the rough screening step is labeled as LLM-RCT and LLM-NonRCT.
[0084] Based on this optimized process, we provide Figure 1 the preferred embodiment shown below, where: For the part with the rough screening result of LLM-RCT, due to the above positive value, we believe that the probability of these literatures belonging to RCT is higher. Therefore, other solutions different from this application are used for secondary verification, which will not be elaborated here.
[0085] For the part with the rough screening result of LLM-NonRCT, the foregoing embodiments of this application can achieve quite ideal effects. Figure 1 The provided data can intuitively compare the performance of the DiTing Model (the code name of the binary classification model after training in this application) and Robotsearch (as a representative of traditional models). In Figure 1 the comparative example, the DeepSeek model is applied to 8325 literatures (the true values of these literatures are: 780 belong to RCT literatures, Figure 1 labeled as 780Include, 7545 belong to non-RCT literatures, Figure 1Among the initially screened ones marked as 7545 (Exclude), after the initial screening, there were 6,695 documents marked as LLM-NonRCT (the true values of these documents were: 16 belonged to RCT documents, and 6,979 belonged to non-RCT documents), and 1,330 marked as LLM-RCT (the true values of these documents were: 764 belonged to RCT documents, Figure 1 marked as 764 (Include), and 566 belonged to non-RCT documents, Figure 1 marked as 566 (Exclude).
[0086] Both models used the 6,695 documents with the initial screening result of LLM-NonRCT as input for comparative testing. As Figure 1 shown, from the test results, although the DiTing Model (the code name of the binary classification model after the training of this application) missed screening 8 Include documents, slightly higher than the 6 of Robotsearch (representing traditional models), the amount of documents mislabeled as included by it was much smaller than the latter (130 < 1,436).
[0087] The false negative rate of the system built by the Diting Model was 1.03%, and the false positive rate was 9.23%. The false negative rate of the system built by Robotsearch was 0.77%, and the false positive rate was 26.53%. Generally speaking, the performance had a significant improvement.
[0088] The embodiment of this application also provides an RCT literature classification model training system based on adversarial training. As Figure 2 shown, it includes: A dataset module 201 for obtaining a specified medical literature dataset with annotations; wherein, the annotations include RCT documents, non-RCT documents, and discrepant documents; the discrepant documents refer to documents for which multiple annotators have different judgments on whether the document belongs to an RCT document; An adversarial module 202 for building an adversarial model based on the documents marked as discrepant documents in the dataset, and training the adversarial model based on the documents marked as RCT documents and non-RCT documents in the data; A training module 203 for running the adversarial model, generating a sample library, and training a binary classification model based on the sample library; wherein, the binary classification model is a machine learning model for identifying whether a document belongs to an RCT document, and the ratio of RCT documents and non-RCT documents in the sample library meets a preset sample balance condition.
[0089] The embodiment of this application also provides a terminal device. As Figure 3As shown, the terminal device 30 includes: at least one processor 301, a memory 302, and a computer program 303 stored in the memory and executable on the at least one processor. When the processor executes the computer program, the steps in any of the above method embodiments are implemented.
[0090] The above are only the preferred embodiments of the present invention and do not impose any restrictive effects on the present invention. Any person skilled in the art within the technical field, without departing from the scope of the technical solution of the present invention, makes any form of equivalent replacement or modification and other changes to the disclosed technical solution and technical content of the present invention, which are all within the content of the technical solution of the present invention and still fall within the protection scope of the present invention.
Claims
1. A training method for RCT literature classification model based on adversarial training, characterized in that: The following steps are involved: Obtain a designated medical literature dataset with annotations; wherein the annotations include RCT literature, non-RCT literature, and divergent literature; the divergent literature refers to literature in which multiple annotators have different judgments on whether the literature is an RCT literature; Building an adversarial model based on documents marked as divergent documents in the data set, and training the adversarial model based on documents marked as RCT documents and non-RCT documents in the data set; The adversarial model is run to generate a sample library, and a binary classification model is trained based on the sample library; wherein the binary classification model is a machine learning model used to identify whether a document is an RCT document, and the ratio of RCT documents to non-RCT documents in the sample library meets a preset sample balance condition.
2. The RCT document classification model training method based on adversarial training according to claim 1 is characterized in that: The binary classification model includes a semantic understanding model and a fully connected prediction layer; wherein the semantic understanding model is used to extract features; the fully connected prediction layer is used to predict the probability that a document belongs to an RCT document based on the features extracted by the semantic understanding model, and output a binary classification result based on the probability.
3. The RCT document classification model training method based on adversarial training according to claim 2 is characterized in that: The semantic understanding model includes: The first feature extraction layer is used to extract the research framework features of the literature.
4. The RCT document classification model training method based on adversarial training according to claim 2 is characterized in that: The semantic understanding model includes: The second feature extraction layer is used to extract the hidden exclusion features of the document.
5. The RCT document classification model training method based on adversarial training according to claim 3 or 4, characterized in that: Outputting a binary classification result according to the probability includes: If it is determined that the probability is greater than the threshold, a binary classification result indicating that the document belongs to an RCT document is output; otherwise, a binary classification result indicating that the document belongs to a non-RCT document is output.
6. The RCT document classification model training method based on adversarial training according to claim 5 is characterized in that: The threshold is a dynamic threshold adjusted according to the year of publication of the literature.
7. A RCT literature classification model training system based on adversarial training, characterized by: include: A dataset module is used to obtain a designated medical literature dataset with annotations; wherein the annotations include RCT literature, non-RCT literature, and divergent literature; the divergent literature refers to literature where multiple annotators have different judgments on whether the literature is an RCT literature; an adversarial module, configured to construct an adversarial model based on documents marked as divergent documents in the dataset, and train the adversarial model based on documents marked as RCT documents and non-RCT documents in the dataset; A training module is used to run the adversarial model, generate a sample library, and train a binary classification model based on the sample library; wherein the binary classification model is a machine learning model used to identify whether a document is an RCT document, and the ratio of RCT documents to non-RCT documents in the sample library meets a preset sample balance condition.
8. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. A computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Construction method of Chinese RCT intelligent classifier based on machine learning
CN109753564A
Intelligent data extraction and quality evaluation method for evidence-based medicine RCT
CN111554405A
Classification model training method, sample classification method, sample classification device and equipment
CN112270379A
Multi-classification method and device based on binary classification model, electronic equipment and medium
CN112801222A
Random contrast test identification method for integrating multiple BERT models based on LightGBM
CN112836772A