Medical first draft generation method and device based on large model and electronic equipment
Through the combination of large language models and medical bias assessment tools, a cross-base causal network is built, dynamic evidence link fusion and standardized conflict analysis are solved, and the problems of traditional meta analysis are time-consuming and manual deviation from the norm are achieved, achieving efficient, complete and credible generation of the first draft of medical treatment.
Patent Information
- Application Number
- CN202510838422.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-23
AI Technical Summary
The traditional meta analysis method takes a long time to generate the first draft of medical treatment, lacks a quantitative arbitration mechanism, resulting in low credibility in the conclusions and incomplete content, and manual operations are prone to deviation from specifications, resulting in omission of key data.
A large language model is used to generate multi-database search, cross-border searches and construct a causal network, and combined with medical bias evaluation tools to screen literature, dynamic evidence link fusion and heterogeneity evaluation, and generate standardized conflict analysis reports, and finally use the medical first draft template to generate the first draft.
Reducing manual intervention improves the integrity and credibility of the first draft of medical treatment, reducing time consumption and manual deviation through automated processing, ensuring the objectivity and standardization of the content.
Smart Images

Figure CN120353848A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technology, and more particularly, to a method, apparatus, and electronic device for generating a medical draft based on a large model. Background Art
[0002] Currently, the Meta (Meta-Analysis) analysis method is commonly used in the medical field to generate a systematic medical draft. A medical draft can only be obtained after covering complex processes such as literature retrieval, data extraction, statistical analysis, and structural interpretation. For obtaining a medical draft through the Meta method, the commonly adopted approach is that it completely relies on manual execution step by step to complete the medical draft.
[0003] However, when obtaining a medical draft in the above manner, the following technical problems often exist: First, in traditional Meta analysis, the literature screening process takes a long time. When the research conclusions included are contradictory, there is a lack of a quantitative arbitration mechanism. Manual reliance on subjective experience to judge conflicts results in low credibility of the conclusions of the medical draft.
[0004] Second, manual operators are prone to deviate from medical report norms, resulting in omission of key data extraction and lack of integrity of the content of the medical draft.
[0005] The above information disclosed in this background art section is only used to enhance the understanding of the background of the inventive concept, and thus, it may include information that does not form the prior art known to ordinary skilled persons in the art of this country. Summary of the Invention
[0006] This content section of the present disclosure is used to briefly introduce concepts that will be described in detail in the following detailed implementation section. This content section of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0007] Some embodiments of the present disclosure propose a method, apparatus, and electronic device for generating a medical draft based on a large model to solve one or more of the technical problems mentioned in the above background art section.
[0008] In a first aspect, some embodiments of the present disclosure provide a method for generating a medical draft based on a large model, including: using a large language model to generate a multi-database retrieval formula corresponding to a medical research direction; executing the multi-database retrieval formula in a database to perform cross-database retrieval to obtain a preliminary retrieved literature set and a cross-database causal network; screening and extracting the preliminary retrieved literature in the preliminary retrieved literature set according to the large language model and a medical bias assessment tool to obtain a draft data literature set and a clinical heterogeneity factor set; performing dynamic evidence chain fusion according to the draft data literature set and the clinical heterogeneity factor set to determine weights to obtain a fusion effect size set and a conflict marker set; performing heterogeneity judgment by executing a heterogeneity judgment script according to the fusion effect size set and the clinical heterogeneity factor set to obtain a visualization chart set; performing standardized conflict resolution according to the cross-database causal network and the conflict marker set to obtain a conflict resolution report; and using the large language model to determine the results of the draft data literature set and the visualization chart set according to a medical draft template and the conflict resolution report to generate a medical draft.
[0009] In a second aspect, some embodiments of the present disclosure provide a device for generating a medical draft based on a large model, including: a retrieval formula generation unit configured to use a large language model to generate a multi-database retrieval formula corresponding to a medical research direction; a literature retrieval unit configured to execute the multi-database retrieval formula in a database to perform cross-database retrieval to obtain a preliminary retrieved literature set and a cross-database causal network; a literature screening unit configured to screen and extract the preliminary retrieved literature in the preliminary retrieved literature set according to the large language model and a medical bias assessment tool to obtain a draft data literature set and a clinical heterogeneity factor set; a conflict marker unit configured to perform dynamic evidence chain fusion according to the draft data literature set and the clinical heterogeneity factor set to determine weights to obtain a fusion effect size set and a conflict marker set; a chart generation unit configured to perform heterogeneity judgment by executing a heterogeneity judgment script according to the fusion effect size set and the clinical heterogeneity factor set to obtain a visualization chart set; a report generation unit configured to perform standardized conflict resolution according to the cross-database causal network and the conflict marker set to obtain a conflict resolution report; and a draft generation unit configured to use the large language model to determine the results of the draft data literature set and the visualization chart set according to a medical draft template and the conflict resolution report to generate a medical draft.
[0010] In a third aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device storing one or more programs thereon, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method described in any implementation manner of the first aspect.
[0011] Fourthly, some embodiments of the present disclosure provide a computer-readable medium, on which a computer program is stored, wherein when the program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.
[0012] The above embodiments of the present disclosure have the following beneficial effects: The medical draft generated by the large model-based medical draft generation method of some embodiments of the present disclosure reduces manual intervention and improves the integrity of the content of the medical draft. Specifically, the reason for the incomplete content of the medical draft is that: the Meta-analysis method is generally used in the medical field to generate a systematic draft, and it completely depends on manual work to gradually execute the complete draft, which takes a long time. Due to manual intervention, it deviates from international standards, resulting in incomplete content of the generated medical draft. Based on this, for the large model-based medical draft generation method of some embodiments of the present disclosure, first, a multi-database retrieval formula corresponding to the medical research direction is generated using a large language model. Thereby, automatic adaptation to the syntax of different databases can be achieved, avoiding potential biases in manual strategy design. Secondly, the above multi-database retrieval formula is executed in the database for cross-database retrieval to obtain a preliminary retrieved literature set and a cross-database causal network. Thereby, the literature retrieval speed is improved through parallel cross-database retrieval, and at the same time, the causal association between literatures is quantified using the cross-database causal network, providing a basis for tracing the reasons for conflicting conclusions. Then, according to the above large language model and medical bias assessment tool, the preliminary retrieved literatures in the above preliminary retrieved literature set are screened and extracted to obtain a draft data literature set and a clinical heterogeneity factor set. Thereby, based on the screening and extraction process, an objective bias assessment result and standardized draft data can be obtained. After that, based on the above draft data literature set and the above clinical heterogeneity factor set, dynamic evidence chain fusion is performed to determine weights, obtaining a fusion effect size set and a conflict marker set. Thereby, conflict conclusions exceeding the threshold can be automatically identified as the basis for subsequent medical draft modification. Secondly, according to the above fusion effect size set and the above clinical heterogeneity factor set, a heterogeneity judgment script is executed for heterogeneity judgment to obtain a visualization chart set. Thereby, the error caused by manual selection of the effect model is reduced, and the standardization specification of chart generation is realized. Then, according to the above cross-database causal network and the above conflict marker set, a standardized conflict resolution is executed to obtain a conflict resolution report. Thereby, the conflict source is located through graph path analysis, and a structured evidence chain is generated. Finally, according to the medical draft template and the above conflict resolution report, the above large language model is used to determine the results of the above draft data literature set and the above visualization chart set to generate a medical draft. Thereby, the integrity of the chapter is restricted using the medical draft template. The conflicting conclusions are forcibly corrected through the conflict report. In summary, using a large model to generate a medical draft reduces the risk of missing content in the medical draft caused by manual intervention. By constructing a causal network, the conflicting conclusions in the medical draft are systematically analyzed, and the subjective experience judgment of manual work is transformed into structured objective analysis, improving the credibility of the medical draft. Brief Description of the Drawings
[0013] In conjunction with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the elements and elements are not necessarily drawn to scale.
[0014] Figure 1 is a flowchart of some embodiments of a large model-based medical draft generation method according to the present disclosure; Figure 2 is a schematic structural diagram of some embodiments of a large model-based medical draft generation device according to the present disclosure; Figure 3 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. Specific Embodiments
[0015] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for illustrative purposes and are not used to limit the protection scope of the present disclosure.
[0016] In addition, it should be noted that for the sake of convenience of description, only the parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.
[0017] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules, or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules, or units.
[0018] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly stated in the context, it should be understood as "one or more".
[0019] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0020] Embodiments of the present disclosure will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0021] Refer to Figure 1, shows a process 100 of some embodiments of the method for generating a medical draft based on a large model according to the present disclosure. The method for generating a medical draft based on a large model comprises the following steps: Step 101, using a large language model, generating a multi-database search formula corresponding to a medical research direction.
[0022] In some embodiments, the execution subject (e.g., electronic device) of the above-mentioned method for generating a medical draft based on a large model can be a medical Meta-analysis automated generation system deployed on a cloud computing platform, including the following core components: a large language model engine, a cross-library retrieval module, a causal network builder, a dynamic evidence fusion processor, and a conflict resolution engine. Among them, the above-mentioned large language model engine can be a model that uses a large language model (e.g., GPT-4 Turbo) as a basic model, superimposes a medical field adapter, and uses the Cochrane System Review Library and PubMed literature abstracts for fine-tuning and training. The above-mentioned cross-library retrieval module can be a module that initiates a query to a medical field database (e.g., PubMed and Embase) to retrieve and screen literature. The above-mentioned causal network builder can be a module that extracts corresponding elements (e.g., research subjects and intervention measures) from the included literature and constructs a causal logic chain between elements. The above-mentioned dynamic evidence fusion processor can be a module that automatically adjusts the weight value of each study through clinical heterogeneity factors (e.g., age deviations and intervention dose differences in the study population). The above-mentioned conflict resolution engine can be a module engine that locates the source of the conflict in the causal network and outputs an interpretability analysis of the cause of the conflict. The above-mentioned large model can be an artificial intelligence model with a large number of parameters and the ability to learn massive data, and can include a large language model and a multimodal large model. The above-mentioned medical research direction can be a systematic research topic that requires a medical draft to be generated. For example, "Analyze the correlation between the incidence of acute delirium after gynecological surgery and the anesthesia method and patient age." The above-mentioned multi-database search formula can be a standardized literature query command executed simultaneously in multiple academic databases. As an example, first, the medical knowledge encoding ability of the large language model is used to extract the elements in the medical research direction. Finally, the keywords in each element are connected using the database's "OR", and different elements are connected using the database's "AND" to obtain a multi-database search formula.
[0023] In some optional implementations of some embodiments, the execution subject may use a large language model to generate a multi-database search formula corresponding to the medical research direction, which may include the following steps: First step, extract keywords from the above-mentioned medical research directions to obtain research keywords. Among them, the above-mentioned research keywords can be the core terms and phrases that can represent the research subject extracted from the medical research directions. As an example, a word segmentation tool (for example, the jieba word segmentation tool.) can be used to perform word segmentation on the medical research directions to obtain research keywords.
[0024] Second step, utilize the above-mentioned large language model to perform the following generation steps: First sub-step, generate a single-database retrieval formula according to the above-mentioned research keywords. Among them, the above-mentioned single-database retrieval formula can be a structured query command designed for an academic database. As an example, use the large language model to perform semantic recognition on the research keywords to obtain research elements, and connect the elements according to the database query rules to generate a database retrieval formula.
[0025] Second sub-step, perform syntax conversion on the above-mentioned single-database retrieval formula according to the database conversion document to generate a multi-database retrieval formula. Among them, the above-mentioned database conversion document can be a mapping file storing the retrieval syntax rules of different academic databases. The above-mentioned multi-database retrieval formula can be a query statement corresponding to the structured query command form designed for multiple academic databases. As an example, perform extraction processing on the single-database retrieval formula to construct a syntax tree. Use the large language model to load the conversion rule library in the database conversion document and perform conversion on the syntax tree to obtain a multi-database retrieval formula.
[0026] Step 102, execute the multi-database retrieval formula in the database to perform cross-database retrieval, and obtain a preliminary retrieval literature set and a cross-database causal network.
[0027] In some embodiments, the above-mentioned execution entity can execute the multi-database retrieval formula in the database to perform cross-database retrieval, and obtain a preliminary retrieval literature set and a cross-database causal network. Among them, the above-mentioned preliminary retrieval literature set can be a literature collection obtained by performing preliminary retrieval from each database. The above-mentioned cross-database causal network can be a causal network diagram constructed by analyzing the associations between different literatures (for example, citation relationships and topic associations). Among them, the nodes represent the research elements of the literature, and the edges represent the causal relationships or strong associations between them.
[0028] As an example, execute the multi-database retrieval formula in the medical field database, and the obtained medical literature can be used as the preliminary retrieval literature set. By extracting the content information of the preliminary retrieved literature, the element information in the literature can be obtained. Compare the relationship of the element information between different literatures as the relationship between the element information and the literature. Use the element information as nodes and the relationship between the element information and the literature as edges. In this way, a graph network structure is constructed to obtain a cross-database causal network.
[0029] In some alternative implementations of some embodiments, the above-mentioned execution entity may execute the above-mentioned multi-database retrieval formula in a database to perform cross-database retrieval, and obtain a preliminary retrieval literature set and a cross-database causal network, which may include the following steps: First step, use an asynchronous function to execute the above-mentioned multi-database retrieval formula in parallel to obtain a retrieval literature set. Among them, the above-mentioned asynchronous function may be a non-blocking execution function, allowing the program to perform other tasks while waiting for I / O operations. The above-mentioned retrieval literature set may be a set of original literature results returned from each database.
[0030] As an example, first, initialize the database connector to establish a cross-platform connection. Then, set the asynchronous execution function of the asynchronous task scheduling mechanism to implement multi-data parallel retrieval. Finally, call the asynchronous execution function to execute the multi-database retrieval formula to retrieve, and obtain and integrate the literature of multiple data sources in real time to obtain the retrieval literature.
[0031] Second step, perform multi-dimensional standard screening on each retrieval literature in the above-mentioned retrieval literature set to generate preliminary retrieval literature and obtain a preliminary retrieval literature set. Among them, the above-mentioned multi-dimensional standard may be a standard for screening literature from multiple dimension directions. For example, in the scenario of generating a medical draft, the multi-dimensional standard may be "The full text of the literature is accessible and retrievable. The literature contains all the keywords extracted from the selected medical research direction. The literature does not belong to the scope of grey literature (such as dissertations and conference records). The research time (based on the literature collection time) is within three years." The above-mentioned preliminary retrieval literature may be the retrieval literature that meets the multi-dimensional standard. The above-mentioned preliminary retrieval literature set may be a set of relevant literature after screening.
[0032] As an example, first, extract the content corresponding to the retrieval standard fields in the retrieval literature according to the screening criteria. Finally, compare the field content with the standard requirements to screen the literature.
[0033] Third step, perform element extraction on each preliminary retrieval literature in the above-mentioned preliminary retrieval literature set to obtain a literature element information set, where the literature element information includes: the research object, intervention measures, result variables, and research design in the literature. Among them, the above-mentioned literature element information set may be a set of all literature element information. The above-mentioned research object in the literature may be the population characteristics involved in the research (such as age, gender, and disease status). The above-mentioned intervention measures may be the specific treatment methods and operations implemented on the research object in the research (such as drug intervention and surgical intervention). The above-mentioned result variables may be the main indicators measured in the research (such as survival rate and complication incidence rate). The above-mentioned research design may be the type of research method (such as cohort study: an observational research design used for analytical epidemiology).
[0034] As an example, first, extract the element information of the corresponding part in the literature according to the information fields of the retrieval elements. Finally, integrate the extracted element information to form a literature element information set.
[0035] In the fourth step, for each piece of literature element information in the above literature element information set, determine the corresponding relationship between the above literature element information and the above preliminary retrieval literature set to obtain a causal network correspondence set. Among them, the above causal network correspondence set can be a set of association relationships between literature elements and literatures.
[0036] As an example, create a mapping from literature elements to literatures, and use the mapping relationship information as the association relationship between literature elements and literatures to obtain a causal network correspondence set.
[0037] In the fifth step, generate the above cross-database causal network according to the literature element information and the corresponding retrieved literature in the above causal network correspondence set. Among them, the above cross-database causal network can be a graph structure representing the causal relationship between different literatures.
[0038] As an example, create a network diagram, use the literature element information as nodes and the relationship between the literature element information and the literature as edges, and add them to the network diagram to obtain a cross-database causal network.
[0039] Step 103, screen and extract the preliminary retrieved literatures in the preliminary retrieval literature set according to the large language model and the medical bias assessment tool to obtain a draft data literature set and a clinical heterogeneity factor set.
[0040] In some embodiments, the above execution subject can screen and extract the preliminary retrieved literatures in the preliminary retrieval literature set according to the large language model and the medical bias assessment tool to obtain a draft data literature set and a clinical heterogeneity factor set. Among them, the above medical bias assessment tool can be a medical research quality evaluation tool standardized for evaluating the bias risk involved in the research (for example, Cochrane ROB2). The above clinical heterogeneity factor set can be a set of factors of heterogeneity factors corresponding to the clinical research objects, intervention programs, and outcome indicators. It includes research object heterogeneity factors (population baseline characteristics: age, gender. Disease status: disease severity), intervention program heterogeneity factors (intervention measure types: drug types, surgical methods. Dosage and treatment course: drug dosage, treatment frequency), and outcome measurement heterogeneity factors (outcome indicator selection: difference between primary outcome and secondary outcome "Study A takes survival rate as the primary outcome, and Study B takes quality of life score as the primary outcome").
[0041] As an example, first, preliminary retrieved literature can be used to construct a medical knowledge graph, and relevant concepts can be extended based on medical research directions, and a semantic relationship network between concepts can be constructed. Second, the knowledge graph and medical bias assessment tools can be used to conduct bias assessment. And key clinical heterogeneity factors are identified to obtain a clinical heterogeneity factor set. Finally, the characteristics of the overall literature are extracted according to the knowledge graph to achieve the extraction of the literature content, and a draft data literature set is obtained.
[0042] In some optional implementation manners of some embodiments, the above-mentioned execution subject can screen and extract the preliminary retrieved literature in the above-mentioned preliminary retrieved literature set according to the above-mentioned large language model and medical bias assessment tool to obtain a draft data literature set and a clinical heterogeneity factor set, which may include the following steps: The first step is to obtain a pre-generated data extraction template. Among them, the data extraction template can be a structured table containing the standard fields required for Meta-analysis or a predefined framework in JSON format.
[0043] The second step is to use the above-mentioned medical bias assessment tool to conduct bias assessment on the above-mentioned preliminary retrieved literature to obtain a bias assessment result. Among them, the above-mentioned bias assessment result scores the literature quality in the specific risk areas of the literature (for example, randomization and data integrity).
[0044] As an example, according to the criteria of medical bias assessment, the medical bias assessment tool is used to score the literature in sequence to obtain the assessment result.
[0045] The third step is to, in response to the above-mentioned bias assessment result reaching a preset value, extract the content of the literature in the above-mentioned preliminary retrieved literature to obtain an extraction content set. Among them, the above-mentioned preset value can be a threshold for literature screening. The above-mentioned content extraction can be a process of obtaining structured data from the literature. The above-mentioned extraction content set can be a set of structured data extracted from the literature according to the template content. The above-mentioned content extraction method can be to extract the corresponding content in the literature according to the template fields.
[0046] The fourth step is to perform the following preprocessing operations on each extraction content in the above-mentioned extraction content set: The first sub-step is to use the above-mentioned large language model to extract the clinical heterogeneity factors in the above-mentioned extraction content to obtain at least one clinical heterogeneity factor. As an example, the large language model can be used to summarize the factors causing heterogeneity, read the literature content, and integrate the heterogeneity factors of the literature to obtain the clinical heterogeneity factor set of the literature.
[0047] The second sub-step is to fill the above-extracted content into the above data extraction template to obtain a draft data document. Among them, the above draft data document can be a standardized and analyzable structured data set that conforms to Meta-analysis. As an example, according to the field specifications defined in the data extraction template, find the corresponding field content in the extracted content and complete the filling of the template.
[0048] Step 104: Based on the draft data document set and the clinical heterogeneity factor set, perform dynamic evidence chain fusion to determine weights, obtaining a fused effect size set and a conflict marker set.
[0049] In some embodiments, the above execution entity can perform dynamic evidence chain fusion based on the draft data document set and the clinical heterogeneity factor set to determine weights, obtaining a fused effect size set and a conflict marker set. Among them, the above dynamic evidence chain fusion can be an evidence fusion method that dynamically adjusts weights according to the research characteristics and heterogeneity factors in the draft data document. The above fused effect size set can be a set of comprehensive effect indicators after integrating research characteristics. The above conflict marker set can be a structural identifier of a study with significant differences. The above weight can be a numerical representation that measures the contribution degree of the research characteristic to the comprehensive effect size according to the research characteristics and clinical heterogeneity factors, and is a quantitative influence on each literature. The above comprehensive effect indicator can be a quantitative indicator that characterizes the intervention measure after integrating the results of multiple literatures. The above research structure identifier can be a structured label used to mark significant differences or conflicts in the literature. For example, in the scenario of generating a medical draft, the research structure identifier can include the conflict type and the conflict location.
[0050] As an example, first, extract the research characteristics (such as age distribution, symptom ratio) from the draft data document set. And perform feature encoding and weight assignment on the clinical heterogeneity factor set. For example, age (weight 0.6), function (weight 0.4). Second, determine the heterogeneity score in the literature according to the research effect size. For example, the research characteristic can be "age > 75, hypertension". The clinical heterogeneity factor set can be {"age threshold": 75, "blood pressure": "hypertension"}. The heterogeneity score can be "age threshold × 0.6 + blood pressure × 0.4" = 1 × 0.6 + 1 × 0.4 = 1.0. Finally, determine the fused effect size according to the heterogeneity score. When the fused effect size is greater than a preset threshold (for example, 0.7). Then the corresponding clinical heterogeneity factor is used as a conflict marker. For example, the heterogeneity score can be 1.0 > 0.7. Then the conflict marker can be {"age threshold": 75, "blood pressure": "hypertension"}.
[0051] In some alternative implementation manners of some embodiments, the above-mentioned execution subject may perform dynamic evidence chain fusion based on the above-mentioned initial draft data literature set and the above-mentioned clinical heterogeneity factor set to determine weights, and obtain a fusion effect size set and a conflict mark set, which may include the following steps: First, extract the structured data corresponding to the research characteristics in the above-mentioned initial draft data literature set to obtain a research characteristic content set. Among them, the above-mentioned research characteristics may be the content of structured variables describing the core attributes of the research. For example, in the scenario of generating a medical initial draft, the description of the core attributes of the research includes: research design, participants, intervention measures, outcome measurement, and time factors. The above-mentioned research characteristic content set may be a set of specific research characteristics extracted from the initial draft data literature set. For example, in the scenario of generating a medical initial draft, the above-mentioned specific research characteristics may be structured variables that can standardize the description of the core attributes of a medical research. For example, the structured variables may include the objects, methods, outcomes, and time ranges of the research.
[0052] As an example, it may be to extract the corresponding content in the initial draft literature data through the mapping of research characteristic fields to obtain the research characteristic content.
[0053] Second, perform standardization processing on each clinical heterogeneity factor in the above-mentioned clinical heterogeneity factor set to obtain a heterogeneity factor data set. Among them, the above-mentioned standardization processing may be to convert the numerical values of different types of heterogeneity factors into a unified standard numerical representation. The above-mentioned heterogeneity factor data set may be a set of structured data after standardization processing.
[0054] As an example, in the scenario of generating a medical initial draft, corresponding processing methods are used for different types of clinical heterogeneity factors to obtain the data corresponding to the heterogeneity factors. For example, continuous factors can use numerical extraction methods. Categorical factors can use one-hot encoding methods. Textual factors can use vector representation methods. The data corresponding to the heterogeneity factors is normalized to obtain standardized heterogeneity factor data.
[0055] Third, use a preset heterogeneity weight function to negatively correlate the above-mentioned research characteristic content set and the above-mentioned heterogeneity factor data set to obtain a weight set corresponding to the research characteristic content set. Among them, the above-mentioned preset heterogeneity weight function may be a mathematical function used to quantify the degree of influence of research characteristics on heterogeneity. The above-mentioned negative correlation may be based on an inverse model of research characteristics and heterogeneity factors (for example, the more ideal the research characteristics, the lower the heterogeneity, and the higher the weight). The weight set corresponding to the above-mentioned research characteristic content set may be a set of relative importance scores of each research characteristic in the Meta-analysis.
[0056] As an example, in the scenario of generating a medical initial draft, the above-mentioned mathematical function may be , where can be the weight of the i-th research feature. can be the data of the i-th heterogeneity factor. can be the heterogeneity sensitivity coefficient in the medical field.
[0057] Step 4: Determine the research effect size according to the above medical research directions. Among them, the above research effect size can be a statistical index that quantifies the effect of the intervention measures in the medical research direction. For example, in the scenario of generating a medical draft, the statistical index can be odds ratio (death / survival), hazard ratio (cohort study), and standardized mean difference (continuous outcome: blood pressure value). In practice, the statistical index can be determined by analyzing the element information in the medical research direction.
[0058] Step 5: Perform dynamic weight weighted average on the above research effect sizes according to the above weight set to obtain the above set of integrated effect sizes. Among them, the above set of integrated effect sizes can be a set of structured data with statistical characteristics as fields and effect sizes as values.
[0059] Step 6: For each integrated effect size in the above set of integrated effect sizes, in response to the above integrated effect size reaching a preset threshold, determine that the research feature content corresponding to the above integrated effect size is a conflict marker. Among them, the above conflict marker can be the research feature corresponding to the integrated effect size exceeding the threshold, as the source of the conflict.
[0060] As an example, according to the value of the integrated effect size, determine the research features in the corresponding literature to obtain the conflict marker. For example, in the scenario of generating a medical draft, the set of integrated effects can be [RR = 0.72 (I 2 = 45%), RR = 1.15 (I 2 = 75%), RR = 0.68 (I 2 = 30%)]. The corresponding research features can be: Study 1: {Age: <70 years old, Dose: 15 mg, Design: RCT}, Study 2: {Age: >75 years old, Dose: 25 mg, Design: Cohort study}, Study 3: {Age: 70 - 75 years old, Dose: 18 mg, Design: RCT}. The preset threshold can be I 2 > 50%, then Study 2 triggers the conflict marker. The research features of Study 2 (Age: >75 years old, Dose: 25 mg, Design: Cohort study) are used as the conflict marker.
[0061] Step 105: According to the set of integrated effect sizes and the set of clinical heterogeneity factors, execute the heterogeneity judgment script to perform heterogeneity judgment and obtain the set of visualization charts.
[0062] In some embodiments, the above-mentioned execution entity may execute a heterogeneity evaluation script based on the set of fusion effect sizes and the set of clinical heterogeneity factors to perform heterogeneity judgment, and obtain a set of visualization charts. Among them, the above-mentioned heterogeneity evaluation script may be an algorithm program for automatically calculating heterogeneity indicators. The above-mentioned set of visualization charts may be the charts required in medical research. For example, in the scenario of generating a medical draft, the visualization charts may include: a forest plot showing the effect sizes of each study, a heterogeneity bar chart showing the sources of heterogeneity, and a funnel plot for detecting publication bias. As an example, the set of fusion effect sizes and the set of clinical heterogeneity factors may be used as data sources, and corresponding medical chart generation tools may be used to generate charts for the required content in the generation of a medical draft.
[0063] In some alternative implementation manners of some embodiments, the above-mentioned execution entity may execute a heterogeneity evaluation script based on the above-mentioned set of fusion effect sizes and the above-mentioned set of clinical heterogeneity factors to perform heterogeneity judgment, and obtain a set of visualization charts, which may include the following steps: First, determine the effect size of the quantitative research results corresponding to the above-mentioned set of fusion effect sizes to obtain an effect size data set. Among them, the effect size of the above-mentioned quantitative research results may be an effect size represented by a specific value (for example, risk ratio, standardized mean difference). The above-mentioned effect size data set may be a structured set of fusion effect size values in the form of an array or a list. As an example, the structured fields (quantitative research results) in the set of fusion effect sizes may be parsed to directly extract the numerical results of the effect sizes.
[0064] Second, determine the heterogeneity indicators of the clinical heterogeneity factors corresponding to the above-mentioned set of clinical heterogeneity factors to obtain a heterogeneity indicator data set. Among them, the above-mentioned heterogeneity indicators may be statistics for quantifying heterogeneity. The above-mentioned heterogeneity indicator data set may be a structured set with heterogeneity indicators as values.
[0065] For example, in the scenario of generating a medical draft, the clinical heterogeneity factors may be ["age difference", "dose difference", "study design difference"]. The heterogeneity indicators in the literature may be ["age difference": 0.58, "dose difference": 0.27, "study design difference": 0.15]. Then the heterogeneity indicator data set may be [0.58, 0.27, 0.15].
[0066] Step 3: Based on the above effect size dataset and the above heterogeneity index dataset, execute the above heterogeneity judgment script pre-generated according to corresponding requirements to obtain a statistic value and a heterogeneity index. Among them, the above heterogeneity judgment script can be a program for calculating heterogeneity statistics. In practice, according to the requirements of heterogeneity judgment, the code generation function in the large language model can be used to obtain the script program. The above statistic value can be a value describing the degree of variation of the effect sizes among studies. The above heterogeneity index can be a percentage value describing the proportion of the heterogeneous part in the total variation part.
[0067] Step 4: Based on the above heterogeneity index and the above effect size dataset, determine a heterogeneity effect model. Among them, the above heterogeneity effect model can be a statistical model selected according to the degree of heterogeneity.
[0068] For example, in the scenario of generating a medical draft, if the heterogeneity index (I 2 ) < 34%, the heterogeneity effect model can be a fixed effect model. If the heterogeneity index (I 2 ) >= 34%, the heterogeneity effect model can be a random effect model. Among them, the above fixed effect model can be a statistical model for Meta-analysis that assumes that the true effect sizes of all included studies are the same, and the differences between studies are only caused by random sampling errors. The above random effect model can be a statistical model for Meta-analysis that assumes that there are essential differences in the true effect sizes of each study, and affected by factors such as the study population, intervention methods, and environment, the effect sizes are normally distributed around the overall mean of the study.
[0069] Step 5: Based on the above heterogeneity effect model and the above effect size dataset, generate a heterogeneity forest plot and a heterogeneity bar plot. Among them, the above heterogeneity forest plot can be a chart showing the effect sizes and confidence intervals of each study. The above heterogeneity bar plot can be a chart showing the distribution of the sources of heterogeneity. In practice, the corresponding functions in the relevant image generation code library can be used to generate the charts.
[0070] Step 6: Based on the above statistic value and the above effect size dataset, generate a significance test chart. Among them, the above significance test chart can be a chart showing the statistical test results. For example, in the scenario of generating a medical draft, the significance test chart can be a heat map.
[0071] Step 7: Determine the above heterogeneity forest plot, the above heterogeneity bar plot, and the above significance test chart as the above visualization chart set.
[0072] Step 106: Based on the cross-database causal network and the conflict marker set, perform standardized conflict resolution to obtain a conflict resolution report.
[0073] In some embodiments, the above-mentioned execution entity may perform standardized conflict resolution based on the cross-database causal network and the conflict tag set to obtain a conflict resolution report. Among them, the above-mentioned standardized conflict resolution may be a process of uniformly processing conflicts according to predefined rules or algorithms to obtain a consistent conclusion. The above-mentioned conflict resolution report may be a report including a detailed description of the conflict, a resolution method, a resolution structure, and a correction suggestion.
[0074] As an example, first, according to the prior knowledge in the medical field, determine the rules for resolving conflicts (in the case of conflicts in studies of different qualities, select the results of high-quality studies. In the case of studies of the same quality, select the results with the majority agreement. In addition, select the conclusions of clinical guidelines). Secondly, analyze each conflict in the conflict tag set and apply the corresponding conflict resolution rule to obtain the analysis result. Finally, modify the cross-database causal network according to the analysis result, add the analyzed causal relationship, and generate a conflict resolution report. For example, in the scenario of generating a medical draft, the conflict resolution report may include: an overview of the number and type of conflicts, conflict elements, specific content of the conflicts, resolution methods, analysis results, and a summary of the updated causal network.
[0075] In the process of adopting technical solutions to solve the above-mentioned second technical problem, the following problems often occur. When multi-source medical literature gives contradictory conclusions on the same clinical problem, traditional manual methods are difficult to quantify the sources of conflicts and cannot extract key content, resulting in the omission of key evidence, leading to contradictory conclusions in the generation of medical drafts and the absence of some content.
[0076] To address these problems, the conventional solution is generally: conventional methods rely on medical guidelines for arbitration. However, the inventor considered that the differences in research methods of different research literatures lead to misjudgment of reasonable errors. We decided to adopt the following solution.
[0077] In some optional implementation manners of some embodiments, the above-mentioned execution entity may perform standardized conflict resolution based on the above-mentioned cross-database causal network and the above-mentioned conflict tag set to obtain a conflict resolution report, which may include the following steps: In the first step, extract the literature nodes and associated edges in the above-mentioned cross-database causal network to obtain graph structure data. Among them, the above-mentioned literature nodes may be network nodes in the cross-database causal network representing a piece of literature. The above-mentioned associated edges may be causal relationships (such as citation relationships and subject element associations) between literatures. The above-mentioned graph structure data may be a graph data structure represented by nodes and edges (such as an adjacency list and an adjacency matrix). As an example, corresponding extraction tools in a graph database (such as Neo4j) may be used to extract nodes and edges in the cross-database causal network.
[0078] In the second step, determine the research features corresponding to the above conflict label set to obtain the research feature content. Among them, the above research features can be the literature research features (research object, intervention measure) corresponding to the conflict labels. The above research feature content can be the specific feature content corresponding to the conflict labels in the literature.
[0079] As an example, based on the PIOCOS (Population, Intervention, Comparison, Outcome, Study design) elements, natural language processing technology can be used to extract the research features in the literature corresponding to the conflict labels.
[0080] In the third step, standardize the above graph structure data according to the preset rules to obtain causal network structured data. Among them, the above causal network structured data can be graph data that is convenient to process after standardization.
[0081] As an example, according to the rules in medical guidelines, unify the node names into common medical terms and the edges into predefined types (for example, "causes", "inhibits").
[0082] In the fourth step, construct hierarchical operation instructions according to the causal network structure of the above cross-database causal network and the above research feature content to obtain enhanced instructions. Among them, the above hierarchical operation instructions can be operation commands organized according to a certain processing order and hierarchical structure. For example, in the scenario of generating a medical draft, the hierarchical operation instructions can be processed from top to bottom. First, process the data, secondly, process the logic, and finally, generate the instructions. The above enhanced instructions can be operation methods for guiding feature enhancement and optimization.
[0083] For example, in the scenario of generating a medical draft, the research feature can be "intrathecal anesthesia". Then generate and analyze the subgraph related to "intrathecal anesthesia" and extract the content related to "intrathecal anesthesia".
[0084] In the fifth step, identify the data structure of the subgraph in the above causal network structured data according to the above preliminary retrieved literature set to obtain the key subgraph structure. Among them, the above subgraph can be a partial graph structure extracted from the causal network related to a specific relevant topic. For example, in the scenario of generating a medical draft, the above key subgraph structure can be a partial graph related to the content corresponding to the conflict label. The above specific relevant topic can be the research features corresponding to the conflict labels.
[0085] For example, in the scenario of generating a medical draft, according to the sequence type numbers of the preliminary retrieved literature, extract the subgraph containing these literature nodes from the causal network structured data.
[0086] Step 6: According to the preset conflict resolution rule library, use the above-mentioned reinforcement instructions to enhance the features of the above-mentioned key sub-graph structure to obtain causal network data. Among them, the above-mentioned preset conflict resolution rule library can be a pre-defined set of rules for resolving conflicts (selecting high-quality research results in studies of different qualities). The above-mentioned causal network data can be graph data obtained after adding new relationships (such as node weights, edge weights) to the sub-graph.
[0087] As an example, in the scenario of generating a medical draft, the rules in the preset conflict resolution rule library can be followed (for example, adding a literature quality score to the node and adding an association strength to the edge). At the same time, specific operations are performed according to the reinforcement instructions (such as merging nodes).
[0088] Step 7: In the above-mentioned causal network data, match the nodes and connection relationships corresponding to the above-mentioned research feature content to obtain adjacent literature nodes and adjacent literature edge weights. Among them, the above-mentioned adjacent literature nodes can be literature nodes directly connected to the literature involved in the conflict mark. The above-mentioned adjacent literature edge weights can be the weight values of the edges connecting the literature nodes.
[0089] As an example, first, in the causal network data, the corresponding nodes can be found according to the keywords of the research feature content. Finally, the weights of the first-order adjacent nodes and the connecting edges are obtained.
[0090] Step 8: According to the preset rules of graph path analysis, determine the relevance of the above-mentioned adjacent literature nodes and the above-mentioned adjacent literature edge weights in the above-mentioned causal network structured data to obtain the cause of the conflict. Among them, the above-mentioned preset rules of graph path analysis can be the method steps for analyzing the type of graph structure. The above-mentioned cause of the conflict can be the reason for the conflict in the draft conclusion. For example, in the scenario of generating a medical draft, the cause of the conflict can be the difference in the research population and the difference in the intervention measures. The above-mentioned preset rules of graph path analysis can be the shortest path.
[0091] As an example, in the scenario of generating a medical draft, graph algorithms (such as community detection) can be used to analyze the paths and weights between adjacent literature nodes to infer the source of the conflict (for example, two literatures with opposite conclusions may belong to different communities).
[0092] Step 9: According to the above-mentioned cause of the conflict, the above-mentioned research feature content, and the above-mentioned causal network data, construct a conflict evidence chain. Among them, the above-mentioned conflict evidence chain can be a chain of evidence that connects the cause of the conflict, relevant literature features, and network relationships.
[0093] As an example, in the scenario of generating a medical draft, the reasons for conflicts can be linked to specific research characteristics (e.g., research subjects, intervention doses) and the paths in the causal network data to form an evidence chain (e.g., the conclusions of Document A and Document B conflict because the ages of the research subjects are different. Document A targets <70 years old. Document B targets >75 years old).
[0094] In the tenth step, according to the preset template inference engine, the above conflict evidence chain is transformed into a structured conflict resolution report as the conflict resolution report. Among them, the above preset template inference engine can be a report template with a standard structure designed in advance. The above conflict resolution report can be a structured data result obtained by filling the conflict resolution structure into a fixed format.
[0095] As an example, according to the preset template, the content in the conflict evidence chain can be filled into the corresponding positions in the template to generate a report including various analysis contents of the conflict.
[0096] The above operation steps, as an inventive point of the present disclosure, solve Technical Problem 2 mentioned in the background art, that is, "manual operators are prone to deviate from medical report specifications, resulting in omission of key data extraction and lack of integrity of the content of the medical draft". In practice, conventional methods rely on arbitration by medical guidelines. The differences in research methods of different research literatures lead to misjudgment of reasonable errors, inability to extract key content, inability to connect the front and back logics, and lack of integrity of the content of the medical draft. The present disclosure designs a scheme for fusing and marking conflict conclusions with a dynamic evidence chain, fusing conflict conclusions through a dynamic evidence chain, and objectively attributing them through a cross-database causal network. Finally, a structured conflict resolution report is generated. As the basis for conflict conclusions, it solves the problem of lack of content in the draft caused by manual deviation from specifications, and ensures the integrity of the upper and lower evidence chains and content of the draft conclusions.
[0097] In the process of adopting the technical solution to solve the above Technical Problem 2, the following problems often accompany. In the process of adopting the technical solution to solve the above Technical Problem 2, the following problems often accompany. Conflict conclusions are scattered in different literatures, and it is difficult for humans to construct a complete evidence chain.
[0098] For these problems, the conventional solution is generally to use a large language model to analyze the connection between the literature and the research topic. However, the inventor considered using a large language model to analyze the correlation between the literature and the research direction. However, the inventor found that due to the wide range of model data sources, it is easy to deviate from the target literature set, resulting in insufficient relevance between the conflict conclusions and the research topic. We decided to adopt the following solution.
[0099] Optionally, the above-mentioned execution entity may construct hierarchical operation instructions based on the causal network structure of the above-mentioned cross-database causal network and the above-mentioned research feature content to obtain enhanced instructions, which may include the following steps: First, construct the subject-object relationship of the above-mentioned research feature content to obtain a set of feature triples. Among them, the above-mentioned subject-object relationship may be the semantic relationship between the subject (executing party) and the object (receiving party). The above-mentioned set of feature triples may be a structured relationship (subject-relationship-object) data set. For example, in the scenario of generating a medical draft, the above-mentioned semantic relationship may be "drug - affects - disease".
[0100] As an example, the research features in the research feature content can be parsed and constructed in the form of triples.
[0101] Second, determine the key nodes and edge weights of the above-mentioned causal network structure to obtain a set of quantization indicators. Among them, the above-mentioned set of quantization indicators may be a set of quantization evaluation results of nodes and edges. As an example, the weight value of the edge can be used as the basis for quantization evaluation. For example, in the scenario of generating a medical draft, the edge weights can be {("intrathecal anesthesia", "incidence of delirium"): 0.82, ("general anesthesia", "incidence of delirium"): 0.76}. Then the quantization indicators can be 0.82 and 0.76.
[0102] Third, for each feature triple in the above-mentioned set of feature triples, execute the following operation instruction generation steps: The first sub-step is to correspond the above-mentioned feature triple to the network nodes of the quantization indicators according to the above-mentioned set of quantization indicators to obtain feature network nodes and conflict quantization indicators. Among them, the above-mentioned feature network nodes may be the nodes corresponding to the triples in the causal network structure. The above-mentioned conflict quantization indicators may be conflict intensity indicators related to the network nodes.
[0103] As an example, first, the subject and object in the feature triple can be corresponded to the corresponding nodes of the causal network to obtain feature network nodes. Second, compare the relationship strength of the triple (for example, the initial relationship strength is a strong association "1") with the edge weight of the network to determine the degree of difference as the conflict quantization indicator.
[0104] The second sub-step is to extract a sub-graph from the above-mentioned causal network structure according to the above-mentioned feature network nodes to obtain a conflict sub-graph. Among them, the above-mentioned conflict sub-graph may be a local network structure including conflict nodes.
[0105] As an example, first, taking the feature network node as the center, directly connected upstream and downstream nodes can be extracted to obtain the first-level nodes. Secondly, the second-level nodes including those causally related to the feature network node are extracted. Finally, a local network structure including the feature network node is generated based on the first-level nodes and the second-level nodes.
[0106] In the third sub-step, according to the preset feature enhancement rule library, the above-mentioned feature network node, the above-mentioned conflict quantification index, and the above-mentioned conflict sub-graph are encapsulated with instructions to obtain enhanced instructions. Among them, the above-mentioned preset feature enhancement rule library can be predefined rules for handling conflicts. The above-mentioned instruction encapsulation can be a process of creating executable operation commands.
[0107] As an example, first, according to the preset rules for matching conflicts with types, the rule type and solution suggestions can be obtained. Secondly, the conflicts and conflict sub-graphs corresponding to the conflict quantification index are added to the instruction template. Finally, executable operation commands are generated. For example, in the scenario of generating a medical draft, the above-mentioned rule types may include: data types for supplementing specific types of research. Method types for adjusting statistical models. Clinical types for adding subgroup analysis.
[0108] The above operation steps, as an inventive point of the present disclosure, solve Technical Problem 2 mentioned in the background art and the problem that "conflict conclusions are scattered in different literatures and it is difficult for humans to construct a complete evidence chain". In practice, using a large language model to analyze the association between literature and research directions may easily deviate from the target literature set, resulting in insufficient relevance between conflict conclusions and research topics. The present disclosure designs a scheme for conflict quantification attribution based on a causal network. By constructing conflict triples, nodes and associated edges in the causal network are located. And sub-graphs are extracted according to the nodes and preset rules are matched to generate executable instructions and encapsulate them. Therefore, by determining the literature connection of conflict conclusions and executable instructions, the problems of fragmented conflict conclusion literature and broken evidence chain are solved.
[0109] Step 107, according to the medical draft template and the conflict resolution report, use a large language model to determine the results of the draft data literature set and the visualization chart set to generate a medical draft.
[0110] In some embodiments, the above-mentioned execution entity can use a large language model to determine the results of the preliminary draft data literature set and the visualization chart set based on the medical preliminary draft template and the conflict resolution report, and generate a medical preliminary draft. As an example, first, use the large language model to analyze the medical preliminary draft template and conflict resolution to obtain a structured medical preliminary draft template (e.g., imaging findings), and at the same time perform cross-modal alignment of the conflict resolution report with the structured medical preliminary draft template. Among them, the medical preliminary draft template includes semantic slots that can be filled (e.g., "method - sample size", "conclusion - contradiction point") to facilitate the filling of the preliminary draft conclusion content and conflict content. Finally, extract keywords (e.g., disease entities) and associated nodes (e.g., medical terms, experimental conclusions) from the preliminary draft data literature set, combine the keywords with the medical preliminary draft template, and use the retrieval enhancement module in the large language model to generate the output medical preliminary draft and automatically insert objective analysis.
[0111] In the process of adopting the technical solution to solve the above-mentioned technical problem 1, the following problems often occur. When using the retrieval enhancement module of the large language model to generate a medical preliminary draft, there are often biases. The modified content breaks the semantic coherence with the context, resulting in a decline in the overall quality level of the medical preliminary draft.
[0112] For these problems, the conventional solution is generally to add manual review and iteratively optimize the medical preliminary draft according to the prompt words. However, when the inventor considers using the prompt words to optimize the entire medical preliminary draft, the large language model may overfit the local standards of the manual prompt words and incorrectly correct the correct information that should be retained in the medical preliminary draft. This may further lead to the overall content loss of the medical preliminary draft and factual errors. We have decided to adopt the following solution.
[0113] Optionally, in some alternative implementation manners of some embodiments, after the above-mentioned execution entity uses the above-mentioned large language model to determine the results of the above-mentioned preliminary draft data literature set and the above-mentioned visualization chart set based on the above-mentioned medical preliminary draft template and the above-mentioned conflict resolution report and generates a medical preliminary draft, the above method may further include the following steps: First step, perform the following preliminary draft optimization steps for the above-mentioned medical preliminary draft: First sub-step, in response to the above-mentioned medical preliminary draft not having a modification suggestion problem feedback from the manual review terminal, perform conflict feature analysis on the above-mentioned conflict resolution report to obtain a conflict feature information set.
[0114] Among them, the above-mentioned conflict feature information set can be a set of structured problem features that can be processed by a machine extracted from the feature report.
[0115] As an example, the natural language description in the conflict report can be parsed to extract key features and stored in a structured form. For example, in the scenario of generating a medical draft, the conflict feature information can be {"conflict type": "data contradiction", "location": "RESULTS_SECTION_p3", "related entities": ["dose A", "1.2%"]}.
[0116] The second sub-step is to perform the following draft generation steps for the conflict feature information set: Sub-step one: Input the above conflict feature information set, the above medical draft, and the adjustment prompt words for the medical draft into a medical draft optimization model pre-trained based on medical knowledge to obtain a set of medical draft modification location information, the modified medical draft, and the adjusted length ratio.
[0117] Among them, the above adjustment prompt words can be instructions used to guide the model's modification direction. The above medical draft optimization model can be a dedicated model pre-trained based on medical knowledge. The above set of medical draft modification location information can be a set of location information of the content to be modified in the medical draft. The above modified medical draft can be the medical draft after being modified and optimized for the modified content under the guidance of the adjustment prompt words. The above adjusted length ratio can be the ratio of the modified content in the medical draft to the full text.
[0118] Sub-step two: According to the above set of medical draft modification location information, the modified medical draft, the above medical draft, and the local semantic adjustment prompt words, use the medical draft semantic verification model to generate the local semantic verification information corresponding to the above modified medical draft, where the above medical draft optimization model and the above medical draft semantic verification model are synchronously trained based on the training set of medical knowledge.
[0119] Among them, the above local semantic adjustment prompt words can be modification guidance instructions focusing on the context of the medical draft content corresponding to the modification location (for example, "check whether the pronoun reference at the modification place is clear"). The above local semantic verification information can be the semantic verification result of the context of the modified content of the medical draft. The above medical draft semantic verification model can extract the modified medical draft and the content at the corresponding local position of the medical draft for the medical modification location information, perform separate expansion of the content (for example, expand the context content), and then compare the expanded content of the two local manuscripts to determine whether the semantics are the same.
[0120] As an example, the context text of the medical draft modification location can be extracted and input into the medical draft semantic verification model to judge the state of the natural connection between the modified sentence and the context.
[0121] Sub-step 3: In response to determining that the above local semantic verification information indicates that the local semantic adjustment is correct and the above adjustment space ratio is higher than the target ratio, according to the above medical draft modification position information set, the modified medical draft, the above medical draft, and the overall semantic adjustment prompt word, use the medical draft semantic verification model to generate the overall semantic verification information corresponding to the above modified medical draft.
[0122] Among them, the above target ratio can be the threshold for triggering the overall modification of the medical draft. The above overall semantic adjustment prompt word can be an instruction for guiding the modification of the entire medical draft. The above overall semantic verification information can be the semantic verification result of the entire modified medical draft.
[0123] Sub-step 4: In response to determining that the above overall semantic verification information indicates that the overall semantic adjustment is correct or the above adjustment space ratio is not higher than the target ratio, generate the conflict resolution report of the above modified medical draft as the target conflict resolution report.
[0124] Among them, the above target conflict resolution report can be a conflict report generated to evaluate the modification effect of the modified medical draft, and detect whether there are conflict contents that have not been modified and optimized.
[0125] As an example, the modified conflict resolution report can be obtained by traversing and comparing the modified parts and the conflict feature information set in the modified medical draft.
[0126] Sub-step 5: In response to determining that the above target conflict resolution report does not indicate the existence of the target conflict feature information set, send the above modified medical draft to the target review terminal for manual review. Among them, the above target review terminal can be the terminal where the review interface used by medical experts is located.
[0127] Sub-step 6: In response to receiving the review correct information sent by the above target review terminal, determine the above modified medical draft as the medical draft. Among them, the above review correct information can be the review opinion given by the medical expert after the review (for example, text information, button confirmation signal).
[0128] Second sub-step: In response to determining that the above target conflict resolution report indicates the existence of the target conflict feature information set, generate the target conflict feature information set. Among them, the above target conflict feature information set can be the unresolved conflict features in the content of the modified medical draft.
[0129] As an example, by traversing the content of the target conflict feature information set, it can be judged whether there are unresolved conflict features. The third sub-step is to use the above-mentioned target conflict feature information set as the conflict feature information set, determine the revised medical draft as the medical draft, and continue to execute the above-mentioned draft generation step.
[0130] The second step is to, in response to receiving the above-mentioned review annotation information sent by the target review terminal, perform problem analysis on the above-mentioned review annotation information to obtain a problem feature information set. Among them, the above-mentioned review annotation information may be the modification opinions given by medical experts after reviewing the revised medical draft. The above-mentioned problem feature information set may be a set of structured problem feature information that can be processed by a machine after analyzing the review annotation information.
[0131] As an example, first, natural language processing technology can be used to analyze the review annotation information to obtain an analysis result. Finally, the analysis result is converted into a structured format of conflict features. For example, in the scenario of generating a medical draft, the review annotation information may be "Discuss the comparison of supplementary drug dose A". The problem feature information may be {"requirement type": "content expansion", "location": "DISCUSSION_SECTION"}.
[0132] The third step is to use the above-mentioned problem feature information set as the conflict feature information set, determine the revised medical draft as the medical draft, and continue to execute the above-mentioned draft generation step.
[0133] Optionally, the above-mentioned execution entity may input the above-mentioned conflict feature information set, the above-mentioned medical draft, and the adjustment prompt words for the medical draft into a medical draft optimization model pre-trained based on medical knowledge to obtain a medical draft modification location information set, a revised medical draft, and an adjusted length ratio.
[0134] Among them, the above-mentioned medical draft optimization model may be based on Transformer and a medical knowledge graph encoder, and includes: a multi-modal input layer, a medical knowledge enhancement layer, an instruction following module, and an output layer. Among them, the multi-modal input layer may be a network layer used to fuse text features (such as draft text, conflict features) and structured data (such as conflict locations, entity relationships). The above-mentioned medical knowledge enhancement layer may be a network layer that injects medical domain knowledge through a knowledge graph to solve the ambiguity problem of entities. The above-mentioned instruction following module may be a human feedback reinforcement learning optimization module that ensures the model can accurately understand the adjustment prompt words. The above-mentioned output layer may be a network layer that generates modification suggestions (such as location information, replacement text) and determines the proportion of the modified length.
[0135] As an example, the conflict feature information set and the medical draft are input into the multi-modal input layer. First, the multi-modal input layer can convert the conflict feature information into a structured conflict feature vector, including conflict type, research design, and effect size. The medical draft can be converted into a sequence of word vectors while preserving the semantic information of the original text. Second, the multi-modal input layer can determine the correlation between the text features in the word vector sequence and the conflict feature vector through the multi-head attention mechanism, and add absolute position encoding and relative position encoding. Finally, a fused feature vector that combines text features and conflict position information can be obtained. The fused feature vector is input into the medical knowledge enhancement layer. The medical knowledge enhancement layer can identify medical entities in the medical draft and the conflict feature information set through the knowledge graph, and make the identified content information correspond to resolve ambiguity. Knowledge-enhanced features are obtained (each knowledge-enhanced feature includes fused knowledge graph information features and medical entity association information). The knowledge-enhanced features and the adjustment prompt words for the medical draft are input into the instruction-following module. The instruction encoder can be used to convert the adjustment prompt words into task vectors, and the model parameters can be dynamically adjusted through human feedback reinforcement learning (e.g., RLHF) to optimize the prompt words. An optimized prompt-word-guided feature vector is obtained. The optimized prompt-word-guided feature vector is input into the output layer. The modified text range can be predicted through a classifier to obtain the modified position information. And the decoder is used to generate optimized content based on the context to obtain the modified medical draft. By determining the ratio of the original text length to the modified text length, the adjusted length ratio is obtained.
[0136] Optionally, the above-mentioned execution entity can use the medical draft semantic verification model to generate local semantic verification information corresponding to the above-mentioned modified medical draft based on the above-mentioned medical draft modification position information set, the modified medical draft, the above-mentioned medical draft, and the local semantic adjustment prompt words. Among them, the above-mentioned medical draft optimization model and the above-mentioned medical draft semantic verification model are synchronously trained based on a medical knowledge training set.
[0137] The above-mentioned medical draft semantic verification model can be a learning model based on double-encoder (e.g., Siamese Network twin neural network) comparison, including: a local encoder, a global encoder, a contrast learning module, and a verification rule engine. Among them, the above-mentioned local encoder can be used to extract the context features before and after the modified position. The above-mentioned global encoder can obtain the semantic features of the document through the BERT architecture. The above-mentioned contrast learning module can be used to determine the semantic similarity before and after modification (e.g., cosine similarity). The above-mentioned verification rule engine can be a built-in medical text specification and generate a structured verification report.
[0138] The above-mentioned shared dataset can be a document sourced from the full text of PubMed medical literature and revision records, Cochrane systematic review reports, expert annotations, clinical research protocols, and modification suggestions given by the committee.
[0139] As an example, preprocess the medical draft modification location information set, the modified medical draft, the medical draft, and the local semantic adjustment prompt words. Extract the original text local fragments and the modified text local fragments from the modified local fragments of the medical draft and the original content of the medical draft according to the medical draft modification location information. And convert them into text feature vectors. Also convert the local semantic adjustment prompt words into feature vectors and fuse them with the text feature vectors to obtain the local semantic prompt word feature vectors. Input the original text local fragments and the modified text local fragments into the local encoder. The local encoder can extract medical academic semantics through a pre-trained BERT model and capture the semantic dependencies between the front and back through the self-attention mechanism to obtain feature vectors, so as to obtain the original local feature vectors and the modified local feature vectors. Input the modified medical draft and the medical draft into the global encoder. The global encoder can encode the full text through a pre-trained BERT model to obtain document-level semantic identifiers, and identify the semantic associations between the modification locations and other chapters and obtain feature vectors, so as to obtain the original full text semantic vectors and the modified full text semantic vectors. Input the local semantic prompt word vectors, the original local feature vectors, the modified local feature vectors, the original full text semantic vectors, and the modified full text semantic vectors into the contrast learning module to determine the similarity of local semantic similarity and global semantic consistency (for example, cosine similarity). Obtain the comprehensive semantic similarity and the local feature similarity. Input the comprehensive semantic similarity, the local feature similarity, and the medical draft modification location information set into the verification rule engine. The verification rule engine can adaptively match rules according to the built-in medical text specifications. According to the verification of the comprehensive semantic similarity and the rule threshold, combined with the corresponding medical draft modification location information, generate structured information including verification conclusions, deviation reasons, and modification suggestions to obtain local semantic verification information.
[0140] The above operation steps, as an inventive point of the present disclosure, solve the first technical problem mentioned in the background art, that is, "when the research conclusions incorporated are contradictory, there is a lack of a quantitative arbitration mechanism, and manual judgment relies on subjective experience to determine conflicts, resulting in a low credibility of the conclusions in the medical first draft." In practice, the retrieval enhancement module of the large language model is used to generate the medical first draft, which may lead to a break in the semantic coherence between the modified content and the context, resulting in a decline in the overall quality level of the medical first draft. The present disclosure designs a three-level verification mechanism and a dynamic feedback optimization first draft solution. By structuring the conflict resolution report, the position of the modified content is obtained. Modify the modified content and verify and judge the semantics of the context, and obtain the modification ratio in the full text. According to the preset full text ratio threshold, perform semantic verification on the full text of the modified medical first draft to judge the structural coherence. And by submitting the first draft that passes the semantic verification to manual review, and continuously iterating and optimizing according to the review opinions and conflict reports, a standard medical first draft is obtained. Therefore, in order to reduce the overfitting of the large model to the manual prompt words and avoid mistakenly deleting the correct information in the medical first draft, the three-level verification mechanism and dynamic feedback are used to optimize the first draft. By verifying the context coherence of the modified content through semantic analysis, triggering the semantic logic verification of the full text based on the proportion threshold of the modified content, and submitting the medical first draft that passes the semantic logic verification of the full text to manual for the final review to iteratively optimize the medical first draft and obtain a medical first draft document with complete content.
[0141] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: The medical draft generated by the large model-based medical draft generation method according to some embodiments of the present disclosure reduces manual intervention and improves the integrity of the content of the medical draft. Specifically, the reason for the incomplete content of the medical draft is that: the Meta-analysis method is generally used in the medical field to generate a systematic draft, and it completely depends on manual work to gradually execute the complete draft, which takes a long time. Due to manual intervention, it deviates from international standards, resulting in incomplete content of the generated medical draft. Based on this, in some embodiments of the present disclosure, the large model-based medical draft generation method first uses a large language model to generate a multi-database retrieval formula corresponding to the medical research direction. Thereby, automatic adaptation to the syntax of different databases can be achieved, and potential biases caused by manual strategy design can be avoided. Secondly, execute the above multi-database retrieval formula in the database for cross-database retrieval to obtain a preliminary retrieval literature set and a cross-database causal network. Thereby, the retrieval speed of the literature is improved through parallel cross-database retrieval, and the causal association between the literatures is quantified by using the cross-database causal network, providing a basis for tracing the reasons for the conflict of conclusions. Then, according to the above large language model and medical bias assessment tool, screen and extract the preliminary retrieval literatures in the above preliminary retrieval literature set to obtain a draft data literature set and a clinical heterogeneity factor set. Thereby, based on the screening and extraction process, an objective bias assessment result and standardized draft data can be obtained. After that, according to the above draft data literature set and the above clinical heterogeneity factor set, perform dynamic evidence chain fusion to determine weights to obtain a fusion effect size set and a conflict marker set. Thereby, conflict conclusions exceeding the threshold can be automatically identified as the basis for subsequent medical draft modification. Secondly, according to the above fusion effect size set and the above clinical heterogeneity factor set, execute a heterogeneity judgment script to perform heterogeneity judgment to obtain a visualization chart set. Thereby, the error caused by the artificial selection of the effect model is reduced, and the standardization specification of chart generation is realized. Then, according to the above cross-database causal network and the above conflict marker set, execute standardized conflict resolution to obtain a conflict resolution report. Thereby, the conflict source is located through graph path analysis, and a structured evidence chain is generated. Finally, according to the medical draft template and the above conflict resolution report, use the above large language model to determine the results of the above draft data literature set and the above visualization chart set to generate a medical draft. Thereby, the integrity of the chapter is restricted by using the medical draft template. The conflicting conclusions are forced to be corrected through the conflict report. In summary, using a large model to generate a medical draft reduces the risk of missing content in the medical draft caused by manual intervention. By constructing a causal network, systematically analyzing the conflict conclusions in the medical draft, and transforming the subjective experience judgment of manual work into a structured objective analysis, the credibility of the medical draft is improved.
[0142] For further reference Figure 2, as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a large model-based medical draft generation device, and these device embodiments correspond to Figure 1 the method embodiments shown. The large model-based medical draft generation device can be specifically applied to various electronic devices.
[0143] As Figure 2 shown, a large model-based medical draft generation device 200 includes: a retrieval-based generation unit 201, a literature retrieval unit 202, a literature screening unit 203, a conflict marking unit 204, a chart generation unit 205, a report generation unit 206, and a draft generation unit 207. Among them, the retrieval-based generation unit 201 is configured to: use a large language model to generate a multi-database retrieval formula corresponding to a medical research direction. The literature retrieval unit 202 is configured to: execute the above multi-database retrieval formula in a database to perform cross-database retrieval, and obtain a preliminary retrieved literature set and a cross-database causal network. The literature screening unit 203 is configured to: screen and extract the preliminary retrieved literatures in the above preliminary retrieved literature set according to the above large language model and a medical bias evaluation tool, and obtain a draft data literature set and a clinical heterogeneity factor set. The conflict marking unit 204 is configured to: perform dynamic evidence chain fusion according to the above draft data literature set and the above clinical heterogeneity factor set to determine weights, and obtain a fusion effect size set and a conflict marking set. The chart generation unit 205 is configured to: execute a heterogeneity judgment script according to the above fusion effect size set and the above clinical heterogeneity factor set to perform heterogeneity judgment, and obtain a visualization chart set. The report generation unit 206 is configured to: perform standardized conflict resolution according to the above cross-database causal network and the above conflict marking set, and obtain a conflict resolution report. The draft generation unit 207 is configured to: determine results for the above draft data literature set and the above visualization chart set according to a medical draft template and the above conflict resolution report, and use the above large language model to generate a medical draft.
[0144] It can be understood that the units described in the large model-based medical draft generation device 200 correspond to the respective steps in the method described with reference to Figure 1 the above. Thus, the operations, features, and beneficial effects described above for the method also apply to the large model-based medical draft generation device 200 and the units included therein, and will not be elaborated herein.
[0145] Next, with reference to Figure 3 , which shows a schematic structural diagram of an electronic device (e.g., an electronic device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not impose any limitation on the functions and usage scopes of the embodiments of the present disclosure.
[0146] As shown Figure 3 in FIG. 335, the electronic device 300 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 301, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0147] Generally, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 3 FIG. 338 shows the electronic device 300 having various devices, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had. Figure 3 Each block shown in FIG. 339 may represent a device or, as needed, multiple devices.
[0148] Specifically, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such some embodiments, the computer program may be downloaded and installed from a network through the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above functions defined in the methods of some embodiments of the present disclosure are performed.
[0149] It should be noted that in some embodiments of the present disclosure, the above-mentioned computer-readable medium may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0150] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0151] The above computer-readable medium may be included in the above electronic device; or it may exist independently without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to: generate a multi-database retrieval formula corresponding to the medical research direction by using a large language model; execute the multi-database retrieval formula in a database to perform cross-database retrieval, and obtain a preliminary retrieval literature set and a cross-database causal network; screen and extract the preliminary retrieval literatures in the preliminary retrieval literature set according to the above large language model and a medical bias assessment tool, and obtain a draft data literature set and a clinical heterogeneity factor set; perform dynamic evidence chain fusion according to the above draft data literature set and the above clinical heterogeneity factor set to determine weights, and obtain a fusion effect size set and a conflict flag set; execute a heterogeneity judgment script according to the above fusion effect size set and the above clinical heterogeneity factor set to perform heterogeneity judgment, and obtain a visualization chart set; execute a standardized conflict resolution according to the above cross-database causal network and the above conflict flag set, and obtain a conflict resolution report; and determine the results of the above draft data literature set and the above visualization chart set by using the above large language model according to a medical draft template and the above conflict resolution report, and generate a medical draft.
[0152] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages - such as Java, Smalltalk, C++; and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or may be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).
[0153] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0154] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes a retrieval formula generation unit, a literature retrieval unit, a literature screening unit, a conflict marking unit, a chart generation unit, a report generation unit, and a draft generation unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the retrieval formula generation unit can also be described as "the unit that generates multi-database retrieval formulas corresponding to medical research directions using large language models".
[0155] The functions described above can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), System on Chip (SOC), Complex Programmable Logic Devices (CPLD), and so on.
[0156] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, technical solutions formed by mutually replacing the above features with (but not limited to) technical features having similar functions disclosed in the embodiments of the present disclosure.
Claims
1. A method for generating a medical draft based on a large model, comprising: Using a large language model to generate a multi-database retrieval formula corresponding to a medical research direction; Executing the multi-database retrieval formula in a database to perform cross-database retrieval, obtaining a preliminary retrieval literature set and a cross-database causal network; According to the large language model and a medical bias assessment tool, screening and extracting the preliminary retrieval literature in the preliminary retrieval literature set to obtain a draft data literature set and a clinical heterogeneity factor set; According to the draft data literature set and the clinical heterogeneity factor set, performing dynamic evidence chain fusion to determine weights, obtaining a fusion effect size set and a conflict marker set; According to the fusion effect size set and the clinical heterogeneity factor set, executing a heterogeneity judgment script to perform heterogeneity judgment, obtaining a visualization chart set; According to the cross-database causal network and the conflict marker set, performing standardized conflict resolution to obtain a conflict resolution report; According to a medical draft template and the conflict resolution report, using the large language model to determine the results of the draft data literature set and the visualization chart set, and generating a medical draft.
2. The method according to claim 1, wherein The step of using a large language model to generate a multi-database retrieval formula corresponding to a medical research direction includes: Extracting keywords from the medical research direction to obtain research keywords; Using the large language model to execute the following generation steps: Generating a single-database retrieval formula according to the research keywords; According to a database conversion document, performing syntax conversion on the single-database retrieval formula to generate a multi-database retrieval formula.
3. The method according to claim 1, wherein, The step of executing the multi-database retrieval formula in a database to perform cross-database retrieval, obtaining a preliminary retrieval literature set and a cross-database causal network, includes: Using an asynchronous function to parallelly execute the multi-database retrieval formula to obtain a retrieval literature set; Performing multi-dimensional standard screening on each retrieval literature in the retrieval literature set to generate preliminary retrieval literature, obtaining a preliminary retrieval literature set; Performing element extraction on each preliminary retrieval literature in the preliminary retrieval literature set to obtain a literature element information set, wherein the literature element information includes: research objects, intervention measures, result variables, and research designs in the literature; For each literature element information in the literature element information set, determining the corresponding relationship between the literature element information and the preliminary retrieval literature set to obtain a causal network corresponding set; Generating the cross-database causal network according to the literature element information and the corresponding retrieval literature in the causal network corresponding set.
4. The method according to claim 1, wherein, The step of screening and extracting the preliminary retrieval literature in the preliminary retrieval literature set according to the large language model and a medical bias assessment tool to obtain a draft data literature set and a clinical heterogeneity factor set includes: Obtaining a pre-generated data extraction template; Using the medical bias assessment tool to perform bias assessment on the preliminary retrieval literature to obtain a bias assessment result; In response to the bias assessment result reaching a preset value, performing content extraction on the literature content in the preliminary retrieval literature to obtain an extraction content set; For each extraction content in the extraction content set, performing the following preprocessing operations: Using the large language model, extract the clinical heterogeneity factors in the extracted content to obtain at least one clinical heterogeneity factor; Fill the extracted content into the data extraction template to obtain a draft data literature.
5. The method according to claim 1, wherein, Performing dynamic evidence chain fusion according to the draft data literature set and the clinical heterogeneity factor set to determine weights, and obtaining a fusion effect size set and a conflict marker set, including: Extract the structured data corresponding to the research characteristics in the draft data literature set to obtain a research characteristic content set; Perform standardization processing on each clinical heterogeneity factor in the clinical heterogeneity factor set to obtain a heterogeneity factor data set; Using a preset heterogeneity weight function, negatively correlate the research characteristic content set and the heterogeneity factor data set to obtain a weight set corresponding to the research characteristic content set; Determine the research effect size according to the medical research direction; According to the weight set, perform dynamic weight weighted average on the research effect size to obtain the fusion effect size set; For each fusion effect size in the fusion effect size set, in response to the fusion effect size reaching a preset threshold, determine that the research characteristic content corresponding to the fusion effect size is a conflict marker.
6. The method according to claim 1, wherein, Performing heterogeneity judgment by executing a heterogeneity judgment script according to the fusion effect size set and the clinical heterogeneity factor set to obtain a visualization chart set, including: Determine the effect size of the quantitative research results corresponding to the fusion effect size set to obtain an effect size data set; Determine the heterogeneity index of the clinical heterogeneity factors corresponding to the clinical heterogeneity factor set to obtain a heterogeneity index data set; According to the effect size data set and the heterogeneity index data set, execute the pre-generated heterogeneity judgment script according to corresponding requirements to obtain a statistic value and a heterogeneity index; Determine a heterogeneity effect model according to the heterogeneity index and the effect size data set; Generate a heterogeneity forest plot and a heterogeneity bar chart according to the heterogeneity effect model and the effect size data set; Generate a significance test chart according to the statistic value and the effect size data set; Determine the heterogeneity forest plot, the heterogeneity bar chart, and the significance test chart as the visualization chart set.
7. A medical draft generation device based on a large model, including: A retrieval-based generation unit configured to use a large language model to generate a multi-database retrieval formula corresponding to a medical research direction; A literature retrieval unit configured to perform cross-database retrieval by executing the multi-database retrieval formula in a database to obtain a preliminary retrieved literature set and a cross-database causal network; A literature screening unit configured to screen and extract the preliminary retrieved literature in the preliminary retrieved literature set according to the large language model and a medical bias assessment tool to obtain a draft data literature set and a clinical heterogeneity factor set; A conflict marking unit configured to perform dynamic evidence chain fusion according to the draft data literature set and the clinical heterogeneity factor set to determine weights, and obtain a fusion effect size set and a conflict marker set; A chart generation unit, configured to execute a heterogeneity evaluation script according to the set of fusion effect amounts and the set of clinical heterogeneity factors to perform heterogeneity judgment, and obtain a set of visualization charts; A report generation unit, configured to execute a standardized conflict resolution according to the cross-database causal network and the set of conflict markers, and obtain a conflict resolution report; A draft generation unit, configured to determine results for the draft data literature set and the set of visualization charts by using the large language model according to a medical draft template and the conflict resolution report, and generate a medical draft.
8. An electronic device, comprising: One or more processors; A storage device having stored thereon one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-6.
9. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, the method according to any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Automatic scientific and technological novelty search method and system
CN115344719A
Corpus database, corpus database maintenance method and device, equipment and medium
CN115495541A
Causal relationship strength evaluation method and device, equipment, storage medium and product
CN118569249A
Automatic paper first draft generation method and system and electronic equipment
CN118966162A
Medical retrieval method, related method, device, equipment and storage medium
CN119964827A
Cited By
Method and device for generating medical report based on evidence-based medicine
CN120932802A
Method and apparatus for generating medical reports based on evidence-based medicine
CN120932802B
Bid inviting and tendering template generation method based on large language model
CN121031563A