Meta analysis method and device, equipment and storage medium
By combining a pre-set large language model and a pre-trained language model, the meta-statistical analysis process is automated, solving the problem of low intelligence in existing methods and achieving efficient and accurate generation of analysis results.
Patent Information
- Application Number
- CN202511259227.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2025-10-28
AI Technical Summary
Existing meta-statistical analysis methods have low intelligence levels, literature retrieval strategies rely on manual labor, and are prone to missing key studies or introducing biases. Data processing consumes a lot of manpower and is highly subjective, and analysis results are static and lagging and cannot reflect the latest evidence in real time.
Through the preset large language model, semantic analysis is performed on the research questions input by the user, the target structured form is generated, retrieval is performed based on the structured form, the pre-trained language model is used to process the literature and extract structured information, statistical analysis is performed and visualization results are generated.
It automates meta-statistical analysis, reduces reliance on professionals, shortens analysis cycles, reduces human error, improves the intelligence, efficiency, and accuracy of the analysis, optimizes retrieval strategies, and enhances the performance and accuracy of the analysis.
Smart Images

Figure CN120849602A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to meta-analysis methods, apparatus, devices and storage media. Background Technology
[0002] Meta-statistical analysis is an important statistical analysis method that improves statistical power and the reliability of conclusions by systematically integrating the results of multiple independent studies. It is widely used in fields such as medicine, public health, and clinical decision-making. Its core value lies in reducing the bias of a single study and providing more comprehensive and objective evidence-based support.
[0003] Currently, traditional meta-statistical analysis processes suffer from numerous bottlenecks, including the reliance on manually constructed literature retrieval strategies, which can easily overlook key studies or introduce biases; the high manpower required for literature screening, data extraction, and quality assessment, which are also highly subjective; the near impracticality of analyzing large-scale literature; and the static lag in analysis results, which cannot reflect the latest evidence in real time.
[0004] Therefore, existing meta-statistical analysis methods suffer from a low level of intelligence. Summary of the Invention
[0005] This application aims to at least solve the technical problems existing in the prior art. To this end, the first aspect of this application proposes a meta-analysis method, which includes: The research questions input by users are semantically parsed using a pre-set large language model to generate a target structured form; Searching based on the target structured form yields the target documents; The target documents are processed by a pre-trained language model to extract structured information, which includes tabular and graphical information. Statistical analysis is performed based on structured information to generate visualization results, and a statistical analysis report is obtained based on the visualization results.
[0006] In one possible implementation, the research question input by the user is semantically parsed using a pre-defined large language model to generate a target structured form, including: The semantic analysis results are obtained by performing semantic analysis on the research questions input by users through a pre-set large language model. Based on a pre-defined knowledge graph, the semantic parsing results are standardized and mapped and expanded with synonyms to generate the target structured form.
[0007] In one possible implementation, the search is performed based on the target structured form to obtain the target documents, including: Based on the target structured form and the preset large language model, a standard Boolean logic retrieval expression is generated; The target documents are retrieved by searching a pre-defined database using Boolean logic search.
[0008] In one possible implementation, based on the target structured form and a pre-defined large language model, a standard Boolean logic retrieval expression is generated, including: Based on the target structured form and the preset large language model, generate intelligent search queries; Based on a pre-defined knowledge graph, the intelligent search query is semantically expanded to generate an expanded search query. The expanded search query is processed by natural language conversion to generate a standard Boolean logic search query.
[0009] In one possible implementation, the target document is processed using a pre-trained language model to extract structured information, including: For each target document, the title and abstract information of the target document are transformed and processed using a pre-trained language model to obtain semantic vectors; Similarity is calculated based on semantic vectors to obtain similarity results; Based on the similarity results, the target documents were deduplicated and clustered to obtain the processing results; Based on the processing results, structured information is extracted.
[0010] In one possible implementation, based on the processing result, structured information is extracted, including: The processing results are analyzed using optical character recognition and natural language processing algorithms to extract document element information, including intervention methods, sample size, research outcomes, follow-up time, and statistical methods. Multimodal structured information extraction is performed based on a pre-defined large language model and document element information to obtain structured information.
[0011] In one possible implementation, statistical analysis based on structured information and the generation of visualization results include: The structured information is transformed and standardized to generate an integrated result; Obtain the identification results after quality assessment and bias risk identification of the target documents; Based on the identification results, statistical analysis is performed on the integrated results and visualization results are generated.
[0012] A second aspect of this application provides a meta-analysis apparatus, the apparatus comprising: The generation module is used to perform semantic parsing on the research questions input by the user through a preset large language model, and generate the target structured form; The retrieval module is used to retrieve target documents based on the target structured form. The extraction module is used to process the target document using a pre-trained language model to extract structured information, including tabular and graphical information. The analysis module is used to perform statistical analysis based on structured information and generate visualization results, and then generate a statistical analysis report based on the visualization results.
[0013] In one possible implementation, the above-mentioned generation module is specifically used for: The semantic analysis results are obtained by performing semantic analysis on the research questions input by users through a pre-set large language model. Based on a pre-defined knowledge graph, the semantic parsing results are standardized and mapped and expanded with synonyms to generate the target structured form.
[0014] In one possible implementation, the retrieval module is specifically used for: Based on the target structured form and the preset large language model, a standard Boolean logic retrieval expression is generated; The target documents are retrieved by searching a pre-defined database using Boolean logic search.
[0015] In one possible implementation, the retrieval module is further used for: Based on the target structured form and the preset large language model, generate intelligent search queries; Based on a pre-defined knowledge graph, the intelligent search query is semantically expanded to generate an expanded search query. The expanded search query is processed by natural language conversion to generate a standard Boolean logic search query.
[0016] In one possible implementation, the extraction module described above is specifically used for: For each target document, the title and abstract information of the target document are transformed and processed using a pre-trained language model to obtain semantic vectors; Similarity is calculated based on semantic vectors to obtain similarity results; Based on the similarity results, the target documents were deduplicated and clustered to obtain the processing results; Based on the processing results, structured information is extracted.
[0017] In one possible implementation, the extraction module is further configured to: The processing results are analyzed using optical character recognition and natural language processing algorithms to extract document element information, including intervention methods, sample size, research outcomes, follow-up time, and statistical methods. Multimodal structured information extraction is performed based on a pre-defined large language model and document element information to obtain structured information.
[0018] In one possible implementation, the analysis module described above is specifically used for: The structured information is transformed and standardized to generate an integrated result; Obtain the identification results after quality assessment and bias risk identification of the target documents; Based on the identification results, statistical analysis is performed on the integrated results and visualization results are generated.
[0019] A third aspect of this application provides an electronic device comprising a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the statistical analysis method as described in the first aspect.
[0020] The fourth aspect of this application provides a computer-readable storage medium storing at least one instruction, at least one program, a code set, or an instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the statistical analysis method as described in the first aspect.
[0021] The embodiments of this application have the following beneficial effects: The meta-analysis method provided in this application includes: semantic parsing of user-inputted research questions using a pre-set large language model to generate a target structured form; searching based on the target structured form to obtain target literature; processing the target literature using a pre-trained language model to extract structured information; wherein the structured information includes tabular and graphical information; performing statistical analysis based on the structured information and generating visualization results; and obtaining a statistical analysis report based on the visualization results. This solution provides efficient and reliable automated tools for evidence-based medicine by automating semantic parsing of user-inputted research questions using a pre-set large language model and processing target literature using a pre-trained language model. This reduces the reliance on professionals in meta-statistical analysis, shortens the analysis cycle, reduces human error, and significantly improves the intelligence, efficiency, accuracy, and reproducibility of meta-statistical analysis, while promoting the standardization and intelligent development of evidence-based medicine research. Furthermore, by searching based on the target structured form, the search strategy is optimized, improving search accuracy, thereby further enhancing the performance and accuracy of meta-statistical analysis. Attached Figure Description
[0022] Figure 1 A block diagram of a computer device provided in an embodiment of this application; Figure 2A flowchart illustrating the steps of a meta-analysis method provided in this application embodiment; Figure 3 A flowchart illustrating the steps for generating a target structured form is provided in this application embodiment. Figure 4 A flowchart illustrating the steps for obtaining a target document is provided in this embodiment of the application. Figure 5 A flowchart illustrating the steps for generating a standard Boolean logic retrieval expression is provided in this application embodiment. Figure 6 A flowchart illustrating the steps for extracting structured information, as provided in this application embodiment; Figure 7 A flowchart illustrating another step for extracting structured information, as provided in an embodiment of this application; Figure 8 A flowchart illustrating the steps for generating a visual result is provided in this application embodiment. Figure 9 This is a structural block diagram of a meta-analysis device provided in an embodiment of this application. Detailed Implementation
[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0024] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of this disclosure, unless otherwise stated, "a plurality of" means two or more. Furthermore, the use of "based on" or "according to" implies openness and inclusiveness, because processes, steps, calculations, or other actions "based on" or "according to" one or more of the stated conditions or values may in practice be based on additional conditions or beyond the stated values.
[0025] The statistical analysis method provided in this application can be applied to computer equipment (electronic devices). The computer equipment can be a server or a terminal. The server can be a single server or a server cluster composed of multiple servers. This application does not specifically limit this. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets and portable wearable devices.
[0026] Taking a computer device as an example, Figure 1 A block diagram of a server is shown, such as Figure 1 As shown, the server may include a processor and memory connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. When the computer program is executed by the processor, it implements a meta-analysis method.
[0027] Those skilled in the art will understand that Figure 1 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the server to which the present application is applied. Optionally, the server may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0028] It should be noted that the execution subject of the embodiments of this application can be a computer device or a statistical analysis device. The following method embodiments will be described with a computer device as the execution subject.
[0029] Figure 2 This is a flowchart illustrating the steps of a meta-analysis method provided in an embodiment of this application. Figure 2 As shown, the method includes the following steps: Step 202: Semantically analyze the research questions input by the user using a pre-set large language model to generate the target structured form.
[0030] Users can describe their research questions in natural language; for example, a research question could be "Analyze whether SGLT2 inhibitors can reduce the hospitalization rate of heart failure in elderly diabetic patients." Then, a pre-trained, pre-defined large language model can be used to semantically parse the user-input research question, thereby generating the target structured form.
[0031] In some alternative embodiments, the preset large language model can be the DeepSeek large language model, so that the user-input research question can be semantically parsed by calling the DeepSeek large language model.
[0032] In some alternative embodiments, such as Figure 3 As shown, Figure 3 A flowchart illustrating the steps for generating a target structured form, as provided in this application embodiment, includes: Step 302: Semantically analyze the research question input by the user using a pre-set large language model to obtain the semantic analysis results.
[0033] Step 304: Based on the preset knowledge graph, standardize the semantic parsing results through mapping and synonym expansion to generate the target structured form.
[0034] Specifically, a pre-defined large language model is used to semantically analyze the research questions input by users, thereby automatically identifying PICOS elements and obtaining semantic analysis results. Here, P represents the research subject, I represents the intervention, C represents the control, O represents the research outcome, and S represents the research design.
[0035] For example, the research question could be "Analyze whether SGLT2 inhibitors can reduce the hospitalization rate of heart failure in elderly diabetic patients," and the semantic parsing results could include: the research subjects are elderly diabetic patients, the intervention is a sodium-dependent glucose transporter 2 (SGLT-2) inhibitor, the control is a placebo or other hypoglycemic drugs, the research outcome is the hospitalization rate of heart failure, and the research design is a randomized controlled trial (RCT).
[0036] Next, the semantic parsing results can be standardized and expanded using synonyms based on a pre-defined knowledge graph to generate the target structured form. Specifically, after standardizing and expanding the semantic parsing results using a pre-defined knowledge graph, a preliminary structured form can be obtained, as shown in the table below.
[0037]
[0038] Therefore, the preliminary structured form can be displayed on the front-end page, where users can modify, supplement, or confirm the various elements. After confirmation, the target structured form can be obtained, and the search process can be triggered.
[0039] Step 204: Perform a search based on the target structured form to obtain the target documents.
[0040] After generating the aforementioned target structured form, a search can be performed based on the target structured form to obtain the target documents. In some optional embodiments, such as... Figure 4 As shown, Figure 4 A flowchart of steps for obtaining a target document provided in this application embodiment includes: Step 402: Generate a standard Boolean logic retrieval expression based on the target structured form and the preset large language model.
[0041] Step 404: Search the preset database based on Boolean logic search to retrieve the target documents.
[0042] Among these, based on the target structured form, a standard Boolean logic retrieval expression can be generated by calling a preset large language model. In some optional embodiments, such as Figure 5 As shown, Figure 5 A flowchart illustrating the steps for generating a standard Boolean logic retrieval expression, provided in this application embodiment, includes: Step 502: Generate intelligent search queries based on the target structured form and the preset large language model.
[0043] Step 504: Based on the preset knowledge graph, semantically expand the intelligent search query to generate an expanded search query.
[0044] Step 506: Perform natural language conversion on the expanded search query to generate a standard Boolean logic search query.
[0045] Specifically, by invoking a pre-defined large language model, intelligent search queries can be generated based on the target structured form. Then, semantic expansion of the intelligent search queries can be performed based on a pre-defined knowledge graph to generate extended search queries, which can dynamically call upon the pre-defined knowledge graph to expand related disease, drug, and intervention terms. Finally, natural language processing can be performed on the extended search queries to generate standard Boolean logic search queries.
[0046] Next, the system can retrieve target documents from the preset databases based on Boolean logic search queries. Specifically, the system can jointly call preset database interfaces and, based on a distributed task queue and parallel scheduling mechanism, can synchronously and efficiently execute retrieval and download tasks across multiple preset databases, supporting the retrieval and preliminary preprocessing of tens of thousands of documents.
[0047] For example, the preset database may include, but is not limited to, core academic databases, authoritative grey literature platforms, and clinical trial registration platforms.
[0048] Step 206: Process the target document using a pre-trained language model to extract structured information.
[0049] The structured information can include tabular and graphical information. After obtaining the target document, a pre-trained language model can be used to process the document and extract the structured information. In some optional embodiments, such as... Figure 6 As shown, Figure 6 A flowchart illustrating the steps for extracting structured information according to an embodiment of this application includes: Step 602: For each target document, the title and abstract information of the target document are transformed and processed by a pre-trained language model to obtain a semantic vector.
[0050] Step 604: Calculate similarity based on semantic vectors to obtain similarity results.
[0051] Step 606: Based on the similarity results, perform deduplication and clustering on each target document to obtain the processing results.
[0052] Step 608: Based on the processing results, extract the structured information.
[0053] The pre-trained language model can be a Bidirectional Encoder Representations from Transformers (BERT) model, or other types of models. This application does not specifically limit this.
[0054] For each target document, the title and abstract information can be transformed using a pre-trained language model to obtain semantic vectors. Next, similarity calculations can be performed based on these semantic vectors to obtain similarity results. Then, based on these results, clustering algorithms are applied to automatically identify highly similar or duplicate documents, performing deduplication and clustering on each target document to obtain the final processing result. The clustering algorithm can be density clustering or hierarchical clustering, among others. Through innovative semantic clustering and deduplication algorithms, efficient processing and intelligent management of massive amounts of documents are achieved, significantly improving the performance and accuracy of meta-statistical analysis. Furthermore, intelligent document screening reduces manual intervention and lowers labor costs.
[0055] Ultimately, structured information can be extracted based on the processing results, such as in some optional embodiments. Figure 7 As shown, Figure 7 Another flowchart of steps for extracting structured information provided in this application embodiment includes: Step 702: Analyze the processing results based on optical character recognition algorithm and natural language processing algorithm to extract document element information.
[0056] Step 704: Based on the preset large language model and document element information, perform multimodal structured information extraction processing to obtain structured information.
[0057] Among them, the Optical Character Recognition (OCR) and Natural Language Processing (NLP) algorithms can perform in-depth analysis of the processing results to extract document element information, which may include intervention methods, sample size, research outcomes, follow-up time and statistical methods.
[0058] Next, multimodal structured information extraction processing can be performed based on a preset large language model and document element information to obtain structured information, which may include tabular information and graphical information.
[0059] Step 208: Perform statistical analysis based on structured information and generate visualization results. Obtain a statistical analysis report based on the visualization results.
[0060] Statistical analysis is performed based on the structured information above, and visualization results are generated. Optionally, after obtaining the structured information above, quality assessment and bias risk identification can be performed to complete a systematic evaluation of the methodological quality of the included target literature and identify potential biases.
[0061] Specifically, the system can first automatically determine the research type and match it with appropriate evaluation tools. Then, it introduces a hierarchical attention mechanism to model the paragraph structure of the target document. By calling a pre-set large language model to assist in understanding ambiguous statements, it extracts quality-related factors such as randomization, blinding, and data integrity. Finally, it automatically generates identification results such as bias charts, scoring matrices, and research quality assessment suggestions. Through intelligent document quality evaluation, manual intervention is reduced, thus lowering labor costs.
[0062] In some alternative embodiments, such as Figure 8 As shown, Figure 8 A flowchart illustrating the steps for generating a visualization result, as provided in this application embodiment, includes: Step 802: Transform and standardize the structured information to generate the integrated result.
[0063] Step 804: Obtain the identification results after quality assessment and bias risk identification of the target documents.
[0064] Step 806: Based on the identification results, perform statistical analysis on the integrated results and generate visualization results.
[0065] This process involves transforming and standardizing structured information to generate integrated results, thereby unifying the types of effect indicators across studies and achieving comparability and integration. Specifically, the effect indicators used in the literature can be identified first, then transformed using an indicator transformation rule base. The transformed results can then be standardized and integrated to finally generate the integrated results.
[0066] Next, the identification results obtained after the above quality assessment and bias risk identification of the target documents can be obtained, and statistical analysis can be performed on the integrated results based on the identification results to generate visualization results.
[0067] Optionally, a meta-statistical analysis tool can be packaged first, and then statistical analysis can be performed on the integrated results, including heterogeneity testing, sensitivity analysis, bias estimation, effect pooling, and subgroup analysis. Finally, forest plots, funnel plots, etc. can be automatically generated, and statistical charts can be interpreted through a preset large language model to perform natural language report interpretation and determine their clinical significance.
[0068] In some optional embodiments, a complete meta-analysis report conforming to standards can also be generated, which can be exported and reused. Specifically, a template engine and a multi-paragraph generation mechanism can be built, and the research background, retrieval strategy, screening flowchart, research feature table, analysis results and interpretations can be automatically written through multimodal data and text collaborative typesetting technology. After embedding charts and structured data, an automated report is obtained. In addition, the automated report can be exported in multiple formats.
[0069] In some alternative embodiments, the system performing the above-described statistical analysis methods can also achieve closed-loop operation of the system process and support embedding into scientific research platforms and deployment of intelligent services. Specifically, the system can also manage and monitor workflow tasks through a process tracking module, access scientific research platform interfaces (such as clinical research systems and literature management systems) through interface and permission management, and achieve parameter debugging and process reuse through system operation logs and error tracking mechanisms.
[0070] This application provides a meta-analysis method, which includes: semantically parsing the user-input research question using a pre-set large language model to generate a target structured form; retrieving target literature based on the target structured form; processing the target literature using a pre-trained language model to extract structured information, wherein the structured information includes tabular and graphical information; performing statistical analysis based on the structured information and generating visualization results; and generating a statistical analysis report based on the visualization results. This approach provides efficient and reliable automated tools for evidence-based medicine by automating the semantic parsing of the user-input research question using a pre-set large language model and processing the target literature using a pre-trained language model. This reduces the reliance on professionals in meta-statistical analysis, shortens the analysis cycle, reduces human error, and significantly improves the intelligence, efficiency, accuracy, and reproducibility of meta-statistical analysis, while promoting the standardization and intelligent development of evidence-based medicine research. Furthermore, by retrieving based on the target structured form, the retrieval strategy is optimized, improving retrieval accuracy, thereby further enhancing the performance and accuracy of meta-statistical analysis.
[0071] Figure 9 This is a structural block diagram of a meta-analysis device provided in an embodiment of this application.
[0072] like Figure 9As shown, the meta-analysis device 900 includes: The generation module 902 is used to perform semantic parsing on the research questions input by the user through a preset large language model and generate the target structured form.
[0073] Search module 904 is used to perform searches based on the target structured form to obtain the target documents.
[0074] The extraction module 906 is used to process the target document through a pre-trained language model to extract structured information, which includes tabular and graphical information.
[0075] Analysis module 908 is used to perform statistical analysis based on structured information and generate visualization results, and to obtain a statistical analysis report based on the visualization results.
[0076] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operations has been described in detail in the embodiments related to the method, and will not be elaborated upon here. Each module in the above statistical analysis apparatus can be implemented entirely or partially through software, hardware, or a combination thereof. Each module can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations of each module.
[0077] In one embodiment of this application, a computer device is provided, the computer device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps: The research questions input by users are semantically parsed using a pre-set large language model to generate a target structured form; Searching based on the target structured form yields the target documents; The target documents are processed by a pre-trained language model to extract structured information, which includes tabular and graphical information. Statistical analysis is performed based on structured information to generate visualization results, and a statistical analysis report is obtained based on the visualization results.
[0078] In one embodiment of this application, the processor further performs the following steps when executing the computer program: The semantic analysis results are obtained by performing semantic analysis on the research questions input by users through a pre-set large language model. Based on a pre-defined knowledge graph, the semantic parsing results are standardized and mapped and expanded with synonyms to generate the target structured form.
[0079] In one embodiment of this application, the processor further performs the following steps when executing the computer program: Based on the target structured form and the preset large language model, a standard Boolean logic retrieval expression is generated; The target documents are retrieved by searching a pre-defined database using Boolean logic search.
[0080] In one embodiment of this application, the processor further performs the following steps when executing the computer program: Based on the target structured form and the preset large language model, generate intelligent search queries; Based on a pre-defined knowledge graph, the intelligent search query is semantically expanded to generate an expanded search query. The expanded search query is processed by natural language conversion to generate a standard Boolean logic search query.
[0081] In one embodiment of this application, the processor further performs the following steps when executing the computer program: For each target document, the title and abstract information of the target document are transformed and processed using a pre-trained language model to obtain semantic vectors; Similarity is calculated based on semantic vectors to obtain similarity results; Based on the similarity results, the target documents were deduplicated and clustered to obtain the processing results; Based on the processing results, structured information is extracted.
[0082] In one embodiment of this application, the processor further performs the following steps when executing the computer program: The processing results are analyzed using optical character recognition and natural language processing algorithms to extract document element information, including intervention methods, sample size, research outcomes, follow-up time, and statistical methods. Multimodal structured information extraction is performed based on a pre-defined large language model and document element information to obtain structured information.
[0083] In one embodiment of this application, the processor further performs the following steps when executing the computer program: The structured information is transformed and standardized to generate an integrated result; Obtain the identification results after quality assessment and bias risk identification of the target documents; Based on the identification results, statistical analysis is performed on the integrated results and visualization results are generated.
[0084] The computer device provided in this application embodiment has a similar implementation principle and technical effect to the above method embodiment, and will not be described again here.
[0085] In one embodiment of this application, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, it performs the following steps: The research questions input by users are semantically parsed using a pre-set large language model to generate a target structured form; Searching based on the target structured form yields the target documents; The target documents are processed by a pre-trained language model to extract structured information, which includes tabular and graphical information. Statistical analysis is performed based on structured information to generate visualization results, and a statistical analysis report is obtained based on the visualization results.
[0086] In one embodiment of this application, the computer program, when executed by a processor, further performs the following steps: The semantic analysis results are obtained by performing semantic analysis on the research questions input by users through a pre-set large language model. Based on a pre-defined knowledge graph, the semantic parsing results are standardized and mapped and expanded with synonyms to generate the target structured form.
[0087] In one embodiment of this application, the computer program, when executed by a processor, further performs the following steps: Based on the target structured form and the preset large language model, a standard Boolean logic retrieval expression is generated; The target documents are retrieved by searching a pre-defined database using Boolean logic search.
[0088] In one embodiment of this application, the computer program, when executed by a processor, further performs the following steps: Based on the target structured form and the preset large language model, generate intelligent search queries; Based on a pre-defined knowledge graph, the intelligent search query is semantically expanded to generate an expanded search query. The expanded search query is processed by natural language conversion to generate a standard Boolean logic search query.
[0089] In one embodiment of this application, the computer program, when executed by a processor, further performs the following steps: For each target document, the title and abstract information of the target document are transformed and processed using a pre-trained language model to obtain semantic vectors; Similarity is calculated based on semantic vectors to obtain similarity results; Based on the similarity results, the target documents were deduplicated and clustered to obtain the processing results; Based on the processing results, structured information is extracted.
[0090] In one embodiment of this application, the computer program, when executed by a processor, further performs the following steps: The processing results are analyzed using optical character recognition and natural language processing algorithms to extract document element information, including intervention methods, sample size, research outcomes, follow-up time, and statistical methods. Multimodal structured information extraction is performed based on a pre-defined large language model and document element information to obtain structured information.
[0091] In one embodiment of this application, the computer program, when executed by a processor, further performs the following steps: The structured information is transformed and standardized to generate an integrated result; Obtain the identification results after quality assessment and bias risk identification of the target documents; Based on the identification results, statistical analysis is performed on the integrated results and visualization results are generated.
[0092] The computer-readable storage medium provided in this embodiment is similar in principle and technical effect to the method embodiment described above, and will not be repeated here.
[0093] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0094] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0095] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A meta-analysis method, characterized in that, The method includes: The research questions input by users are semantically parsed using a pre-set large language model to generate a target structured form; The target documents are obtained by searching based on the target structured form. The target document is processed by a pre-trained language model to extract structured information; wherein, the structured information includes table information and diagram information; Statistical analysis is performed based on the structured information, and visualization results are generated. A statistical analysis report is then obtained based on the visualization results.
2. The method according to claim 1, characterized in that, The step of semantically parsing the research question input by the user through a preset large language model to generate a target structured form includes: The semantic analysis results are obtained by performing semantic analysis on the research questions input by users through a pre-set large language model. Based on a pre-defined knowledge graph, the semantic parsing results are standardized and mapped and expanded with synonyms to generate the target structured form.
3. The method according to claim 1 or 2, characterized in that, The search based on the target structured form to obtain target documents includes: Based on the target structured form and the preset large language model, a standard Boolean logic retrieval expression is generated; The target document is retrieved from the preset database based on the Boolean logic search expression.
4. The method according to claim 3, characterized in that, The step of generating a standard Boolean logic retrieval expression based on the target structured form and the preset large language model includes: Based on the target structured form and the preset large language model, an intelligent search query is generated; The intelligent search query is semantically expanded based on a preset knowledge graph to generate an expanded search query. The extended search expression is processed by natural language conversion to generate the standard Boolean logic search expression.
5. The method according to claim 1 or 2, characterized in that, The process of processing the target document using a pre-trained language model to extract structured information includes: For each of the target documents, the title and abstract information of the target documents are transformed and processed using a pre-trained language model to obtain semantic vectors; Similarity is calculated based on the semantic vectors to obtain similarity results; Based on the similarity results, the target documents are deduplicated and clustered to obtain the processing results; Based on the processing results, the structured information is extracted.
6. The method according to claim 5, characterized in that, The extraction of structured information based on the processing result includes: The processing results are analyzed using optical character recognition and natural language processing algorithms to extract document element information; wherein, the document element information includes intervention method, sample size, research outcome, follow-up time and statistical method; Based on the preset large language model and the document element information, multimodal structured information extraction processing is performed to obtain the structured information.
7. The method according to claim 1 or 2, characterized in that, The statistical analysis based on the structured information and the generation of visualization results include: The structured information is transformed and standardized to generate an integrated result; Obtain the identification results after performing quality assessment and bias risk identification on the target documents; Based on the identification results, statistical analysis is performed on the integrated results, and the visualization results are generated.
8. A meta-analysis apparatus, characterized in that, The device includes: The generation module is used to perform semantic parsing on the research questions input by the user through a preset large language model, and generate the target structured form; The retrieval module is used to retrieve target documents based on the target structured form. The extraction module is used to process the target document through a pre-trained language model to extract structured information; wherein, the structured information includes tabular information and graphical information; The analysis module is used to perform statistical analysis based on the structured information and generate visualization results, and to obtain a statistical analysis report based on the visualization results.
9. An electronic device, characterized in that, The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the meta-analysis method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the meta-analysis method as described in any one of claims 1-7.
Citation Information
Cited By
Meta analysis method, system and equipment based on artificial intelligence and storage medium
CN121935265A
Meta analysis RCT outcome index structure analysis and computable input method and Meta analysis RCT outcome index structure analysis and computable input system
CN122334235A
A Meta-analysis Method and System for Analyzing and Computable Input of Outcome Indicators in RCTs
CN122334235B