A system and method for automatically generating reviews based on artificial intelligence
Through an automatic generation review system based on artificial intelligence, the problem of low efficiency and accuracy of medical literature processing is solved, and the full process automation from retrieval to analysis is achieved, and scientific research efficiency and clinical transformation capabilities are improved.
Patent Information
- Application Number
- CN202411865771.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2044-12-18
AI Technical Summary
The existing technology is difficult to extract high-quality information from a large number of medical literature quickly and accurately, and there is a lag in the transformation of scientific research results into clinical applications, the literature review generation efficiency is low, and the scientific research efficiency and accuracy are insufficient.
An automatic generation review system based on artificial intelligence is adopted, including a search-based construction module, a literature screening module, a scale evaluation module, a data extraction and information summary module and a system maintenance module, and automated processing is used by PubMed's MeSH interface, Langchain-ChatChat's knowledge base, OCR technology, and Python's Matplotlib and Python-docx technology.
It has achieved the automation of the entire process from literature retrieval to data analysis, improved the efficiency and accuracy of literature processing, and promoted the clinical transformation of scientific research results and cross-disciplinary cooperation.
Smart Images

Figure CN119322842B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data classification, and in particular to a system and method for automatically generating summaries based on artificial intelligence. Background Art
[0002] There are strong barriers between disciplines, and medical students are unfamiliar with and lack understanding of professional knowledge in other branches; there is a large amount of literature, and the quality varies greatly. How to quickly extract and scientifically analyze high-quality information and data has become a puzzle in scientific research; scientific research lags behind clinical needs, and there is also a phenomenon of low conversion of scientific research into clinical applications.
[0003] Teaching - Difficulty in student training: The number of medical literature has increased dramatically. According to PubMed statistics, the number of new medical literature has increased by about 40,000 per year in the past year. It is difficult for students to read and organize the literature and integrate knowledge!
[0004] Scientific research - difficult, large in volume, high in error rate, and time-consuming: Traditional splicing of subject terms, free words, and manual truncation of words in literature retrieval is time-consuming and has a high error rate; massive amounts of articles are retrieved during literature screening, and manual screening to eliminate redundant articles has a high error rate; scale evaluation requires detailed reading of the background methods and conclusions of a large number of documents, resulting in poor accuracy in summarization and organization; conclusion / data summary design involves multiple types of data, which is difficult, large in volume, and lacks scientificity.
[0005] Clinical - The conversion of scientific research results lags behind, and precision medicine has a long way to go: There is a lag in the conversion of scientific research results into clinical applications, the diagnosis and treatment plans are updated slowly, and it is difficult to be personalized and advanced. Scientific research results are difficult to apply to precision medicine, and patient benefits are low. In 2019, the National Institutes of Health of the United States invested more than US$100 million in the field of precision medicine, but the conversion rate was only 10%.
[0006] Although document management software such as Endnote and Noteexpress have emerged, they can only manage documents but are still unable to classify and identify document contents.
[0007] Although people have proposed literature reviews and used them to allow other scholars to quickly understand the latest progress in a certain field, and review generation software has emerged, this is obviously an inefficient method given the ever-changing scope of literature that needs to be summarized and the huge workload.
[0008] Although project collaboration software supports document collaboration and exchange and circulation within the team, and improves efficiency to a certain extent, the degree of improvement is still inseparable from manual labor.
[0009] Therefore, there is an urgent need for a product with strong content aggregation capabilities, accurate and comprehensive data extraction, analysis and summary capabilities, which can improve research efficiency and the accuracy of scientific decision-making. Summary of the Invention
[0010] The present invention aims to overcome the shortcomings of the prior art and provide a system for automatically generating reviews based on artificial intelligence, comprising: a search formula construction module, a literature screening module, a scale evaluation module, a data extraction and information aggregation module, an output and application module, and a system maintenance and update module; the search formula construction module, the literature screening module, the scale evaluation module, the data extraction and information aggregation module, the output and application module, and the system maintenance and update module are connected in sequence;
[0011] The search formula building module is used to receive keywords and search conditions input by the user, generate a search formula, and search for relevant documents based on the search formula;
[0012] The literature screening module is used to evaluate and screen the quality of retrieved literature;
[0013] The scale evaluation module is used to read and extract key data from the screened literature and conduct scale evaluation;
[0014] The data extraction and information summary module is used to summarize the extracted information and generate reports;
[0015] Output and apply module output generated reports;
[0016] The system maintenance and update module is used to perform system maintenance and updates based on user feedback and system evaluation results.
[0017] Preferably, the search formula building module is used to receive keywords and search conditions input by the user, generate a search formula, and search for relevant documents according to the search formula, including:
[0018] The search formula construction module receives keywords and search conditions input by the user, automatically generates a search formula through PubMed's MeSH interface and natural language processing model, and searches for corresponding documents in the medical database based on the search formula;
[0019] The search formula construction module includes an optimization unit, which is used to dynamically adjust the natural language processing model parameters based on the user's historical search records and feedback.
[0020] Preferably, the document screening module is used to perform quality assessment and screening on the retrieved documents, including:
[0021] Through the knowledge base and vector database of Langchain-ChatChat, combined with the fine-tuning of OCR and deep learning models, the retrieved documents are quality assessed and screened to select the corresponding medical literature.
[0022] Preferably, the scale evaluation module is used to read and extract key data from the screened literature and perform scale evaluation, including:
[0023] Through table OCR and semantic recognition, based on the chatpdf and chatdoc interfaces, key data from the screened literature is read and extracted; the key data includes the title, abstract, publication date, author, research method, research sample size, research population characteristics, corresponding dependent variable indicator information of the research method, and the effect size of the research results;
[0024] The scale evaluation module includes a verification unit for cross-validating the extracted key data.
[0025] Preferably, the data extraction and information aggregation module is used to aggregate the extracted information and generate a report, including:
[0026] The extracted effect sizes of each research result are summarized, along with the groups and specific scopes of application of the effect sizes. The summarized information is formatted using regular expressions, and an automated report with charts and multi-style themes is generated using Python's Matplotlib and Python-docx.
[0027] The data extraction and information aggregation module includes a custom report generation unit, which is used to generate corresponding reports based on the report style, chart type and content layout selected by the user.
[0028] Preferably, the system maintenance and update module further comprises a collaboration platform for multi-user online collaborative editing, annotation and sharing of reports.
[0029] A method for automatically generating a review based on artificial intelligence, using a system for automatically generating a review based on artificial intelligence, comprises the following steps:
[0030] S1, search formula construction, through the multi-task design mode of PubMed's MeSH interface and the language output of the NLP model, AI automatically constructs the search formula and formulates the inclusion and exclusion criteria;
[0031] S2, literature screening, manually created dataset screening software was used to fine-tune the 7BBaichuan language model to filter literature types by abstract;
[0032] S3, scale evaluation, uses Langchain.Chatchat's knowledge base and vector database technology, OCR technology, and deep learning to fine-tune a large language model for literature screening;
[0033] S4, data extraction, using table OCR technology and language recognition technology on the basis of ChatPDF and ChatDoc interfaces to condense and extract data from the selected articles;
[0034] S5, information summary, summarizes the extracted literature data information, uses regular formatting to apply Python's Matplotlib and Python-docxE operation technology, and outputs charts and multi-style theme automated PDF reports.
[0035] The beneficial effects of the present invention are: through the MeSH interface of PubMed and the language output multi-task design mode of the NLP model, AI is used to automatically construct search formulas, formulate inclusion and exclusion criteria, and improve search efficiency and accuracy.
[0036] Improve the efficiency of literature screening: Through Langchain-ChatChat's knowledge base and vector database technology, OCR technology, and deep learning to fine-tune the large language model, the screening function is realized, which improves the efficiency and accuracy of literature screening.
[0037] Improve data extraction efficiency: Through table OCR technology and semantic recognition technology, based on the chatpdf and chatdoc interfaces, AI can read the selected articles and perform data concentration and extraction, improving data extraction efficiency and accuracy.
[0038] Improve report generation efficiency: Through Python's Matplotlib and Python-docx interoperability technology, output charts and automated PDF reports with multiple styles and themes, improving report generation efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 A schematic diagram of the principle of a system for automatically generating reviews based on artificial intelligence;
[0040] Figure 2 Schematic diagram of the implementation of a system for automatically generating reviews based on artificial intelligence. DETAILED DESCRIPTION
[0041] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings, but the protection scope of the present invention is not limited to the following.
[0042] The features and performance of the present invention are further described in detail below with reference to the embodiments.
[0043] like Figure 1As shown, a system for automatically generating reviews based on artificial intelligence includes: a search formula construction module, a literature screening module, a scale evaluation module, a data extraction and information aggregation module, an output and application module, and a system maintenance and update module;
[0044] The search formula building module is used to receive keywords and search conditions input by the user, generate a search formula, and search for relevant documents based on the search formula;
[0045] The literature screening module is used to evaluate and screen the quality of retrieved literature;
[0046] The scale evaluation module is used to read and extract key data from the screened literature and conduct scale evaluation;
[0047] The data extraction and information summary module is used to summarize the extracted information and generate reports;
[0048] Output and apply module output generated reports;
[0049] The system maintenance and update module is used to perform system maintenance and updates based on user feedback and system evaluation results.
[0050] The search formula building module is used to receive keywords and search conditions input by the user, generate a search formula, and search for relevant documents based on the search formula, including:
[0051] The search formula construction module receives keywords and search conditions input by the user, automatically generates a search formula through PubMed's MeSH interface and natural language processing model, and searches for corresponding documents in the medical database based on the search formula;
[0052] The search formula construction module includes an optimization unit that dynamically adjusts natural language processing model parameters based on user search history and feedback. For example, if a user's historical search commands contain a large number of free words, the basic model's temperature value will be increased; otherwise, the temperature value will be automatically lowered. If user feedback indicates an excessive number of free words, the basic model's temperature value will be automatically lowered; otherwise, the temperature value will be automatically raised.
[0053] The document screening module is used to evaluate and screen the quality of the retrieved documents, including:
[0054] By leveraging Langchain-ChatChat's knowledge base and vector database, combined with OCR and fine-tuning of deep learning models, we assess and screen the retrieved literature for quality, identifying relevant medical literature. This involves first excluding online literature based on given criteria, then integrating these into the exclusion data. The corresponding inclusion criteria and articles are then fed into the model for fine-tuning. In practice, OCR is used to identify tables and text in the papers, which are then fed into the fine-tuned model to determine inclusion or exclusion based on the given criteria. Quality assessment is then provided through prompts with pre-set scales such as NOS, providing feedback on the model's score for each item and outputting it in JSON format.
[0055] The scale evaluation module is used to read and extract key data from the screened literature and perform scale evaluation, including:
[0056] Through table OCR and semantic recognition, based on chatpdf and chatdoc interfaces, key data from the screened literature is automatically read and extracted for scale evaluation;
[0057] The scale evaluation module includes a validation unit for cross-validating extracted key data. This includes simultaneously evaluating the scale using fine-tuned models based on multiple different benchmark models trained in clinical, basic, and bioinformatics research, and adopting a majority-first approach to achieve results that are recognized by more models.
[0058] The data extraction and information aggregation module is used to aggregate the extracted information and generate a report, including:
[0059] Summarize the extracted information, format it using regular expressions, and generate automated reports with charts and multi-style themes using Python's Matplotlib and Python-docx.
[0060] The data extraction and information aggregation module includes a custom report generation unit, which is used to generate corresponding reports based on the report style, chart type and content layout selected by the user.
[0061] The system maintenance and update module also includes a collaboration platform for multi-user online collaborative editing, annotation and sharing of reports.
[0062] A method for automatically generating a review based on artificial intelligence, using a system for automatically generating a review based on artificial intelligence, comprises the following steps:
[0063] S1, search formula construction, through the multi-task design mode of PubMed's MeSH interface and the language output of the NLP model, AI automatically constructs the search formula and formulates the inclusion and exclusion criteria;
[0064] S2, literature screening, manually created dataset screening software was used to fine-tune the 7BBaichuan language model to filter literature types by abstract;
[0065] S3, scale evaluation, uses Langchain.Chatchat's knowledge base and vector database technology, OCR technology, and deep learning to fine-tune a large language model for literature screening;
[0066] S4, data extraction, using table OCR technology and language recognition technology on the basis of ChatPDF and ChatDoc interfaces to condense and extract data from the selected articles;
[0067] S5, information summary, summarizes the extracted literature data information, uses regular formatting to apply Python's Matplotlib and Python-docxE operation technology, and outputs charts and multi-style theme automated PDF reports.
[0068] Specifically, the present invention uses an artificial intelligence-based system to accurately screen and analyze medical literature. The following is its specific working principle:
[0069] Search query construction: The system automatically generates search queries using PubMed's MeSH interface and natural language processing (NLP) models. This process improves search efficiency and accuracy through AI's multi-task design model.
[0070] Literature screening: The system uses Langchain-ChatChat's knowledge base and vector database technology, combined with OCR technology and deep learning model fine-tuning, to automatically screen out high-quality literature. AI can understand and filter out irrelevant or low-quality literature.
[0071] Scale evaluation: Through table OCR technology and semantic recognition technology, the system can automatically read and extract key data from documents based on the chatpdf and chatdoc interfaces.
[0072] Data extraction and information aggregation: The system aggregates the extracted information, formats it using regular expressions, and generates charts and automated PDF reports with multiple themes using Python's Matplotlib and Python-docx technologies.
[0073] Output and application: Ultimately, the reports generated by the system can be used for scientific research, teaching, and clinical applications, helping researchers quickly obtain high-quality information, promoting interdisciplinary collaboration, improving scientific research efficiency, and promoting the clinical transformation of scientific research results.
[0074] Through these principles, the system realizes the automation of the entire process from literature retrieval to data analysis, greatly improving the efficiency and accuracy of literature processing.
[0075] Working process:
[0076] Step 1: Search query construction
[0077] The system receives keywords and search conditions entered by the user
[0078] The system uses PubMed's MeSH interface and NLP model to automatically generate search terms
[0079] The system searches for relevant literature in databases such as PubMed based on the search formula
[0080] Step 2: Literature screening
[0081] The system uses Langchain-ChatChat's knowledge base and vector database technology to screen the retrieved documents
[0082] The system combines OCR technology and deep learning model fine-tuning to assess the quality of documents
[0083] The system selects high-quality literature based on the evaluation results
[0084] Step 3: Scale Evaluation
[0085] The system uses table OCR technology and semantic recognition technology to evaluate the selected documents.
[0086] The system extracts key data from the literature based on the scale evaluation results
[0087] Step 4: Data extraction and information aggregation
[0088] The system summarizes the extracted information
[0089] The system uses regular expressions for formatting
[0090] The system generates charts and automated PDF reports with multiple themes using Python's Matplotlib and Python-docx technologies.
[0091] Step 5: Output and Application
[0092] System output generated reports
[0093] The system provides reports in different formats (such as PDF, Word, etc.) according to user needs
[0094] The system supports users to browse and download reports online
[0095] Step 6: System maintenance and updates
[0096] The system regularly updates literature information in databases such as PubMed
[0097] The system is maintained and updated based on user feedback and evaluation results
[0098] Through these steps, the system realizes the automation of the entire process from literature retrieval to data analysis, greatly improving the efficiency and accuracy of literature processing.
[0099] Examples, such as Figure 2 As shown:
[0100] Step 01: Search formula construction
[0101] Objective: To automatically construct search queries based on given topics and formulate inclusion and exclusion criteria.
[0102] Technology: Utilizes PubMed's MeSH interface and natural language processing (NLP) models.
[0103] Operation: AI automatically generates search expressions based on preset tasks and instructions.
[0104] Step 02: Literature Screening
[0105] Objective: To determine whether the literature should be included in the study through abstract screening.
[0106] Technology: Dataset screening software based on the 7B Baichuan language model.
[0107] Operation: Conduct preliminary screening of a large number of documents to determine whether they meet the research requirements.
[0108] Step 03: Scale Evaluation
[0109] Objective: To evaluate the quality of the literature.
[0110] Technology: Uses Langchain-ChatChat’s knowledge base, vector database technology, and OCR technology.
[0111] Operation: Fine-tune the large language model through deep learning to achieve automatic evaluation of document quality.
[0112] Step 4: Data Extraction
[0113] Objective: To extract the required data from the selected literature.
[0114] Technology: Combining table OCR technology and semantic recognition technology.
[0115] Operation: Using the chatpdf and chatdoc interfaces, AI understands the article content and extracts key data.
[0116] Step 05: Information Summary
[0117] Purpose: To summarize and format the extracted data.
[0118] Technique: Use regular expressions for formatting, combined with Python's Matplotlib and Python-docx libraries.
[0119] Operation: Generate automated PDF summary reports with charts and multiple themes.
[0120] Specific implementation cases
[0121] Conduct a literature review on “Applications of Artificial Intelligence in Medicine”.
[0122] 1. Construct search formula:
[0123] Use AI tools to enter keywords such as "artificial intelligence", "oralmedicine", etc., and AI will automatically generate PubMed search terms.
[0124] 2. Literature screening:
[0125] Run the generated search formula in the database to obtain a large number of documents.
[0126] Use AI screening software to filter out literature related to the topic based on the abstract content.
[0127] 3. Scale evaluation:
[0128] The quality of the screened literature is evaluated, and the AI tool automatically scores it according to the preset evaluation criteria.
[0129] 4. Data extraction:
[0130] For literature evaluated as high-quality, AI tools are used to extract key data, such as research methods and results.
[0131] 5. Information Summary:
[0132] The extracted data is organized, and the AI tool generates charts based on the data characteristics.
[0133] Use the Python-docx library to generate a formatted PDF report containing the main content and conclusions of the literature review.
Claims
1. A system for automatically generating reviews based on artificial intelligence, characterized in that: include: A search formula construction module, a literature screening module, a scale evaluation module, a data extraction and information aggregation module, an output and application module, and a system maintenance and update module; the search formula construction module, the literature screening module, the scale evaluation module, the data extraction and information aggregation module, the output and application module, and the system maintenance and update module are connected in sequence; The search formula construction module is used to receive keywords and search conditions input by the user, generate a search formula, and search for relevant documents based on the search formula. The search formula construction module includes an optimization unit for dynamically adjusting the natural language processing model parameters based on the user's historical search records and feedback. If the user's formal search command in the history contains a large number of free words, the temperature value of the basic model is increased; otherwise, the temperature value is automatically decreased. If the user feedback indicates that there are too many free words, the temperature value of the basic model is automatically decreased; otherwise, the temperature value of the basic model is automatically increased. The literature screening module is used to evaluate and screen the retrieved literature. Through the knowledge base and vector database of Langchain-ChatChat, combined with the fine-tuning of OCR and deep learning models, the retrieved literature is quality evaluated and screened to select the corresponding medical literature. First, the online literature is excluded online under given criteria, and these are included in the exclusion data. The corresponding inclusion criteria and articles are given to the model for fine-tuning. In the actual environment, the paper is subjected to OCR to identify tables and text information. This information is given to the fine-tuned model to determine whether to include or exclude it under the given criteria. The quality assessment is feedback through prompts preset NOS and other scales. The score of each model is output in JSON format. The scale evaluation module is used to read and extract key data from the screened literature and conduct scale evaluation. It uses table OCR and semantic recognition based on the chatpdf and chatdoc interfaces to read and extract key data from the screened literature. The key data includes the literature title, abstract, publication date, author, literature research method, research sample size, research population characteristics, research method corresponding dependent variable indicator information, and the effect size of the research results. The scale evaluation module includes a validation unit for cross-validating extracted key data. This includes: Simultaneous scale evaluation using fine-tuned models based on multiple different benchmark models trained in clinical, basic, and bioinformatics research, with a majority-first approach to achieve results recognized by more models. The data extraction and information summary module is used to summarize the extracted information and generate reports; Output and apply module output generated reports; The system maintenance and update module is used to perform system maintenance and updates based on user feedback and system evaluation results.
2. The system for automatically generating a review based on artificial intelligence according to claim 1, characterized in that: The search formula building module is used to receive keywords and search conditions input by the user, generate a search formula, and search for relevant documents based on the search formula, including: The search formula construction module receives keywords and search conditions input by the user, automatically generates a search formula through PubMed's MeSH interface and natural language processing model, and searches for corresponding documents in the medical database based on the search formula.
3. The system for automatically generating a review based on artificial intelligence according to claim 1, characterized in that: The data extraction and information aggregation module is used to aggregate the extracted information and generate a report, including: The extracted effect sizes of each research result are summarized, along with the groups and specific scopes of application of the effect sizes. The summarized information is formatted using regular expressions, and an automated report with charts and multi-style themes is generated using Python's Matplotlib and Python-docx. The data extraction and information aggregation module includes a custom report generation unit, which is used to generate corresponding reports based on the report style, chart type and content layout selected by the user.
4. The system for automatically generating a review based on artificial intelligence according to claim 3, characterized in that: The system maintenance and update module also includes a collaboration platform for multi-user online collaborative editing, annotation and sharing of reports.
5. A method for automatically generating a review based on artificial intelligence, characterized in that: A system for automatically generating a review based on artificial intelligence according to any one of claims 1 to 4, comprising: S1, search formula construction, through the multi-task design mode of PubMed's MeSH interface and the language output of the NLP model, AI automatically constructs the search formula and formulates the inclusion and exclusion criteria; S2, literature screening, manually created dataset screening software was used to fine-tune the 7BBaichuan language model to filter literature types by abstract; S3, scale evaluation, uses Langchain.Chatchat's knowledge base and vector database technology, OCR technology, and deep learning to fine-tune a large language model for literature screening; S4, data extraction, using table OCR technology and language recognition technology on the basis of ChatPDF and ChatDoc interfaces to condense and extract data from the selected articles; S5, information summary, summarizes the extracted literature data information, uses regular formatting to apply Python's Matplotlib and Python-docxE operation technology, and outputs charts and multi-style theme automated PDF reports.
Citation Information
Patent Citations
Variable auto-encoder balanced Hash remote sensing image retrieval method based on multi-channel feature fusion
CN114090813A
Data mining and analysis method based on LLM large model
CN118210914A
Assistant decision-making system and method for water transportation engineering project management based on intelligent agent
CN119107047A