Medical scientific research auxiliary method and system based on large model, terminal and medium
Through large-scale auxiliary medical research methods, natural language processing and knowledge graphs are used to improve the efficiency of literature screening and experimental design, and a scientific paper framework is generated, which solves the inefficiency and complexity of traditional medical research and achieves the improvement of scientific research efficiency and quality.
Patent Information
- Application Number
- CN202510255686.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-25
AI Technical Summary
Traditional medical research faces the problems of inefficient literature screening, complex experimental design and heavy writing. The existing large-model agent auxiliary tools lack systematicity and intelligence, making it difficult to improve scientific research efficiency and quality.
A large-model-based medical research auxiliary method is used to obtain research topics through natural language processing, combine TF-IDF and text classification to generate literature reviews, use medical knowledge graphs to determine experimental types, select appropriate analysis methods, and generate a preliminary paper framework based on standard structural templates.
It improves the efficiency of literature acquisition and processing, ensures scientific and reasonable experimental design, reduces writing time, improves the efficiency and quality of medical research, and enhances systematicity and intelligence.
Smart Images

Figure CN120371987A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical data analysis and processing, and particularly relates to a medical research assistance method, system, terminal and medium based on a large model. Background Art
[0002] Medical research is a key force in promoting medical progress and improving medical standards. However, traditional medical research faces many challenges. On the one hand, the medical knowledge system is huge and complex, and new research results emerge continuously. It is difficult for researchers to comprehensively and timely master the latest information. When conducting literature research, a large amount of time and energy are required to screen a vast amount of literature, resulting in low efficiency and easy omission of important information. On the other hand, the design and implementation of medical experiments involve many professional knowledge and technologies. From the formulation of experimental plans, the collection and processing of samples to the selection of data analysis methods, high requirements are imposed on the professional abilities of researchers. In addition, the writing of medical research results also needs to follow strict academic norms, which is also a heavy task for researchers.
[0003] With the development of artificial intelligence technology, large model agents have demonstrated powerful capabilities in natural language processing, knowledge reasoning, etc. However, at present, the systems applying large model agents to the field of medical research are not yet mature enough to meet the complex needs of medical research. Some existing auxiliary tools have single functions, lack systematicness and intelligence, and are difficult to effectively improve the efficiency and quality of medical research. Therefore, it is of great practical significance to develop an efficient and intelligent medical research assistance system based on large model agents. Summary of the Invention
[0004] In view of the above deficiencies of the prior art, the present invention provides a method, system, terminal and medium to solve the above technical problems.
[0005] In a first aspect, the present invention provides a medical research assistance method based on a large model, including: S1, obtaining a medical research topic, extracting the topic meaning based on natural language processing technology, and sending a retrieval request to multiple medical literature databases based on the topic meaning and returning the retrieved relevant literature; S2, extracting features from the relevant literature based on the TF-IDF algorithm, classifying the relevant literature based on the extracted features in combination with a text classification algorithm, and generating a literature review according to the classified relevant literature; S3, obtaining a medical research question and an expected goal, determining an experiment type based on the relevant literature and a medical knowledge graph in combination with the research question and the expected goal, determining an experiment process and input requirements based on the experiment type, obtaining accurate search conditions based on natural language processing technology in combination with the input requirements, and obtaining corresponding research data in a case database based on the accurate search conditions; S4. Determine the type of research data, select the corresponding analysis method based on the type of research data for analysis and processing, and obtain the analysis results; S5. Obtain the standard structure template of the medical paper, and generate the preliminary framework of the paper based on the research data, analysis results, and literature review in combination with the standard structure template.
[0006] In an optional implementation manner, step S1 specifically includes: Preprocess the technical theme; Load the pre-trained BERT model, input the preprocessed theme text into the model. The model, based on the language knowledge and semantic representations it learned on a large-scale corpus, performs deep encoding on the theme text through a multi-layer Transformer structure, captures the context-dependent relationships between words, and outputs the understood theme semantics; Convert the theme semantics into a structured query statement recognizable by the database as a retrieval request; After receiving the retrieval request, the database performs a retrieval operation and returns the retrieved relevant literature.
[0007] In an optional implementation manner, step S2 specifically includes: Split the relevant literature into multiple words, calculate the word frequency and inverse document frequency of each word, multiply the word frequency and inverse document frequency of each word to obtain the TF-IDF value, and extract the features of the relevant literature based on the comparison of the TF-IDF value with a preset threshold; Annotate the degree of association of the relevant literature based on the features of the relevant literature in combination with disease type relevance, research method relevance, and research topic relevance; Train the SVM model based on the annotated literature and the corresponding features; Input the unannotated literature and the corresponding features into the SVM model to obtain the types corresponding to the unannotated literature; Extract key information based on the different types corresponding to the literature, and generate a literature review according to the preset review template based on the type and key information of the literature.
[0008] In an optional implementation manner, step S3 specifically includes: Analyze the application frequency of different experimental types in the relevant literature; Extract the keywords in the research questions and expected goals, match the keywords with the nodes in the knowledge graph, and obtain the recommended experimental types associated with them; Based on the analysis results of the relevant literature and the recommendations of the knowledge graph, determine the optimal experimental type.
[0009] In an optional implementation manner, in step S4, determining the type of research data specifically includes: Determine whether the research data is a number and has numerical significance. If so, it is determined as numerical data. Further, if the data can take any value within a certain range, it is determined as continuous numerical data. If the data can only take specific discrete values, it is determined as discrete numerical data; Determine whether the research data is a category label. If so, it is determined as categorical data. Further, if there is a relationship between the categorical data, it is determined as ordinal categorical. If there is no relationship between the categorical data, it is determined as nominal categorical; Determine whether the research data is composed of text. If so, it is determined as text data; Determine whether the research data has a time sequence. If so, it is determined as time series data.
[0010] In an alternative embodiment, selecting a corresponding analysis method based on the type of research data for analysis and obtaining an analysis result specifically includes: For continuous numerical data, use analysis of variance method or linear regression analysis method to obtain the analysis result between variables; For discrete data and nominal categorical, both use chi-square test method or logistic regression method to obtain the analysis result between variables; For ordinal categorical, use rank sum test to analyze the rank difference between variables as the analysis result; For text data, classify the text based on machine learning algorithms to obtain the analysis result between variables; For time series data, based on the autoregressive integrated moving average model, model and predict the time series with stationarity or after being processed by differencing to obtain the analysis result between variables.
[0011] In an alternative embodiment, the preliminary framework includes a plurality of preset parts. After step S5, it further includes: Extract the number of parts and keywords included in the target literature, match the number of parts and keywords of the target literature with the preliminary framework. When the match is unsuccessful, adjust the format of the preliminary framework according to the format requirements of the target literature, and proofread the content of the preliminary framework based on a grammar checking tool.
[0012] In a second aspect, the present invention provides a medical research assistance system based on a large model, including: When the system executes, it implements the above-mentioned medical research assistance method based on a large model, including: A literature retrieval module, which obtains a medical research topic, extracts the topic meaning based on natural language processing technology, and sends a retrieval request to multiple medical literature databases based on the topic meaning and returns the retrieved relevant literature; A review generation module extracts features from relevant literature based on the TF-IDF algorithm, classifies the relevant literature based on the extracted features in combination with a text classification algorithm, and generates a literature review according to the classified relevant literature; A data generation module obtains a medical research question and an expected goal, determines an experiment type based on the relevant literature and a medical knowledge graph in combination with the research question and the expected goal, determines an experiment process and input requirements based on the experiment type, obtains accurate search conditions based on the input requirements in combination with natural language processing technology, and obtains corresponding research data in a case database based on the accurate search conditions; An analysis and acquisition module determines the type of research data, selects a corresponding analysis method for analysis and processing based on the type of research data, and obtains an analysis result; A framework generation module obtains a standard structure template of a medical paper, and generates a preliminary framework of the paper based on the research data, the analysis result, and the literature review in combination with the standard structure template.
[0013] In a third aspect, a terminal is provided, including: A processor and a memory, wherein, The memory is used to store a computer program, The processor is used to call and run the computer program from the memory, so that the terminal executes the method of the above terminal.
[0014] In a fourth aspect, a computer storage medium is provided. Instructions are stored in the computer-readable storage medium, and when it runs on a computer, the computer is made to execute the methods described in the above aspects.
[0015] The beneficial effects of the present invention are as follows. The medical research assistance method, system, terminal and medium based on a large model provided by the present invention use technologies such as natural language processing to efficiently obtain relevant literature on medical research topics and generate a review, providing a comprehensive knowledge base for scientific research; combining literature and knowledge graphs to accurately determine experiment types, processes and data requirements, obtaining research data and selecting appropriate methods for analysis to ensure the scientific and effective nature of the experiment; generating a preliminary framework according to the standard structure template of medical papers and existing results, improving the efficiency and quality of paper writing, accelerating the medical scientific research process as a whole, effectively improving the efficiency and quality of medical scientific research, reducing costs, and enhancing the intelligence and systematicness of scientific research.
[0016] In addition, the design principle of the present invention is reliable, the structure is simple, and it has a very wide application prospect. Description of the Drawings
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 is a schematic flowchart of a large model-based medical research assistance method according to an embodiment of the present invention.
[0019] Figure 2 is a schematic block diagram of a large model-based medical research assistance system according to an embodiment of the present invention.
[0020] Figure 3 is a schematic structural diagram of a terminal provided by an embodiment of the present invention. Detailed implementation manners
[0021] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0023] The following explains the key terms that appear in the present invention.
[0024] The large model-based medical research assistance method provided by the embodiments of the present invention is executed by a computer device. Correspondingly, the large model-based medical research assistance system runs in the computer device.
[0025] Figure 1 is a schematic flowchart of the method according to an embodiment of the present invention. Among them, Figure 1 The execution subject can be a large model-based medical research assistance system. According to different requirements, the order of the steps in this flowchart can be changed, and some can be omitted.
[0026] As Figure 1 shown, this method includes: S1. Obtain a medical research topic, extract the topic meaning based on natural language processing technology, send a retrieval request to multiple medical literature databases based on the topic meaning, and return the retrieved relevant literature. Researchers input a medical research topic in the system interface. The system uses lexical, syntactic, and semantic analysis in natural language processing technology to understand the topic meaning. It converts the topic into a keyword combination that conforms to the retrieval syntax of each medical literature database, sends a retrieval request to multiple medical literature databases through the database interface, and receives and organizes the returned relevant literature.
[0027] It changes the cumbersome way of manual literature retrieval by researchers, improves the literature acquisition efficiency, can quickly obtain a large number of relevant literatures, avoids possible omissions in manual retrieval, and helps researchers comprehensively understand the existing achievements in the research field.
[0028] S2. Extract features from the relevant literature based on the TF-IDF algorithm, classify the relevant literature based on the extracted features combined with text classification algorithms, and generate a literature review according to the classified relevant literature. For the relevant literature retrieved in S1, use the TF-IDF algorithm to calculate the word frequency and inverse document frequency of the words in each literature, and extract the words and corresponding weights that can represent the literature features. Select a suitable text classification algorithm, such as a classifier based on support vector machines or convolutional neural networks, input the extracted features into the model for training and classification. According to the classification results, extract key information to generate a literature review according to a certain logic, such as the relevance of the research topic, chronological order, etc.
[0029] It automatically screens and classifies a large number of literatures, reduces the workload of researchers reading irrelevant literatures, and improves the literature processing efficiency. The generated literature review has clear logic and can help researchers quickly grasp the key points and trends in the research field.
[0030] S3. Obtain a medical research question and expected goal, determine the experiment type based on the relevant literature and medical knowledge graph combined with the research question and expected goal, determine the experiment process and input requirements based on the experiment type, obtain precise search conditions based on natural language processing technology combined with the input requirements, and obtain the corresponding research data in the case database based on the precise search conditions. Researchers input a medical research question and expected goal. The system integrates the relevant literature obtained in S1 and the information in the medical knowledge graph, analyzes the characteristics of the research question and goal, and recommends a suitable experiment type. According to the experiment type, determine the specific experiment process and the required input requirements. Researchers input relevant requirements. The system uses natural language processing technology to analyze the requirements, converts them into precise search conditions, and queries and obtains the corresponding research data in the case database.
[0031] Provide scientific and reasonable guidance on experimental design for researchers to avoid research biases caused by inappropriate selection of experimental types.
[0032] S4. Determine the type of research data, select the corresponding analysis method based on the type of research data for analysis and processing, and obtain the analysis results. For the research data obtained in S3, first judge its type. If it is numerical data, further analyze the distribution characteristics of the data, etc. According to the data type and research purpose, select a suitable method from the preset analysis method library. For example, when the numerical data conforms to the normal distribution, select the parametric test method, and then use the selected method to process the data to obtain the analysis results.
[0033] Ensure that the data analysis method matches the data characteristics and can deeply explore the information behind the data.
[0034] S5. Obtain the standard structure template of medical papers, and generate the preliminary framework of the paper based on the research data, analysis results, and literature review in combination with the standard structure template.
[0035] The system has a built-in standard structure template of medical papers. Researchers input the research data obtained in S3, the analysis results obtained in S4, and the literature review generated in S2 into the system. The system fills the relevant content into the corresponding paper sections according to the template requirements. For example, put the research data into the results section and the analysis results into the discussion section to generate the preliminary framework of the paper.
[0036] Provide a clear structural framework for writing medical papers, standardize the paper writing process, and reduce the time and energy consumption of researchers in constructing the paper structure. It helps to ensure the completeness and logical coherence of the paper content.
[0037] Optionally, as an embodiment of the present invention, step S1 specifically includes: Preprocess the technical theme. Load the pre-trained BERT model, input the preprocessed theme text into the model. The model, based on the language knowledge and semantic representations learned on a large-scale corpus, deeply encodes the theme text through a multi-layer Transformer structure, captures the context-dependent relationships between words, and outputs the understood theme semantics. Convert the theme semantics into a structured query statement recognizable by the database as a retrieval request. After receiving the retrieval request, the database performs a retrieval operation and returns the retrieved relevant literature.
[0038] Optionally, as an embodiment of the present invention, step S2 specifically includes: Split the relevant literature into multiple words, calculate the word frequency and inverse document frequency of each word, multiply the word frequency and inverse document frequency of each word to obtain the TF-IDF value, and extract the features of the relevant literature based on the comparison between the TF-IDF value and the preset threshold; Annotate the degree of association of the relevant literature based on the features of the relevant literature in combination with the disease type relevance, research method relevance, and research topic relevance; Train an SVM model based on the annotated literature and the corresponding features; Input the unannotated literature and the corresponding features into the SVM model to obtain the type corresponding to the unannotated literature; Extract key information based on the different types corresponding to the literature, and generate a literature review according to the preset review template based on the type and key information corresponding to the literature.
[0039] Optionally, as an embodiment of the present invention, the generation of the literature review is specifically as follows: When the system is started, the literature review agent module loads a pre-trained natural language processing model (such as the BERT model based on the Transformer architecture), and imports the professional vocabulary and knowledge graph in the medical field. These resources are used to improve the agent's ability to understand and analyze medical literature.
[0040] After the scientific research personnel input the research topic in the system interface, the agent preprocesses the topic text and converts it into a format suitable for database retrieval. For example, convert the natural language topic into a combination of keywords and adjust it according to the retrieval syntax of different databases. Then, send a retrieval request to the medical literature database through the database interface to obtain the index information of the relevant literature (such as title, author, abstract, published journal, etc.).
[0041] The agent makes a preliminary screening of the retrieved literature index information. Using a text classification model (such as an SVM based on support vector machines), judge the relevance of the literature to the research topic based on the title and abstract content of the literature. The literature with high relevance enters the next step of processing, and the literature with low relevance is excluded. For the screened literature, the agent further uses a deep learning model (such as a text classifier based on Transformer) for detailed classification, such as classification according to research methods (clinical trials, basic research, etc.) and disease fields (cardiovascular diseases, tumors, etc.).
[0042] The agent extracts key information from the classified literature. It uses named entity recognition (NER) technology to identify medical entities in the literature (such as disease names, drug names, gene names, etc.), and uses information extraction algorithms to extract key contents such as research purposes, methods, results, and conclusions. Finally, based on the classification and key information of the literature, the agent generates a literature review report according to specific templates and logical structures. For example, organize and summarize the literature in chronological order or according to the importance of research hotspots, generate a literature review report with clear structure and comprehensive content, and display it on the system interface for scientific researchers.
[0043] Optionally, as an embodiment of the present invention, step S3 specifically includes: The data analysis agent module receives various types of medical data (such as structured clinical data, gene sequencing data, etc.) from the experimental execution phase or the case database. For clinical data, the agent first performs data cleaning to remove duplicate records and correct incorrect data (such as incorrect date formats, unreasonable indicator values, etc.). Then, perform standardization processing on the data to convert data with different units or magnitudes into a unified standard form (such as unifying blood pressure values to mmHg units). For gene sequencing data, the agent performs quality control, removes low-quality sequencing reads, and normalizes gene expression data to eliminate experimental errors and batch effects.
[0044] The agent automatically selects appropriate analysis methods according to the type of data, research purpose, and data characteristics. For example, for between-group comparisons in clinical data (such as comparing the treatment effect differences between the experimental group and the control group), the agent selects statistical methods such as t-tests and analysis of variance; for exploring the relationships between multiple variables, it selects regression analysis methods (such as linear regression, logistic regression, etc.). For gene data, if the research purpose is to find genes related to diseases, the agent selects methods such as gene set enrichment analysis (GSEA) and genome-wide association studies (GWAS).
[0045] The agent converts the analysis results into intuitive visual charts. For example, bar charts are used to display the comparison results of between-group data, line charts are used to show the change trend of data over time, heat maps are used to present the differences in gene expression, and 3D reconstruction images are used to display the three-dimensional structure of medical images, etc. At the same time, the agent uses natural language generation technology to interpret the analysis results. Specifically: analyze the application frequency of different experimental types in relevant literature; extract keywords in research questions and expected goals, match the keywords with the nodes in the knowledge graph, and obtain recommended experimental types associated with them; comprehensively determine the optimal experimental type based on the analysis results of relevant literature and the recommendations of the knowledge graph.
[0046] For example, if the research question is to explore the therapeutic effect of a new drug on lung cancer patients, the agent may recommend a randomized controlled trial. After determining the type of experiment, the agent assists the researchers in designing the experimental protocol. For a randomized controlled trial, the agent helps calculate the sample size, considering factors such as the incidence of the disease, the difference in expected treatment effects, and the power of the test, and uses the corresponding sample size calculation formula (such as the sample size calculation formula based on rates) to obtain the required sample size. The agent also plans the experimental process, including how to group (such as randomly dividing into an experimental group and a control group), the implementation method of the intervention measures (such as using the new drug in the experimental group and using a placebo or traditional treatment drug in the control group), the time points and indicators for data collection, etc.
[0047] Natural language case search and inclusion / exclusion criteria conversion: The researchers input the case screening requirements in natural language in the system, such as "search for patient cases who have had lung cancer and received targeted therapy in the past three years, aged between 40 - 70 years old, without severe liver and kidney dysfunction". The experimental design and execution agent module first uses natural language processing technology to parse the input content. Through lexical analysis, syntactic analysis, and semantic understanding, it identifies medical entities (such as "lung cancer", "targeted therapy", "age", "liver and kidney function dysfunction", etc.) and semantic relationships (such as "has", "received", "without", etc.). Then, the agent converts the natural language into search conditions recognizable by the database. Assuming the case database uses MySQL, the search statement generated by the agent may be: "SELECT * FROM patient_records WHERE disease = 'lung cancer' AND treatment LIKE '%targeted therapy%' AND age BETWEEN 40 AND 70 AND (liver_function NOT LIKE '%severe dysfunction%' OR kidney_function NOT LIKE '%severe dysfunction%') AND diagnosis_year BETWEEN '2020' AND '2023'". The agent executes the search operation, filters out the eligible case data from the case database, and displays the results to the researchers, while generating a detailed inclusion and exclusion criteria report.
[0048] Optionally, as an embodiment of the present invention, selecting a corresponding analysis method for analysis and processing based on the type of research data and obtaining the analysis result specifically includes: For continuous numerical data, the analysis result between variables is obtained by using the analysis of variance method or the linear regression analysis method, where: a. Ensure that the continuous numerical variables in the dataset meet the basic assumptions of analysis of variance, namely independence (each observation is independent of each other), normality (each group of data follows a normal distribution), and homogeneity of variance (the variances of each group of data are equal). These assumptions can be verified by plotting histograms, performing normality tests (such as the Shapiro-Wilk test) and homogeneity of variance tests (such as the Levene test).
[0049] Identify the independent variables (factors) in the study and their different values (levels). For example, when studying the effect of different drug treatment regimens (factor) on the blood glucose levels of patients (continuous numerical dependent variable), the different drug treatment regimens are the different levels of the factor.
[0050] Perform analysis of variance using statistical analysis software (such as SPSS, R, or relevant libraries in Python). In the software, specify the dependent variable as the continuous numerical variable and the independent variable as the factor variable. The results of the analysis of variance mainly focus on the F-value and the corresponding p-value. If the p-value is less than the pre-set significance level (usually 0.05), it indicates that the factor at different levels has a significant effect on the dependent variable. Further, multiple comparisons (such as the Tukey test) can be used to determine which specific levels are significantly different from each other.
[0051] b. Check whether there is a linear relationship between the independent variable and the dependent variable. A preliminary observation can be made by plotting a scatter plot. At the same time, ensure that the data meets the basic assumptions of linear regression, including linear relationship, independence, normality, and homogeneity of variance (similar to analysis of variance), as well as no multicollinearity (if there are multiple independent variables). For multicollinearity, it can be tested by calculating the variance inflation factor (VIF). If the VIF value is greater than 10, there may be a serious multicollinearity problem.
[0052] Select appropriate independent and dependent variables, and use statistical analysis software to construct a linear regression model. Check the summary output of the regression model, focusing on the coefficient estimates, standard errors, t-values, and p-values. The coefficient estimate represents the average change in the dependent variable for each one-unit change in the independent variable. A p-value less than the significance level indicates that the independent variable has a significant effect on the dependent variable. At the same time, the goodness of fit of the model can be evaluated by the R-squared value. The closer the R-squared is to 1, the better the model fits the data.
[0053] For discrete data and unordered classifications, both the chi-square test method or the logistic regression method are used to obtain the analysis results between variables; a. Organize discrete data or unordered categorical data into the form of a contingency table. For example, to study the relationship between gender (male, female) and a certain disease (diseased, non-diseased), a 2×2 contingency table can be constructed, with rows representing gender and columns representing disease status.
[0054] The null hypothesis is that there is no association between the two variables, and the alternative hypothesis is that there is an association between the two variables.
[0055] Use statistical analysis software to calculate the chi-square value and the corresponding p-value. In SPSS, the chi-square test can be performed through the "Crosstabs" function; If the p-value is less than the significance level (usually 0.05), the null hypothesis is rejected, and it is considered that there is a significant association between the two variables.
[0056] b. Convert the discrete dependent variable into a binary variable (if it is not originally binary), and use statistical analysis software or machine learning libraries to construct and train a logistic regression model. Evaluate the performance of the model by calculating indicators such as accuracy, recall, and F1 value. Check the coefficients of the model. The coefficient represents the change in the log odds ratio of the dependent variable taking the value of 1 for every one-unit change in the independent variable. For example, if the independent variable is a certain lifestyle habit (coded as 0 and 1), a positive coefficient indicates an increase in the log odds ratio of this lifestyle habit being associated with the disease (the dependent variable being 1 indicates having the disease), that is, this lifestyle habit may increase the risk of getting the disease.
[0057] For ordinal classification, use the rank sum test to analyze the rank differences between variables as the analysis result; Ensure that the data is ordinal classification, such as the severity of the disease (mild, moderate, severe). Clearly define the two or more groups of data to be compared.
[0058] If comparing two independent samples, the Mann - Whitney U test can be used; if comparing multiple independent samples, the Kruskal - Wallis H test can be used.
[0059] Check the p-value in the test result. If the p-value is less than the significance level (usually 0.05), it indicates that there are significant differences in the rank distributions of the two or more groups of data. For example, when comparing the improvement effects of two treatment methods on the severity of patients' diseases, if the p-value is less than 0.05, it means that there are significant differences between the two treatment methods in improving the severity of the disease.
[0060] For text data, use machine learning algorithms to classify the text to obtain the analysis result between variables; Remove noise from the text, such as HTML tags (if the text is obtained from a web page), special characters, stop words (such as "的", "是", "在" and other common words with no actual meaning), and convert the text to lowercase.
[0061] Use word segmentation tools (such as NLTK, Jieba word segmentation, etc.) to split the text into individual words or phrases.
[0062] A commonly used method is the Bag-of-Words model, which represents the text as a vector where each element of the vector represents the number of times a word appears in the text; or the TF-IDF (Term Frequency - Inverse Document Frequency) method, which considers not only the word frequency but also the rarity of the word in the entire document collection.
[0063] Dataset division: The preprocessed text data is divided into training set, validation set and test set, usually in a ratio of 70%, 15%, and 15%. The training set is used to train the model, the validation set is used to adjust the model parameters, and the test set is used to evaluate the final performance of the model.
[0064] Model selection and training: Choose the appropriate machine learning algorithm, such as Naive Bayes, Support Vector Machine (SVM), Decision Tree, etc.
[0065] The performance of the model is evaluated using the test set. Common evaluation indicators include accuracy, recall, F1 value, etc. For example, the accuracy is obtained by calculating the ratio of the number of samples correctly predicted by the model on the test set to the total number of samples, and the accuracy of the model for text classification is evaluated.
[0066] For time series data, the autoregressive integrated moving average model is used to model and predict time series that are stationary or stationary after difference processing, and the analysis results between variables are obtained.
[0067] Use unit root test (such as Augmented Dickey-Fuller test, ADF test for short) to determine whether the time series data is stationary. If the p value is greater than the significance level (usually 0.05), it means that the data is non-stationary and needs to be differentiated.
[0068] Differentiate the non-stationary time series to make it stationary. The first-order difference is to calculate the difference between the data at adjacent time points, and the second-order difference is to differentiate the data after the first-order difference. Through continuous trial and error, find the appropriate difference order to make the data stationary.
[0069] Determine the three parameters (p, d, q) of the ARIMA model, where p is the autoregressive order, d is the differencing order, and q is the moving average order. The parameter values can be initially determined by observing the graphs of the autocorrelation function (ACF) and the partial autocorrelation function (PACF).
[0070] Use the trained model for prediction to predict the time series values for a period in the future. Evaluate the prediction performance of the model by calculating metrics such as the mean squared error (MSE) and the root mean squared error (RMSE). For example, RMSE measures the average error degree between the predicted values and the actual values, and the smaller the RMSE value, the better the prediction effect of the model.
[0071] Optionally, as an embodiment of the present invention, step S5 specifically includes: Researchers input research data, analysis results, and literature review content into the intelligent agent module for thesis writing. The intelligent agent generates a preliminary framework for the thesis based on the standard structure (such as the IMRAD structure: Introduction, Methods, Results, Discussion) and templates of medical theses. For example, in the Introduction section, the intelligent agent combines the literature review content to elaborate on the research background and purpose, point out the problems existing in the current research field and the innovation points of this research; in the Methods section, the intelligent agent details the experimental objects, experimental methods, data collection and analysis methods, etc. according to the process of experimental design and execution; in the Results section, the intelligent agent lists the main results obtained from data analysis, including the data values of key indicators, statistical test results, etc.; in the Discussion section, the intelligent agent combines the Introduction and Results to conduct an in-depth analysis of the research results, discuss the significance of the results, the similarities and differences with previous studies, the limitations of the research, etc.
[0072] Content filling and optimization: The intelligent agent fills the content of each part of the thesis according to the research data and analysis results. In the Results section, the intelligent agent sorts out and describes the data and charts obtained from data analysis to ensure the accuracy and clarity of the results. In the Discussion section, the intelligent agent combines the literature review to interpret and discuss the results, cites relevant research results to support its own views, and analyzes the potential impact of the research results on the medical field. Researchers review and modify the generated content, and the intelligent agent further optimizes the thesis content according to the modification opinions. For example, if researchers think that the logic of a certain paragraph is unclear, the intelligent agent optimizes the content by adjusting the sentence order, adding conjunctions, etc.; if researchers hope to supplement more references, the intelligent agent searches for and cites relevant literature based on the literature review.
[0073] Format Specification and Proofreading: After the paper content is completed, the agent adjusts the format of the paper according to the format requirements of the target academic journal. For example, for the submission format of The New England Journal of Medicine, the agent adjusts format parameters such as font (e.g., using Times New Roman), font size (e.g., 12-point for the main text and different font sizes for headings according to their levels), line spacing (e.g., 1.5 times line spacing), standardizes the citation format (e.g., adjusts the format of references according to the citation style of this journal), numbers and labels the figures and tables (e.g., "Figure 1. [Figure Title]"). The agent also proofreads the paper for grammar, spelling, and punctuation, using a language model (e.g., a fine-tuned version of the GPT series models) to check for language errors in the paper, ensuring the language quality and academic standardization of the paper. Finally, the agent generates a paper document (e.g., in PDF format) that meets the requirements of the target journal for researchers to submit.
[0074] Optionally, as an embodiment of the present invention, the preliminary framework includes a plurality of preset parts, and after step S5, it further includes: Extract the number of parts and keywords contained in the target document, match the number of parts and keywords of the target document with the preliminary framework. When the match is unsuccessful, adjust the format of the preliminary framework according to the format requirements of the target document, and proofread the content of the preliminary framework based on a grammar checking tool.
[0075] In some embodiments, the large model-based medical research assistance system may include a plurality of functional modules composed of computer program segments. The computer programs of each program segment in the large model-based medical research assistance system can be stored in the memory of the computer device and executed by at least one processor to execute (see details in Figure 1 description) the functions of large model-based medical research assistance.
[0076] In this embodiment, the large model-based medical research assistance system can be divided into a plurality of functional modules according to the functions it performs, such as Figure 2 shown. The functional modules of system 200 may include: a literature retrieval module, a review generation module, a data generation module, an analysis and acquisition module, and a framework generation module. The module referred to in the present invention means a series of computer program segments that can be executed by at least one processor and can complete a fixed function, and is stored in the memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments. The system includes: A literature retrieval module, which obtains a medical research topic, extracts the topic meaning based on natural language processing technology, and sends a retrieval request to multiple medical literature databases based on the topic meaning and returns the retrieved relevant literature; A review generation module extracts features from relevant literature based on the TF-IDF algorithm, classifies the relevant literature based on the extracted features in combination with a text classification algorithm, and generates a literature review according to the classified relevant literature. A data generation module obtains a medical research question and an expected goal, determines an experiment type based on the relevant literature and a medical knowledge graph in combination with the research question and the expected goal, determines an experiment process and input requirements based on the experiment type, obtains accurate search conditions based on natural language processing technology in combination with the input requirements, and obtains corresponding research data from a case database based on the accurate search conditions. An analysis and acquisition module determines the type of research data, selects a corresponding analysis method based on the type of research data for analysis and processing, and obtains an analysis result. A framework generation module obtains a standard structure template of a medical paper, and generates a preliminary framework of the paper based on the research data, the analysis result, and the literature review in combination with the standard structure template.
[0077] Figure 3 The figure is a schematic structural diagram of a terminal provided by an embodiment of the present invention, and the terminal can be used to execute the medical research assistance method based on a large model provided by an embodiment of the present invention.
[0078] Among them, the terminal may include: a processor, a memory, and a communication unit. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the server shown in the figure does not constitute a limitation to the present invention. It can be a bus structure, a star structure, and may also include more or fewer components than shown in the figure, or combine some components, or have different component arrangements.
[0079] Among them, the memory can be used to store the execution instructions of the processor. The memory can be implemented by any type of volatile or non-volatile storage terminal or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. When the execution instructions in the memory are executed by the processor, the terminal can execute some or all of the steps in the above method embodiments.
[0080] The processor is the control center of the storage terminal, connecting various parts of the entire electronic terminal through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory, and by calling data stored in the memory, it performs various functions of the electronic terminal and / or processes data. The processor may be composed of an integrated circuit (IC for short), for example, it may be composed of a single packaged IC, or it may be composed of multiple packaged ICs with the same or different functions connected together. For example, the processor may only include a central processing unit (CPU for short). In the embodiments of the present invention, the CPU may be a single arithmetic core or may include multiple arithmetic cores.
[0081] The communication unit is used to establish a communication channel so that the storage terminal can communicate with other terminals. It receives user data sent by other terminals or sends user data to other terminals.
[0082] The present invention also provides a computer storage medium. Among them, the computer storage medium can store a program, and when the program is executed, it may include some or all of the steps in the embodiments provided by the present invention. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM for short), a random access memory (RAM for short), etc.
[0083] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disk, etc., which can store program codes, and includes several instructions to enable a computer terminal (which may be a personal computer, a server, or a second terminal, a network terminal, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0084] For the same or similar parts among the various embodiments in this specification, reference can be made to each other. In particular, for the terminal embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the descriptions in the method embodiments.
[0085] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of systems or modules can be in electrical, mechanical or other forms.
[0086] The modules described as separate components may or may not be physically separated. The components displayed as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0087] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.
[0088] Although the present invention has been described in detail by referring to the drawings and in combination with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, those of ordinary skill in the art can make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions should all be within the scope of the present invention. / Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, and they should all be covered within the protection scope of the present invention.
Claims
1. A medical research assistance method based on a large model, characterized in that, Including the following steps: S1. Obtain a medical research topic, extract the topic meaning based on natural language processing technology, and send a retrieval request to multiple medical literature databases based on the topic meaning and return the retrieved relevant literature; S2. Extract features from the relevant literature based on the TF-IDF algorithm, classify the relevant literature based on the extracted features combined with the text classification algorithm, and generate a literature review according to the classified relevant literature; S3. Obtain a medical research question and expected goals, determine the experimental type based on the relevant literature and medical knowledge graph combined with the research question and expected goals, and determine the experimental process and input requirements based on the experimental type. Obtain accurate search conditions based on natural language processing technology combined with the input requirements, and obtain corresponding research data in the case database based on the accurate search conditions; S4. Determine the type of research data, select the corresponding analysis method based on the type of research data for analysis and processing and obtain the analysis result; S5. Obtain the standard structure template of a medical paper, and generate a preliminary framework of the paper based on the research data, analysis result and literature review combined with the standard structure template.
2. The medical research assistance method based on a large model according to claim 1, wherein, Step S1 specifically includes: Preprocess the technical topic; Load the pre-trained BERT model, input the preprocessed topic text into the model. The model, based on the language knowledge and semantic representation it learned on a large-scale corpus, deeply encodes the topic text through a multi-layer Transformer structure, captures the context dependency relationships between words, and outputs the understood topic semantics; Convert the topic semantics into a structured query statement recognizable by the database as a retrieval request; After receiving the retrieval request, the database performs a retrieval operation and returns the retrieved relevant literature.
3. The medical research assistance method based on a large model according to claim 1, wherein, Step S2 specifically includes: Split the relevant literature into multiple words, calculate the word frequency and inverse document frequency of each word, multiply the word frequency and inverse document frequency of each word to obtain the TF-IDF value, and extract the features of the relevant literature based on the comparison between the TF-IDF value and a preset threshold; Mark the degree of association of the relevant literature based on the features of the relevant literature combined with disease type relevance, research method relevance and research topic relevance; Train an SVM model based on the marked literature and corresponding features; Input the unmarked literature and corresponding features into the SVM model to obtain the type corresponding to the unmarked literature; Extract key information respectively based on different types corresponding to the literature, and generate a literature review according to the type and key information of the literature according to a preset review template.
4. The medical research assistance method based on a large model according to claim 1, wherein Step S3 specifically includes: Analyze the application frequency of different experimental types in the relevant literature; Extract keywords in the research question and expected goals, match the keywords with the nodes in the knowledge graph, and obtain recommended experimental types associated with them; Comprehensively determine the optimal experimental type based on the analysis results of the relevant literature and the recommendations of the knowledge graph.
5. The medical research assistance method based on a large model according to claim 1, wherein In step S4, determining the type of research data specifically includes: Determine whether the research data is a number and has numerical significance. If so, it is determined as numerical data. Further, if the data can take any value within a certain range, it is determined as continuous numerical data. If the data can only take discrete values, it is determined as discrete numerical data; Determine whether the research data is a categorical label. If so, it is determined as categorical data. Further, if there is a relationship between the categorical data, it is determined as ordered classification. If there is no relationship between the categorical data, it is determined as unordered classification; Determine whether the research data is composed of text. If so, it is determined as text data; Determine whether the research data has a time sequence. If so, it is determined as time series data.
6. The medical research assistance method based on a large model according to claim 1, wherein Based on the type of research data, select the corresponding analysis method for analysis and obtain the analysis results, specifically including: For continuous numerical data, use analysis of variance method or linear regression analysis method to obtain the analysis results between variables; For discrete data and unordered classification, both use chi-square test method or logistic regression method to obtain the analysis results between variables; For ordered classification, use the rank sum test to analyze the rank difference between variables as the analysis result; For text data, classify the text based on machine learning algorithms to obtain the analysis results between variables; For time series data, based on the autoregressive integrated moving average model, model and predict the time series with stationarity or after being made stationary through differencing processing to obtain the analysis results between variables.
7. The medical research assistance method based on a large model according to claim 1, wherein The preliminary framework includes multiple preset parts. After step S5, it also includes: Extract the number of parts and keywords included in the target document, match the number of parts and keywords of the target document with the preliminary framework. When the match is unsuccessful, adjust the format of the preliminary framework based on the format requirements of the target document, and proofread the content of the preliminary framework based on a grammar checking tool.
8. A medical research assistance system based on a large model, characterized in that, When the system executes, it implements the medical research assistance method based on the large model as described in any one of claims 1-7, characterized by including: A literature retrieval module, which obtains the medical research topic, extracts the topic meaning based on natural language processing technology, and sends a retrieval request to multiple medical literature databases based on the topic meaning and returns the retrieved relevant literature; A review generation module, which extracts features from the relevant literature based on the TF-IDF algorithm, classifies the relevant literature based on the text classification algorithm combined with the extracted features, and generates a literature review based on the classified relevant literature; A data generation module, which obtains the medical research question and expected goal, determines the experiment type based on the relevant literature and medical knowledge graph combined with the research question and expected goal, determines the experimental process and input requirements based on the experiment type, obtains accurate search conditions based on natural language processing technology combined with the input requirements, and obtains the corresponding research data in the case database based on the accurate search conditions; An analysis and acquisition module, which determines the type of research data, selects the corresponding analysis method based on the type of research data for analysis and obtains the analysis results; A framework generation module, which obtains the standard structure template of a medical paper, and generates a preliminary framework of the paper based on the research data, analysis results, and literature review combined with the standard structure template.
9. A terminal, characterized in that, Including: A memory, which is used to store the medical research assistance program based on the large model; A processor for implementing the steps of the large model-based medical research assistance method according to any one of claims 1-7 when executing the large model-based medical research assistance program.
10. A computer-readable storage medium storing a computer program, characterized in that, The readable storage medium stores a large model-based medical research assistance program, and when the large model-based medical research assistance program is executed by a processor, it implements the steps of the large model-based medical research assistance method according to any one of claims 1-7.
Citation Information
Cited By
Literature review generation method and system and electronic equipment
CN121118882A
A literature review generation method, system and electronic device
CN121118882B
Clinical research information batch extraction system and extraction method
CN122117190A
Medical automation scientific research method and device
CN122154632A
Data generation method and system based on social simulation experiment and electronic equipment
CN122242303A