Medical scientific research index recommendation method based on large model literature learning and related device
Through the big model literature learning method, medical research indicators are identified in multi-field scientific research literature and clinical diagnosis and treatment data, and a multi-field index database is built, which solves the problem of difficulty in providing cross-field medical research indicators in the existing technology, and achieves efficient indicator recommendation and database richness.
Patent Information
- Application Number
- CN202510118821.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-06
AI Technical Summary
The existing technology is difficult to provide cross-field medical research indicators, and it is less efficient to determine medical research indicators by manually collecting and processing data by medical researchers.
The large-model literature learning method is adopted to identify medical research indicators through a large language model for scientific research literature materials in multiple fields and clinical diagnosis and treatment data in multiple departments, and an index database containing medical research indicators in multiple fields is constructed, and indicator recommendations are made based on prompt information and index databases.
It has achieved cross-field recommendation of medical research indicators, broken the field barriers of medical experts, improved the efficiency of determining medical research indicators for medical research, and enriched the medical research indicators in the index database.
Smart Images

Figure CN120104866A_ABST
Abstract
Description
Technical Field
[0001] The present application is applied to the field of artificial intelligence technology, and in particular relates to a medical research indicator recommendation method based on large-model literature learning and related devices. Background Art
[0002] Scientific research indicators in the medical field can provide research directions for medical researchers.
[0003] In related technologies, medical researchers determine the research objectives and scope, select medical research indicators from routine clinical work indicators based on the research objectives and scope, and conduct some data analysis and research around the medical research indicators to form scientific research innovation results.
[0004] However, different medical researchers may focus on studying different medical fields. Due to the limitations of the medical fields that medical researchers focus on, the above method is difficult to provide cross-field scientific research indicators. Summary of the invention
[0005] In order to solve the above problems, this application proposes a medical research indicator recommendation method and related devices based on large-model literature learning, which can realize cross-domain scientific research indicator recommendation.
[0006] The first aspect of the present application provides a method for recommending medical research indicators based on large-model literature learning, including: obtaining prompt information related to medical research; determining target medical research indicators based on the prompt information and a pre-constructed indicator database, wherein the indicator database includes multiple first medical research indicators, and the multiple first medical research indicators are obtained by identifying medical research indicators of scientific research literature materials in multiple fields through a large language model; and outputting recommendation information, wherein the recommendation information indicates the target medical research indicator.
[0007] In some embodiments, the first medical research indicator is obtained by the following operations: acquiring the scientific research literature; parsing the scientific research literature to obtain parsed data; formatting and denoising the parsed data according to the input data requirements of the large language model to obtain target text data; and identifying the medical research indicators of the target text data through the large language model to obtain the first medical research indicator.
[0008] In some embodiments, the process of formatting the parsed data includes: converting the format of the parsed data into text data that meets the requirements of the input data by performing data normalization on the parsed data; wherein the data normalization includes at least one of the following: batch normalization, instance normalization, layer normalization, group normalization, and position normalization.
[0009] In some embodiments, the process of denoising the parsed data includes: identifying data noise in the parsed data through a data discriminator; removing data noise in the parsed data; wherein the data discriminator includes at least one of the following: a language discriminator, a quality discriminator, a privacy discriminator, and a security discriminator.
[0010] In some embodiments, the indicator database also includes multiple second medical research indicators, and the second medical research indicators are obtained by using the large language model to perform medical research indicator identification on clinical diagnosis and treatment data of multiple departments.
[0011] In some embodiments, the second medical research indicator is obtained by the following operations: acquiring the clinical diagnosis and treatment data, the clinical diagnosis and treatment data including clinical documents corresponding to multiple patients, the clinical documents including at least one of the following information: medical consultation information, physical examination information and physiological sample test information; using the large language model, performing medical research indicator identification on the clinical diagnosis and treatment data to obtain the second medical research indicator.
[0012] In some embodiments, the large language model is used to identify the medical research indicators of the clinical diagnosis and treatment data to obtain the second medical research indicator, including: performing synonym processing and image text encoding on the clinical diagnosis and treatment data through a convolutional neural network and a language-image pre-training model based on contrastive learning to obtain a feature vector representation corresponding to the clinical diagnosis and treatment data; and the large language model is used to identify the medical research indicators of the feature vector representation to obtain the second medical research indicator.
[0013] In some embodiments, the prompt information includes a disease identifier of the target disease. After determining the target medical research indicator based on the prompt information and a pre-constructed indicator database, the method further includes: obtaining clinical documents related to the target disease; identifying the indicator value of the target medical research indicator on the clinical document to obtain a special disease database for the target disease, wherein the special disease database contains the indicator value; performing data analysis on the special disease database to obtain an impact factor of the target medical research indicator, wherein the impact factor indicates the degree of impact of the target medical research indicator on the target disease; wherein the recommendation information also indicates the special disease database and the impact factor.
[0014] In some embodiments, after performing data analysis on the special disease database to obtain the influencing factor of the target medical research indicator, it also includes: adjusting the target medical research indicator according to the influencing factor; wherein the recommendation information indicates the adjusted target medical research indicator, the indicator value of the adjusted target medical research indicator and the influencing factor of the adjusted target medical research indicator.
[0015] The second aspect of the present application provides a medical research indicator recommendation device based on large model literature learning, including: an acquisition unit, used to obtain prompt information related to medical research; a processing unit, used to determine the target medical research indicator based on the prompt information and a pre-constructed indicator database, the indicator database contains multiple first medical research indicators, and the multiple first medical research indicators are obtained by identifying medical research indicators of scientific research literature materials in multiple fields through a large language model; an output unit, used to output recommendation information, the recommendation information indicates the target medical research indicator.
[0016] The third aspect of the present application provides an electronic device, including a memory and a processor; the memory is connected to the processor and is used to store programs; the processor is used to implement the medical research indicator recommendation method for large-model literature learning as described in the first aspect or any embodiment of the first aspect by running the program in the memory.
[0017] The fourth aspect of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the medical research indicator recommendation method for large-model literature learning as described in the first aspect or any embodiment of the first aspect.
[0018] In a fifth aspect, the present application provides a storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for recommending medical research indicators through large-model literature learning as described in the first aspect or any embodiment of the first aspect is implemented.
[0019] The present application proposes a method and related device for recommending medical research indicators based on large-model literature learning. Through a large language model, medical research indicators are identified for scientific research literature materials in multiple fields, and medical research indicators in multiple fields are identified. Medical research indicators are recommended based on indication information related to medical research and an indicator database containing medical research indicators in multiple fields. This can break the limitations of human cognition and realize cross-domain medical research indicator recommendations. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0021] Figure 1 A schematic diagram of an implementation environment involved in an embodiment of the present application;
[0022] Figure 2 Schematic diagram of the process of recommending medical research indicators based on large-scale model literature learning provided in the embodiment of the present application Figure 1 ;
[0023] Figure 3 A schematic flow chart of a process for determining a first medical research indicator in a method for recommending medical research indicators through large-model literature learning provided in an embodiment of the present application;
[0024] Figure 4 A schematic flow chart of a process for determining a second medical research indicator in a method for recommending medical research indicators through large-model literature learning provided in an embodiment of the present application;
[0025] Figure 5 Schematic diagram of the process of recommending medical research indicators based on large-scale model literature learning provided in the embodiment of the present application Figure 2 ;
[0026] Figure 6 A schematic diagram of the structure of a medical research indicator recommendation device based on large model literature learning provided according to an embodiment of the present application;
[0027] Figure 7 It is a schematic diagram of the structure of an electronic device provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0028] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0029] In the related technologies, the application of large language models in the medical field is mainly concentrated in various links before, during and after the consultation, including electronic medical record understanding, medical Q&A, medical education and training, disease auxiliary diagnosis, drug development, virtual hospitals and medical virtual digital human interaction. However, there is a lack of application of large language models in medical scientific research (hereinafter referred to as medical research).
[0030] In the core work of medical research, a medical research database is constructed through continuous medical experiments and data processing. In the process of constructing a medical research database, relying on the knowledge and experience of medical researchers, indicators that may affect the research goals are selected from clinical work indicators (research goals such as studying hypertension, indicators that may affect research goals such as whether the patient smokes, the patient's hemoglobin test results), and indicators that may affect the research goals are determined as medical research indicators.
[0031] However, on the one hand, there are many medical specialties, and medical researchers have field boundaries and limited energy. It is difficult to establish a multi-field (or multi-disciplinary) medical research database and to find multi-field (i.e., cross-field) related medical research indicators for medical research. On the other hand, it is inefficient to manually collect and process data to determine medical research indicators through medical researchers.
[0032] To solve the above problems, the embodiment of the present application provides a method and related devices for recommending medical research indicators based on large model literature learning, which are applied to the field of artificial intelligence technology. In the embodiment of the present application, the large language model is used to identify indicators of scientific research documents in multiple fields to obtain medical research indicators, so that the indicator database contains medical research indicators in multiple fields; when the indicator database contains medical research indicators in multiple fields, the indicator is recommended based on the indication information related to medical research and the indicator database, breaking the field barriers of medical experts, realizing the recommendation of cross-field medical research indicators, and improving the efficiency of determining medical research indicators for medical research.
[0033] Exemplary Implementation Environment
[0034] Please refer to Figure 1 , Figure 1 The implementation environment involved in the embodiment of the present application may include: a data processing device 101, an indicator recommendation device 102 and a database 103, and the database 103 may include an indicator database.
[0035] On the data processing device 101, a large language model is used to identify medical research indicators in scientific research literature in multiple fields to construct an indicator database in the database 103. The indicator recommendation device 102 can recommend medical research indicators based on the indicator database in the database 103.
[0036] The data processing device 101 and the indicator recommendation device 102 may include a terminal and / or a server. The terminal may be a personal digital assistant (PDA), a handheld device with wireless communication function (such as a smart phone, a tablet), a computing device (such as a personal computer (PC), a wearable device (such as a smart watch, a smart bracelet), and a smart home device (such as a smart speaker, a smart display device). The server may be an independent server or a server cluster, and may be a local server or a cloud server.
[0037] Exemplary Methods
[0038] See also Figure 2 In an exemplary embodiment, a method for recommending medical research indicators based on large-scale model literature learning is provided, and the method comprises the following steps:
[0039] S201, obtaining prompt information related to medical research.
[0040] The prompt information may include key information of medical research, such as the purpose and scope of medical research.
[0041] In one example, the prompt information includes a disease identifier of the target disease, such as a disease name and a disease code of the target disease.
[0042] Further, the prompt information indicates that medical research indicators are recommended for the target disease.
[0043] For example, the prompt message is "Please recommend medical research indicators related to hypertension to me."
[0044] In this embodiment, prompt information input by the user or sent by the terminal where the user is located may be obtained.
[0045] In one example, a user enters prompt information on a terminal. After the terminal obtains the prompt information entered by the user, it generates an indicator recommendation request message and sends the indicator recommendation request message to the current device. After the current device receives the indicator recommendation request message, it obtains the prompt information from the indicator recommendation request message.
[0046] S202, determining the target medical research indicator according to the prompt information and a pre-constructed indicator database, wherein the indicator database includes a plurality of first medical research indicators, and the plurality of first medical research indicators are obtained by performing medical research indicator recognition on scientific research literature materials in multiple fields through a large language model.
[0047] Among them, scientific research literature materials include public papers, journals, conference documents, patent documents, scientific research reports, books, etc. Multi-fields can include multiple medical fields, such as clinical medicine, ophthalmology, stomatology, anesthesiology, immunology, etc. Multi-fields can also include multiple non-medical fields. In the scientific research literature materials in non-medical fields, there may also be content related to diseases in medical research, which can increase the quantity and richness of scientific research literature materials.
[0048] Among them, medical research indicators can also be called medical research variables, which refer to the patient's disease-related indicators, such as the patient's blood test indicators (platelets, hemoglobin and other examination items in the blood), lifestyle indicators (such as whether to smoke), etc.
[0049] The target medical research indicator refers to the medical research indicator to be recommended. The target medical research indicator is related to the prompt information, that is, the target medical indicator is related to the medical research associated with the prompt information.
[0050] In one example, when the prompt information includes a disease identifier of a target disease, the target medical research indicator is a medical research indicator related to the target disease. For example, if the target disease is hypertension, the target medical research indicator may be the patient's blood pressure, whether he or she smokes, whether he or she is obese, etc.
[0051] In this embodiment, since the plurality of first medical research indicators in the indicator database are from scientific research literature materials in multiple fields, the plurality of first medical research indicators are medical research indicators in multiple fields. After obtaining the prompt information, the medical research indicator related to the prompt information is determined among the plurality of medical research indicators included in the indicator database, i.e., the target medical research indicator.
[0052] In one example, the prompt information can be input into a large language model, and the large language model can be used to determine the target medical research indicator from multiple medical research indicators included in the indicator database.
[0053] Among them, the big language model used to recommend medical research indicators and the big language model used to identify medical research indicators from scientific research literature can be the same model or different models.
[0054] In this example, the large language model can be used to extract features from the prompt information, and based on the extracted feature information, the target medical research indicator is determined from the multiple medical research indicators contained in the indicator database. Thus, the feature learning ability of the large language model is used to learn the relationship between diseases and indicators, and the accuracy and efficiency of selecting target medical research indicators from the indicator database are improved.
[0055] In another example, when the prompt information includes a disease identifier of the target disease, the disease description information of the target disease can be obtained based on the disease identifier of the target disease (for example, based on the disease identifier of the target disease, the disease description information of the target disease is searched from a disease description database, an encyclopedia entry, or a disease dictionary). The target medical research indicator is determined by matching the disease description information of the target disease with a plurality of medical research indicators contained in an indicator database. For example, a medical research indicator whose degree of matching with the disease description information is greater than a degree threshold is determined as the target medical research indicator. Alternatively, the target medical research indicator is determined by inputting the disease description information into a large language model through the plurality of medical research indicators contained in the indicator database by the large language model. Thus, the accuracy of the target medical research indicator is further improved through the disease description information.
[0056] S203, outputting recommendation information, where the recommendation information indicates target medical research indicators.
[0057] Among them, the recommendation information may include identification information of the target medical research indicator (such as indicator name, indicator number, etc.) to indicate the target medical research indicator through the identification information.
[0058] In this embodiment, the recommendation information can be output through a display device, or the recommendation information can be sent to a terminal where the user is located, so as to recommend the target medical research indicator to the user.
[0059] In an embodiment of the present application, a large language model is used to replace medical researchers to identify medical research indicators from massive scientific research literature in multiple fields, and an indicator database is constructed. On the one hand, the indicator database contains medical research indicators in multiple fields, which improves the richness of the medical research indicators in the indicator database, and on the other hand, the efficiency of constructing the indicator database is improved; based on the prompt information and the indicator database constructed as above, the target medical research indicators to be recommended are determined, so that the target research indicators come from medical research indicators in multiple fields and are related to the indication information, breaking the constraints of personal cognition on subject research and providing inspiration and discovery for medical research.
[0060] See also Figure 3 In another exemplary embodiment, a process for determining a first medical research indicator in a medical research indicator recommendation method for large-model literature learning is provided, wherein the first medical research indicator is obtained by the following operations:
[0061] S301, obtain scientific research literature.
[0062] Among them, the scientific research literature and data can refer to the description of the aforementioned embodiments and will not be repeated here.
[0063] In this embodiment, scientific research literature in multiple fields can be collected from scientific research literature sharing websites (websites that disclose scientific research literature and authorize the download of scientific research literature).
[0064] In one example, data extraction tools can be used to collect scientific research literature in multiple fields from scientific research literature sharing websites, thereby achieving automatic collection of scientific research literature in multiple fields and improving the efficiency of collecting scientific research literature in multiple fields.
[0065] Furthermore, the data extraction tool may be at least one of the following: an open source data crawling framework (such as Scrapy), an extensible markup language (XML) path language (XPath), and a cascading style sheets (CSS) selector. Thus, the efficiency of obtaining scientific research literature can be improved through these data extraction tools.
[0066] S302, analyzing scientific research literature to obtain analysis data.
[0067] The parsed data includes the literature content in the scientific research literature, such as the title, experimental data, graphics, text, tables, etc. of the scientific research literature.
[0068] In this embodiment, the data formats of the scientific research literature data collected from the sharing website are various, for example, the data formats of the scientific research literature data are as follows: XML format, hypertext markup language (HTML) format, document (DOC) format, portable document format (PDF). For scientific research literature data in different data formats, different parsing methods can be used to parse them, and the parsed data corresponding to the scientific research literature data can be obtained, thereby improving the accuracy of parsing the scientific research literature data.
[0069] For example, for scientific research documents in XML format, the document object model (DOM) is used for parsing. For scientific research documents in HTML format, the regular expression-based method is used for parsing. I will not list them all here.
[0070] S303: According to the input data requirements of the large language model, the parsed data is formatted and denoised to obtain target text data.
[0071] The input data requirements include requirements on one or more aspects such as data length and word segmentation.
[0072] The order of the format processing and the denoising processing is not limited here. The format processing may be performed first and then the denoising processing, or the denoising processing may be performed first and then the format processing.
[0073] In this embodiment, since the scientific research literature has different sources (from different websites), the formats of the scientific research literature are varied, and the parsed data obtained by parsing are also different, such as the data volume, data length, symbols, etc. of the parsed data are different. It is necessary to format the parsed data according to the input data requirements of the large language model, such as data clipping, numerical transformation, etc., to obtain the format-processed parsed data. The scientific research literature may also contain some redundant information, such as data unrelated to medicine, repeated data, etc., and the parsed data of the scientific research literature can be denoised to improve the quality of the input data of the large language model.
[0074] In one example, the process of formatting the parsed data includes: normalizing the parsed data to convert the format of the parsed data into text data that meets the input data requirements.
[0075] Among them, data normalization includes at least one of the following: batch normalization (BN), instance normalization (IN), layer normalization (LN), group normalization (GN), and positional normalization (PONO). These normalization methods belong to the field of deep learning. By processing parsed data through these normalization methods, the accuracy and efficiency of format processing of parsed data can be improved.
[0076] In this example, in batch normalization, the variance and mean of the parsed data can be calculated, and the parsed data can be normalized according to the variance and mean of the parsed data; in instance normalization, the mean and standard deviation of each feature dimension (or feature channel) of the parsed data are calculated respectively to obtain the mean and standard deviation corresponding to each feature dimension, and each feature dimension is normalized according to the mean and standard deviation corresponding to each feature dimension; the layer normalization process is similar to the instance normalization process, and each feature dimension of the parsed data is normalized separately; in group normalization, multiple feature dimensions of the parsed data are grouped, and the feature dimensions in each group are normalized; in position normalization, the feature representation of the parsed data is normalized according to the position information of each element in the feature representation of the parsed data, and the influence of the data position is considered in the normalization process.
[0077] In one example, the process of denoising the parsed data includes: identifying data noise in the parsed data through a data discriminator; and removing the data noise in the parsed data. The data discriminator includes at least one of the following: a language discriminator, a quality discriminator, a privacy discriminator, and a security discriminator. Thus, noise is identified on the parsed data from one or more aspects of language, quality, privacy, and security, thereby improving the denoising effect of the parsed data.
[0078] In this embodiment, a language discriminator can be used to identify whether there is data in the parsed data whose language does not meet the input data requirements. If so, the data can be determined as data noise in the parsed data, or the data can be translated into the target language in the input data requirements, and then other discriminators (quality discriminators, privacy discriminators and / or security discriminators) can be used to determine whether the translated data is data noise in the parsed data. A quality discriminator can be used to identify whether there is low-quality data in the parsed data. If so, the low-quality data in the parsed data is determined as data noise. Low-quality data, such as duplicate data, and data not related to medicine (such as website addresses). A privacy discriminator can be used to identify data involving personal or institutional privacy in the parsed data (such as name, address, etc.), and this part of the data can be determined as data noise to protect privacy data by denoising; a security discriminator can be used to identify data with security risks or hazards in the parsed data, such as some obviously wrong views, slogans, etc.
[0079] Among them, the language discriminator, quality discriminator, privacy discriminator, and security discriminator can be neural networks. The language, quality, privacy and / or security discrimination can be achieved through multiple neural networks, or through one neural network.
[0080] S304, using a large language model, identifying medical research indicators for the target text data to obtain a first medical research indicator.
[0081] Among them, the large language model is a pre-trained model.
[0082] In one example, the training data may include training text and training labels corresponding to the training text, the training labels indicating the medical research indicators contained in the training text, and the medical research indicators are identified by inputting the training text into a large language model through the large language model to obtain the identified medical research indicators, and the error value of the large language model is obtained by comparing the identified medical research indicators with the training labels; and the parameters of the large language model are adjusted according to the training error values.
[0083] In this embodiment, multiple target text data can be obtained by parsing, formatting and denoising the scientific research literature. For each target text data, feature extraction can be performed on the target text data through a large language model, and medical research indicators can be identified based on the extracted features, and finally the first medical research indicator can be identified.
[0084] In one example, the target text data can be encoded through a word embeddings (also known as Word2Vec) model to obtain a feature vector representation of the target text data, and the feature vector representation is input into a large language model. The feature vector representation is extracted through the large language model to obtain text features, and medical research indicators are identified based on the extracted features, and finally the first medical research indicator is identified.
[0085] In one example, considering that there may be multiple different pronouns for the same entity in scientific research literature, and these multiple different pronouns are synonyms, in order to make one entity correspond to one pronoun, a synonym mining model can be used to perform synonym mining on the feature vector representation corresponding to the target text data to obtain synonyms of the same entity in the target text data, and the word vectors corresponding to the synonyms of the same entity are unified into the same word vector, so that one entity corresponds to one pronoun.
[0086] Optionally, the synonym mining model adopts a distributional and pattern integrated embedding framework (DPE) model (also known as a DFE algorithm) to improve the accuracy of synonym mining through the DFE model.
[0087] In this optional solution, the DPE model includes a first network and a second network, and the first network and the second network share the feature vector representation of the target text data. In the process of synonym mining, the feature vector representation of the target text data is input into the first network and the second network respectively; in the first network, the distribution characteristics of the words in the target text data are extracted from the feature vector representation of the target text data, and synonym recognition is performed based on the distribution characteristics of the words in the target text data, and the first probability that multiple pairs of words in the target text data are synonyms is obtained and output. For example, based on the characteristic that synonyms are likely to be distributed in similar contextual content, synonym recognition is performed. If there are multiple words in the target text data whose distribution characteristics meet this characteristic, then the first network may recognize the multiple words as synonyms; the second network learns the pattern characteristics of sentences containing synonyms (for example, A is also called B, and A and B have the same meaning). In the second network, the feature vector representation of the target text data can be processed. The grammatical and semantic features are extracted to obtain the lexical features of the words in the target text data and the syntactic features of the target text data. Synonym recognition is performed in combination with the lexical features of the words in the target text data and the syntactic features of the target text data to obtain and output the second probability that multiple pairs of words in the target text data are synonyms. Afterwards, the output data of the first network (i.e., the first probability that multiple pairs of words are synonyms) and the output data of the second network (i.e., the second probability that multiple pairs of words are synonyms) are weighted and summed to obtain the target probability that multiple pairs of words in the target text data are synonyms. Based on the target probability that multiple pairs of words in the target text data are synonyms, it is determined whether the multiple pairs of words are synonyms. If the target probability that at least one pair of words in the target text data are synonyms is greater than or equal to a probability threshold, it can be determined that the at least one pair of words are synonyms.
[0088] In the embodiment of the present application, the target text data is obtained by parsing, formatting and denoising the scientific research literature data in multiple fields, thereby improving the quality of the text data provided to the large language model for processing; the target text data is identified by the large language model for medical research indicators to obtain the first medical research indicator. Thus, the medical research indicators in multiple fields are identified from the scientific research literature data in multiple fields by the large language model, thereby improving the efficiency of the medical research indicators and preparing sufficient medical research indicators for the recommendation of medical research indicators in multiple fields.
[0089] In some embodiments, the indicator database also includes a plurality of second medical research indicators, which are obtained by using a large language model to identify medical research indicators from clinical diagnosis and treatment data of multiple departments. Thus, the medical research indicators in the indicator database are obtained from the scientific research data and actual clinical data, thereby improving the richness of the medical research indicators in the indicator database.
[0090] Below, the implementation process of extracting the second medical research indicator from clinical diagnosis and treatment data is provided.
[0091] See also Figure 4 In another exemplary embodiment, a process for determining a second medical research indicator in a medical research indicator recommendation method for large model literature learning is provided, and the second medical research indicator is obtained by the following operations:
[0092] S401, obtaining clinical diagnosis and treatment data of multiple departments, wherein the clinical diagnosis and treatment data includes clinical documents corresponding to multiple patients, and the clinical documents include at least one of the following information: consultation information, physical examination information, and physiological sample test information.
[0093] Among them, medical consultation information refers to the data recorded during the medical consultation process, such as the patient's gender, age, symptoms, and living habits. Physical examination information refers to the examination results obtained by performing physical examinations on patients through instruments or based on the doctor's experience, such as the patient's medical imaging results (which may include images and text), the results of auscultation through a stethoscope, and the body temperature through a thermometer. Physiological sample test information refers to the test results obtained by testing the patient's physiological samples, such as the patient's blood test results, urine test results, etc.
[0094] In this embodiment, clinical medical records that are allowed to be used for medical research can be obtained. Using the patient's identification information (such as medical treatment number) as an index, the patient's medical records in the past can be obtained from the clinical medical records (if multiple consultations have been conducted, multiple medical records of the patient can be obtained), and the patient's clinical documents can be constructed based on the patient's medical records in the past.
[0095] In one example, in the process of constructing a patient's clinical document, the patient's identification information is used as the root node, and the patient's consultation information, physical examination information, and physiological sample examination information are used as leaf nodes, so that the patient's diagnosis and treatment conditions are recorded in a tree structure in the clinical document.
[0096] S402, using a large language model, identifying medical research indicators for clinical diagnosis and treatment data to obtain a second medical research indicator.
[0097] The large language model and the large language model in the aforementioned embodiment may be the same model.
[0098] In this embodiment, unlike scientific research literature in multiple fields, clinical diagnosis and treatment data contains multiple clinical documents with a unified format and most of the content is related to medicine, so it is not necessary to perform parsing, format processing and denoising like scientific research literature. For each clinical document in the clinical diagnosis and treatment data, the large language model can be used to extract features from the clinical document, and based on the extracted features, the medical research indicators are identified, and finally the second medical research indicator is identified.
[0099] In one example, there may be synonyms in clinical diagnosis and treatment data. Before using a large language model to identify medical research indicators in clinical diagnosis and treatment data, the clinical diagnosis and treatment data may be processed for synonyms, which may include identifying synonyms in the clinical diagnosis and treatment data and unifying the identified synonyms into the same words. Thus, the quality of clinical diagnosis and treatment data can be improved through synonym processing.
[0100] Optionally, a synonym mining model is used to perform synonym mining (ie, identification) on clinical diagnosis and treatment data to improve the accuracy of synonym identification.
[0101] Furthermore, the synonym mining model is a DPE model. The process of performing synonym mining by using the DPE model can refer to the above-mentioned embodiment and will not be described in detail.
[0102] In one example, clinical diagnosis and treatment data may contain text and images. Before using a large language model to identify medical research indicators for the clinical diagnosis and treatment data, the clinical diagnosis and treatment data may be encoded into images and texts to obtain a feature vector representation corresponding to the clinical diagnosis and treatment data. Afterwards, the feature vector representation may be identified as a medical research indicator using a large language model to obtain a second medical research indicator.
[0103] In one example, S402 includes: performing synonym processing and image text encoding on clinical diagnosis and treatment data through convolutional neural networks and contrastive language-image pretraining (CLIP) models based on contrastive learning, and obtaining feature vector representations corresponding to the clinical diagnosis and treatment data; performing medical research indicator recognition on the feature vector representation through a large language model, and obtaining a second medical research indicator. Thus, the accuracy of synonym processing of clinical diagnosis and treatment data is improved by using convolutional neural networks, and the accuracy of encoding clinical diagnosis and treatment data is improved by using the CLIP model.
[0104] In this example, a trained convolutional neural network (for example, a convolutional neural network used for vocabulary classification in text, which can be trained to identify words with similar semantic features through classification) can be used to perform synonym recognition on clinical diagnosis and treatment data, and the identified synonyms can be converted into the same words to obtain clinical diagnosis and treatment data after synonym processing; the CLIP model (after inputting text and / or images, it can output a feature vector representation of the text and / or a feature vector representation of the image) can be used to perform image-text encoding on the clinical diagnosis and treatment data after synonym processing to obtain a feature vector representation corresponding to the clinical diagnosis and treatment data.
[0105] In the embodiment of the present application, a large language model is used to identify medical research indicators for clinical diagnosis and treatment data of multiple departments to obtain a second medical research indicator. Since the clinical diagnosis and treatment data is data from multiple departments, the second medical research indicator obtained is also a medical research indicator for multiple departments, that is, a multi-field and cross-field medical research indicator. Thus, the efficiency of medical research indicators is improved, and sufficient medical research indicators are prepared for the recommendation of medical research indicators in multiple fields.
[0106] See also Figure 5 In another exemplary embodiment, a method for recommending medical research indicators based on large-model literature learning is provided, the method comprising the following steps:
[0107] S501, obtaining prompt information related to medical research.
[0108] In this embodiment, prompt information related to the target disease can be obtained.
[0109] S502, determining the target medical research indicator according to the prompt information and the indicator database.
[0110] The implementation principles and technical effects of S501 to S502 may refer to the aforementioned embodiments and will not be described in detail.
[0111] S503, obtaining clinical documents related to the target disease.
[0112] In this embodiment, clinical documents of patients with the target disease can be obtained from the clinical diagnosis and treatment data to obtain clinical documents related to the target disease. For example, if the target disease is hypertension, clinical documents of patients with hypertension can be obtained. Since there may be multiple patients with the target disease, multiple clinical documents related to the target disease can be obtained from the clinical diagnosis and treatment data.
[0113] Among them, the clinical diagnosis and treatment data and clinical documents can refer to the description of the aforementioned embodiments and will not be repeated here.
[0114] S504, identifying the index value of the target medical research index for the clinical document, and obtaining a disease-specific database of the target disease, wherein the disease-specific database contains the index value of the target medical research index.
[0115] The large language model and the large language model in the aforementioned embodiment may be the same model.
[0116] In this embodiment, the index value of the target medical research indicator can be identified for clinical documents related to the target disease through a large language model. For example, in the clinical document, the index value of the medical research indicator of hemoglobin is identified, and the index value of the medical research indicator of whether the user smokes is identified. A disease-specific database for the target disease is constructed, and after the index value of the target medical research indicator is obtained, the index value of the target medical research indicator can be stored in the disease-specific database.
[0117] Among them, one clinical document can be identified to obtain a set of indicator data, and multiple sets of indicator data can be identified by identifying multiple clinical documents. Therefore, multiple sets of indicator data can be included in the special disease database, and one set of indicator data includes the indicator values of the target medical research indicators.
[0118] In one example, S502 and S504 can be performed together by a large language model. Specifically, after obtaining prompt information related to the target disease and clinical documents related to the target disease, the prompt information and clinical documents can be input into the large language model; in the large language model, according to the prompt information, the target medical research indicator is determined from multiple medical research indicators included in the indicator database; through the large language model, the indicator value of the target medical research indicator is identified in the clinical document.
[0119] S505, performing data analysis on the disease-specific database to obtain the impact factor of the target medical research indicator, where the impact factor indicates the degree of impact of the target medical research indicator on the target disease.
[0120] Among them, the impact factor indicates the degree of influence of the target medical research indicator on the target disease, that is, the impact factor indicates the degree of correlation between the target medical research indicator and the target disease.
[0121] In one example, the data in the disease-specific database can be analyzed through the corresponding indicator impact analysis model to obtain the impact factor of the target medical research indicator. For example, the larger the impact factor of the target medical research indicator, the higher the impact of the target medical research indicator on the target disease.
[0122] In this example, the indicator impact analysis model is a pre-trained machine learning model. During the training process, the training data contains multiple groups of indicator data used for training. The indicator impact analysis model is used to perform influence analysis on the multiple groups of indicator data used for training to obtain predicted impact factors corresponding to multiple medical research indicators. According to the predicted impact factors corresponding to the multiple medical research indicators and the label data corresponding to the training data (the actual impact factors corresponding to the multiple medical research indicators), the model parameters of the indicator impact analysis model are adjusted.
[0123] In another example, a special disease database contains multiple groups of indicator data, and one group of indicator data includes the indicator value of the target medical research indicator. The influencing factor of the target medical research indicator can be determined by performing statistical analysis on the indicator values of the target medical research indicator in the multiple groups of indicator data.
[0124] In this example, the indicator values of the target medical research indicators in multiple groups of indicator data can be compared with the normal range corresponding to the target medical research indicators, and the influencing factors of the target medical research indicators can be determined based on the comparison results. If the indicator values of the target medical research indicator in multiple groups of indicator data are all outside the normal range, then the influencing factor of the target medical research indicator can be determined to be a first value, and the first value indicates that the target medical research indicator has a greater influence or relationship on the target disease; if a part of the indicator values of the target medical research indicator in multiple groups of indicator data are outside the normal range, and the proportion of the indicator values outside the normal range in the indicator values of the target research indicator in multiple groups of indicator data is greater than the set proportion value, then the influencing factor of the target medical research indicator can be determined to be a second value, and the second value is less than the first value; if a part of the indicator values of the target medical research indicator in multiple groups of indicator data are outside the normal range, and the proportion of the indicator values outside the normal range in the indicator values of the target research indicator in multiple groups of indicator data is less than the first proportion value, then the influencing factor of the target medical research indicator can be determined to be a third value, and the third value is less than the second value; if the indicator values of the target medical research indicator in multiple groups of indicator data are all within the normal range, then the influencing factor of the target medical research indicator can be determined to be a fourth value, and the fourth value is less than the second value and the fourth value is greater than or equal to 0.
[0125] S506, output recommendation information, the recommendation information indicates the target medical research indicators, special disease database and impact factor.
[0126] In this embodiment, after determining the target medical research indicators, the special disease database, and the influencing factors of the target medical research indicators, the recommendation information can be output through the display device, or the recommendation information can be sent to the terminal where the user is located to recommend the target medical research indicators, the special disease database, and the influencing factors of the target medical research indicators to the user as reference data for medical research.
[0127] In one example, the target medical research indicator can be adjusted according to the influencing factor of the target medical research indicator to obtain the adjusted target medical research indicator. Therefore, the recommendation information can indicate the adjusted target medical research indicator, the indicator value of the adjusted target medical research indicator, and the influencing factor of the adjusted target medical research indicator. Thus, the accuracy of the recommendation information is improved.
[0128] In this example, according to the impact factor of the target medical research indicator, the position order of the target medical research indicator in the recommended information can be adjusted, and / or the target medical research indicator can be deleted. For example, in the recommended information, the target medical research indicator with a larger impact factor is arranged in front; for another example, the target medical research indicator with an impact factor less than a set threshold is deleted.
[0129] In the embodiment of the present application, a large language model is used to replace medical researchers, and medical research indicators are identified from massive scientific research literature in multiple fields to construct an indicator database. On the one hand, the indicator database contains medical research indicators in multiple fields, which improves the richness of medical research indicators in the indicator database, and on the other hand, it improves the efficiency of constructing the indicator database. Based on the prompt information and the indicator database constructed above, the target medical research indicator to be recommended is determined; based on the clinical documents related to the target disease and the target medical research indicator, a special disease database containing the indicator value of the target medical research indicator is constructed; data analysis is performed on the special disease database to determine the influencing factor of the target medical research indicator; output indicates the target medical research indicator, the influencing factor of the target medical research indicator, and the recommendation information of the special disease database, which calmly breaks the constraints of personal cognition on subject research, provides cross-domain medical research indicators for medical research, and also provides a special disease database and the influencing factor of medical research indicators, providing inspiration discovery and data assistance for medical research.
[0130] Exemplary Devices
[0131] Correspondingly, the embodiment of the present application also provides a medical research indicator recommendation device based on large model literature learning.
[0132] See also Figure 6 In an exemplary embodiment, a medical research indicator recommendation device 600 based on large model literature learning is provided. The medical research indicator recommendation device 600 based on large model literature learning includes: an acquisition unit 601, a processing unit 602 and an output unit 603, wherein:
[0133] The acquisition unit 601 is used to acquire prompt information related to medical research; the processing unit 602 is used to determine the target medical research indicator based on the prompt information and a pre-constructed indicator database, the indicator database includes multiple first medical research indicators, and the multiple first medical research indicators are obtained by identifying medical research indicators of scientific research literature materials in multiple fields through a large language model; the output unit 603 is used to output recommendation information, and the recommendation information indicates the target medical research indicator.
[0134] In some embodiments, the first medical research indicator is obtained by the following operations: acquiring scientific research literature; parsing the scientific research literature to obtain parsed data; formatting and denoising the parsed data according to the input data requirements of the large language model to obtain target text data; and identifying the medical research indicators of the target text data through the large language model to obtain the first medical research indicator.
[0135] In some embodiments, the process of formatting the parsed data includes: converting the format of the parsed data into text data that meets the input data requirements by normalizing the parsed data; wherein data normalization includes at least one of the following: batch normalization, instance normalization, layer normalization, group normalization, and position normalization.
[0136] In some embodiments, the process of denoising the parsed data includes: identifying data noise in the parsed data through a data discriminator; removing the data noise in the parsed data; wherein the data discriminator includes at least one of the following: a language discriminator, a quality discriminator, a privacy discriminator, and a security discriminator.
[0137] In some embodiments, the indicator database also includes multiple second medical research indicators, and the second medical research indicators are obtained by identifying medical research indicators of clinical diagnosis and treatment data of multiple departments through a large language model.
[0138] In some embodiments, the second medical research indicator is obtained by the following operations: obtaining clinical diagnosis and treatment data, the clinical diagnosis and treatment data includes clinical documents corresponding to multiple patients, and the clinical documents include at least one of the following information: medical consultation information, physical examination information and physiological sample test information; through a large language model, the medical research indicator is identified for the clinical diagnosis and treatment data to obtain the second medical research indicator.
[0139] In some embodiments, a large language model is used to identify medical research indicators of clinical diagnosis and treatment data to obtain a second medical research indicator, including: performing synonym processing and image text encoding on the clinical diagnosis and treatment data through a convolutional neural network and a language-image pre-training model based on contrastive learning to obtain a feature vector representation corresponding to the clinical diagnosis and treatment data; and a large language model is used to identify medical research indicators of the feature vector representation to obtain a second medical research indicator.
[0140] In some embodiments, the prompt information includes a disease identifier of the target disease. The acquisition unit is further used to: acquire clinical documents related to the target disease. The processing unit 602 is further used to: identify the index value of the target medical research index in the clinical document, obtain a disease-specific database of the target disease, and the disease-specific database contains the index value; perform data analysis on the disease-specific database to obtain an impact factor of the target medical research index, and the impact factor indicates the degree of impact of the target medical research index on the target disease. Among them, the prompt information also indicates the disease-specific database and the impact factor.
[0141] In some embodiments, the processing unit 602 is also used to: adjust the target medical research indicator according to the impact factor; wherein the prompt information indicates the adjusted target medical research indicator, the indicator value of the adjusted target medical research indicator and the impact factor of the adjusted target medical research indicator.
[0142] The medical research indicator recommendation device based on large model literature learning provided in this embodiment belongs to the same application concept as the medical research indicator recommendation method based on large model literature learning provided in the above embodiments of this application, and can execute the medical research indicator recommendation method based on large model literature learning provided in any of the above embodiments of this application, and has the corresponding functional modules and beneficial effects of the execution method. For the technical details not described in detail in this embodiment, please refer to the specific processing content of the medical research indicator recommendation method based on large model literature learning provided in the above embodiments of this application, which will not be repeated here.
[0143] The functions implemented by each unit in the device provided in the above embodiment (medical research indicator recommendation device based on large model literature learning) can be implemented by the same or different processors respectively, and the embodiments of this application are not limited thereto.
[0144] It should be understood that each unit in the device provided in the above embodiment can be implemented in the form of a processor calling software. For example, the device includes a processor, the processor is connected to a memory, and instructions are stored in the memory. The processor calls the instructions stored in the memory to implement any of the above methods or realize the functions of each unit of the device, wherein the processor can be a general-purpose processor, such as a CPU or a microprocessor, etc., and the memory can be a memory in the device or a memory outside the device. Alternatively, the unit in the device can be implemented in the form of a hardware circuit, and the functions of some or all units can be realized by designing the hardware circuit. The hardware circuit can be understood as one or more processors; for example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above units are realized by designing the logical relationship of the components in the circuit; for another example, in another implementation, the hardware circuit can be implemented by PLD, taking FPGA as an example, which can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by a configuration file, so as to realize the functions of some or all of the above units. All units of the above device can be implemented in the form of a processor calling software, or in the form of a hardware circuit, or in part by a processor calling software, and the remaining part is implemented in the form of a hardware circuit.
[0145] In an embodiment of the present application, a processor is a circuit with the ability to process signals. In one implementation, the processor may be a circuit with the ability to read and run instructions, such as a CPU, a microprocessor, a GPU, or a DSP; in another implementation, the processor may implement certain functions through the logical relationship of a hardware circuit, and the logical relationship of the hardware circuit is fixed or reconfigurable, such as a hardware circuit implemented by an ASIC or PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the hardware circuit configuration can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as an NPU, TPU, DPU, etc.
[0146] It can be seen that each unit in the above device can be one or more processors (or processing circuits) configured to implement the above method, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.
[0147] In addition, all or part of the units in the above device can be integrated together, or can be implemented independently. In one implementation, these units are integrated together and implemented in the form of a SOC. The SOC may include at least one processor for implementing any of the above methods or implementing the functions of each unit of the device. The type of the at least one processor may be different, for example, including a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.
[0148] Exemplary Electronic Devices
[0149] Another embodiment of the present application also provides an electronic device. Figure 7 As shown, the electronic device includes: a memory 700 and a processor 710; wherein the memory 700 is connected to the processor 710 for storing programs; the processor 710 is used to implement the medical research indicator recommendation method for large model literature learning disclosed in any of the above embodiments by running the program stored in the memory 700.
[0150] Specifically, the electronic device may further include: a bus, a communication interface 720 , an input device 730 and an output device 740 .
[0151] The processor 710, the memory 700, the communication interface 720, the input device 730 and the output device 740 are connected to each other via a bus.
[0152] A bus may include a pathway that transfers information between components of a computer system.
[0153] The processor 710 may be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present application. It may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0154] The processor 710 may include a main processor, and may also include a baseband chip, a modem, and the like.
[0155] The memory 700 stores a program for executing the technical solution of the present application, and may also store an operating system and other key services. Specifically, the program may include a program code, and the program code includes a computer operation instruction. More specifically, the memory 700 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk storage, a flash, and the like.
[0156] The input device 730 may include a device for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor.
[0157] Output device 740 may include devices that allow information to be output to a user, such as a display screen, printer, speaker, etc.
[0158] The communication interface 720 may include any transceiver or the like to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.
[0159] The processor 710 executes the program stored in the memory 700 and calls other devices, which can be used to implement the various steps of the medical research indicator recommendation method for large-model literature learning provided in any of the above embodiments of the present application.
[0160] An embodiment of the present application also proposes a chip, which includes a processor and a data interface. The processor reads and runs the program stored in the memory through the data interface to execute the medical research indicator recommendation method based on large model literature learning introduced in any of the above embodiments. The specific processing process and its beneficial effects can be found in the embodiment introduction of the medical research indicator recommendation method based on large model literature learning mentioned above.
[0161] Exemplary computer program products and storage media
[0162] In addition to the above-mentioned methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the medical research indicator recommendation method based on large model literature learning according to various embodiments of the present application described in any of the above embodiments of this specification.
[0163] The computer program product may be written in any combination of one or more programming languages to write program codes for performing the operations of the embodiments of the present application, including object-oriented programming languages, such as Java, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0164] In addition, an embodiment of the present application may also be a storage medium on which a computer program is stored, and the computer program is executed by a processor to execute the steps of the medical research indicator recommendation method based on large model literature learning according to various embodiments of the present application described in any of the above embodiments of this specification.
[0165] For the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the order of the actions described, because according to the present application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0166] It should be noted that each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the embodiments can be referred to each other. For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0167] The steps in the methods of each embodiment of the present application can be adjusted in order, combined and deleted according to actual needs, and the technical features recorded in each embodiment can be replaced or combined.
[0168] The units in the devices in the embodiments of the present application can be combined, divided and deleted according to actual needs.
[0169] In the several embodiments provided in the present application, it should be understood that the disclosed terminals, devices and methods can be implemented in other ways. For example, the terminal embodiments described above are only schematic, for example, the division of modules or submodules is only a logical function division, and there may be other division methods in actual implementation, for example, multiple submodules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.
[0170] The modules or submodules described as separate components may or may not be physically separated, and the components of the modules or submodules may or may not be physical modules or submodules, that is, they may be located in one place, or they may be distributed on multiple network modules or submodules. Some or all of the modules or submodules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0171] In addition, each functional module or submodule in each embodiment of the present application may be integrated into one processing module, or each module or submodule may exist physically separately, or two or more modules or submodules may be integrated into one module. The above-mentioned integrated modules or submodules may be implemented in the form of hardware or in the form of software functional modules or submodules.
[0172] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0173] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly by hardware, software units executed by a processor, or a combination of the two. The software units may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0174] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0175] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for recommending medical research indicators based on large model literature learning, characterized in that: include: Get tips related to medical research; Determine a target medical research indicator according to the prompt information and a pre-built indicator database, wherein the indicator database includes a plurality of first medical research indicators, and the plurality of first medical research indicators are obtained by performing medical research indicator recognition on scientific research literature materials in multiple fields through a large language model; The recommendation information is output, where the recommendation information indicates the target medical research indicator.
2. The method for recommending medical research indicators based on large model literature learning according to claim 1 is characterized in that: The first medical research indicator is obtained by the following operation: Obtaining the scientific research literature; Analyze the scientific research literature to obtain analytical data; According to the input data requirements of the large language model, formatting and denoising the parsed data to obtain target text data; The target text data is subjected to medical research indicator identification through the large language model to obtain the first medical research indicator.
3. The method for recommending medical research indicators based on large model literature learning according to claim 2 is characterized in that: The process of formatting the parsed data includes: By performing data normalization on the parsed data, the format of the parsed data is converted into text data that meets the requirements of the input data; The data normalization includes at least one of the following: batch normalization, instance normalization, layer normalization, group normalization and position normalization.
4. The method for recommending medical research indicators based on large model literature learning according to claim 2, characterized in that: The process of performing denoising processing on the analyzed data includes: identifying data noise in the analyzed data by a data discriminator; removing data noise from the analyzed data; The data discriminator includes at least one of the following: a language discriminator, a quality discriminator, a privacy discriminator and a security discriminator.
5. The method for recommending medical research indicators based on large model literature learning according to any one of claims 1 to 4, characterized in that: The indicator database also includes multiple second medical research indicators, which are obtained by using the large language model to identify medical research indicators on clinical diagnosis and treatment data of multiple departments.
6. The method for recommending medical research indicators based on large model literature learning according to claim 5 is characterized in that: The second medical research indicator is obtained by the following operation: Acquire the clinical diagnosis and treatment data, wherein the clinical diagnosis and treatment data includes clinical documents corresponding to a plurality of patients, and the clinical documents include at least one of the following information: consultation information, physical examination information, and physiological sample test information; The large language model is used to identify medical research indicators on the clinical diagnosis and treatment data to obtain the second medical research indicators.
7. The method for recommending medical research indicators based on large model literature learning according to claim 6 is characterized in that: The step of identifying medical research indicators on the clinical diagnosis and treatment data by using the large language model to obtain the second medical research indicator includes: Through a convolutional neural network and a language-image pre-training model based on contrastive learning, the clinical diagnosis and treatment data are subjected to synonym processing and image-text encoding to obtain a feature vector representation corresponding to the clinical diagnosis and treatment data; The medical research indicator is identified by using the large language model for the feature vector representation to obtain the second medical research indicator.
8. The method for recommending medical research indicators based on large model literature learning according to any one of claims 1 to 4, characterized in that: The prompt information includes a disease identifier of the target disease. After determining the target medical research indicator according to the prompt information and the pre-built indicator database, the method further includes: obtaining clinical documentation related to the target disease; Identify the index value of the target medical research index on the clinical document to obtain a disease-specific database of the target disease, wherein the disease-specific database contains the index value; Performing data analysis on the disease-specific database to obtain an impact factor of the target medical research indicator, wherein the impact factor indicates the degree of impact of the target medical research indicator on the target disease; The recommendation information further indicates the disease-specific database and the impact factor.
9. The method for recommending medical research indicators based on large model literature learning according to claim 8, characterized in that: After analyzing the data of the disease-specific database to obtain the impact factor of the target medical research indicator, the method further includes: According to the impact factor, the target medical research indicator is adjusted; Among them, the recommendation information indicates the adjusted target medical research indicator, the indicator value of the adjusted target medical research indicator and the influencing factor of the adjusted target medical research indicator.
10. A medical research indicator recommendation device based on large model literature learning, characterized in that: include: An acquisition unit, used for acquiring prompt information related to medical research; A processing unit, configured to determine a target medical research indicator according to the prompt information and a pre-constructed indicator database, wherein the indicator database includes a plurality of first medical research indicators, and the plurality of first medical research indicators are obtained by performing medical research indicator recognition on scientific research literature materials in multiple fields through a large language model; An output unit is used to output recommendation information, wherein the recommendation information indicates the target medical research indicator.
11. An electronic device, characterized in that: including memory and processor; The memory is connected to the processor and is used to store programs; The processor is used to implement the medical research indicator recommendation method based on large model literature learning as described in any one of claims 1 to 9 by running the program in the memory.
12. A computer program product, characterized in that It includes a computer program, which, when executed by a processor, implements the medical research indicator recommendation method based on large model literature learning as described in any one of claims 1 to 9.
Citation Information
Cited By
Complication risk assessment method and system for multiple myeloma
CN120299722A