Method and device for fine-tuning pre-trained model and extracting scientific hypothesis information, medium and product
By fine-tuning and training the pre-trained model, a scientific hypothesis extraction model is generated, which solves the problem of inaccurate scientific hypothesis information extraction in the existing technology and achieves the effect of efficiently extracting scientific hypothesis information from literature.
Patent Information
- Application Number
- CN202511062303.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-07-31
AI Technical Summary
The existing technology lacks a dedicated model for extracting scientific hypothesis information, resulting in insufficient accuracy and completeness of scientific hypothesis information extraction.
By fine-tuning the pre-trained model, injecting the trainable adapter and training it with a sample dataset, the weight matrix is updated to generate a scientific hypothesis extraction model, which can automatically identify and structuredly output scientific hypothesis information from document titles and abstracts.
It improves the accuracy and completeness of scientific hypothesis information extraction, helps researchers quickly extract innovative points from literature, and improves the efficiency of project design.
Smart Images

Figure CN120562409B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to a pre-training model fine-tuning and scientific hypothesis information extraction method, device, equipment, medium and product. BACKGROUND
[0002] In the field of biomedical literature mining, researchers often need to quickly extract scientific hypothesis information from a large number of literatures to quickly understand the overall design and ideas of the research.
[0003] The prior art lacks a special model for scientific hypothesis information extraction, and usually relies on a general natural language processing model to extract key sentences as scientific hypothesis information, which has the problem of insufficient accuracy and completeness of scientific hypothesis information extraction. SUMMARY
[0004] The present application provides a pre-training model fine-tuning and scientific hypothesis information extraction method, device, equipment, medium and product to solve the problem of insufficient accuracy and completeness of scientific hypothesis information extraction in the prior art.
[0005] According to an aspect of the present application, a pre-training model fine-tuning method is provided, which comprises:
[0006] determining at least one to-be-trained linear layer from candidate linear layers included in a target pre-training model, and injecting a trainable adapter into the to-be-trained linear layer; wherein the trainable adapter has a trainable weight matrix;
[0007] inputting a sample data set into the target pre-training model, determining sample prompt word information and sample text information in the sample data set through the target pre-training model, and outputting predicted scientific hypothesis information according to the sample prompt word information and the sample text information; wherein the sample prompt word information is used to prompt the target pre-training model to extract a target text element included in the sample text information, and to generate the predicted scientific hypothesis information according to the target text element; the sample text information is literature title information and / or literature abstract information of a sample literature; the target text element includes at least one of a clinical problem, a scientific problem, a molecular target, an action mechanism and an interaction mechanism;
[0008] determining labeled scientific hypothesis information in the sample data set through the target pre-training model, and training the trainable adapter according to the predicted scientific hypothesis information and the labeled scientific hypothesis information, for updating the trainable weight matrix to obtain an updated weight matrix;
[0009] obtaining a scientific hypothesis extraction model of the target pre-training model after fine-tuning according to the updated weight matrix.
[0010] Optionally, the sample data set contains at least one type of candidate structured data;
[0011] The target pre-training model determines the sample prompt word information and sample text information in the sample data set, including:
[0012] Determine the structured role label corresponding to each candidate structured data;
[0013] Obtain the candidate structured data whose structured role label is a system role as the sample prompt word information;
[0014] Obtain the candidate structured data whose structured role label is a user role as the sample text information.
[0015] Optionally, the target pre-training model determines the labeled scientific hypothesis information in the sample data set, including:
[0016] Obtain the candidate structured data whose structured role label is an assistant role as the labeled scientific hypothesis information.
[0017] Optionally, the trainable adapter is trained according to the predicted scientific hypothesis information and the labeled scientific hypothesis information, including:
[0018] Determine at least one candidate word element contained in the labeled scientific hypothesis information, and determine the candidate mask information added to each candidate word element;
[0019] Obtain the candidate word element whose candidate mask information is a target mask as a to-be-learned word element, and generate to-be-learned scientific hypothesis information according to each to-be-learned word element;
[0020] The trainable adapter is trained according to the predicted scientific hypothesis information and the to-be-learned scientific hypothesis information.
[0021] Optionally, the trainable adapter is trained according to the predicted scientific hypothesis information and the to-be-learned scientific hypothesis information, including:
[0022] According to the predicted scientific hypothesis information and the to-be-learned scientific hypothesis information, a target loss value is determined by loss value calculation;
[0023] The trainable adapter is trained according to the target loss value.
[0024] Optionally, the scientific hypothesis extraction model of the target pre-training model after fine-tuning is obtained according to the updated weight matrix, including:
[0025] determine information similarity between the predicted scientific hypothesis information and the labeled scientific hypothesis information, and determine element completeness and element accuracy of a text element contained in the predicted scientific hypothesis information, wherein the information similarity is text similarity and / or semantic similarity;
[0026] In a case where the information similarity, the element completeness, and the element accuracy all satisfy respective threshold values, obtain a scientific hypothesis extraction model after fine-tuning of the target pre-training model according to the updated weight matrix.
[0027] Optionally, obtaining the scientific hypothesis extraction model after fine-tuning of the target pre-training model according to the updated weight matrix comprises:
[0028] obtain an original weight matrix corresponding to the target pre-training model, and perform weight matrix merging according to the original weight matrix and the updated weight matrix to generate a merged weight matrix;
[0029] obtain the scientific hypothesis extraction model after fine-tuning of the target pre-training model according to the merged weight matrix.
[0030] Optionally, inputting the sample data set into the target pre-training model comprises:
[0031] determine an input data length threshold value corresponding to the target pre-training model, determine a sample data length corresponding to the sample data set, and compare the input data length threshold value and the sample data length;
[0032] In a case where the sample data length is less than or equal to the input data length threshold value, input the sample data set into the target pre-training model.
[0033] According to another aspect of the present application, a scientific hypothesis information extraction method is provided, and the method comprises:
[0034] obtain current prompt word information and current text information;
[0035] input the current prompt word information and the current text information into a scientific hypothesis extraction model, and output target scientific hypothesis information according to the current prompt word information and the current text information through the scientific hypothesis extraction model;
[0036] The scientific hypothesis extraction model is obtained by using the fine-tuning method of the pre-training model disclosed in the present application.
[0037] The current prompt word information is used to prompt the scientific hypothesis extraction model to extract a target text element contained in the current text information, and generate the target scientific hypothesis information according to the target text element; the current text information is literature title information and / or literature abstract information of a current literature; and the target text element includes at least one of a clinical problem, a scientific problem, a molecular target, an action mechanism and an interaction mechanism.
[0038] According to another aspect of the present application, a device for fine-tuning a pre-trained model is provided, the device comprising:
[0039] an adapter injection module configured to determine at least one to-be-trained linear layer from candidate linear layers included in the target pre-trained model, and inject a trainable adapter in the to-be-trained linear layer; wherein the trainable adapter has a trainable weight matrix;
[0040] a scientific hypothesis information prediction module configured to input a sample data set into the target pre-trained model, determine sample prompt word information and sample text information in the sample data set by the target pre-trained model, and output predicted scientific hypothesis information according to the sample prompt word information and the sample text information; wherein the sample prompt word information is used to prompt the target pre-trained model to extract a target text element contained in the sample text information, and generate the predicted scientific hypothesis information according to the target text element; the sample text information is literature title information and / or literature abstract information of a sample literature; and the target text element includes at least one of a clinical problem, a scientific problem, a molecular target, an action mechanism and an interaction mechanism;
[0041] an adapter training module configured to determine labeled scientific hypothesis information in the sample data set by the target pre-trained model, and train the trainable adapter according to the predicted scientific hypothesis information and the labeled scientific hypothesis information, to update the trainable weight matrix to obtain an updated weight matrix;
[0042] a scientific hypothesis extraction model generation module configured to obtain a scientific hypothesis extraction model of the fine-tuned target pre-trained model according to the updated weight matrix.
[0043] According to another aspect of the present application, a device for extracting scientific hypothesis information is provided, the device comprising:
[0044] an information acquisition module configured to acquire current prompt word information and current text information;
[0045] The scientific hypothesis information extraction module is configured to input the current prompt word information and the current text information into a scientific hypothesis extraction model, and output target scientific hypothesis information from the scientific hypothesis extraction model according to the current prompt word information and the current text information.
[0046] The scientific hypothesis extraction model is obtained by using the fine-tuning method of the pre-training model.
[0047] The current prompt word information is used to prompt the scientific hypothesis extraction model to extract a target text element contained in the current text information, and generate the target scientific hypothesis information according to the target text element. The current text information is literature title information and / or literature abstract information of a current literature. The target text element includes at least one of a clinical problem, a scientific problem, a molecular target, an action mechanism and an interaction mechanism.
[0048] According to another aspect of the present application, an electronic device is provided, which includes:
[0049] at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the method of any one of the present application.
[0050] According to another aspect of the present application, a computer readable storage medium is provided, which stores computer instructions for enabling a processor to perform the method of any one of the present application.
[0051] According to another aspect of the present application, a computer program product is provided, which includes a computer program for enabling a processor to perform the method of any one of the present application.
[0052] The present application obtains a scientific hypothesis extraction model which can be used for extracting scientific hypothesis information by fine-tuning a pre-training model. The scientific hypothesis extraction model can automatically identify and structure scientific hypothesis information containing core elements such as clinical problems, scientific problems, molecular targets, action mechanisms and interaction mechanisms from literature title information and / or literature abstract information, thereby improving the accuracy and completeness of scientific hypothesis information extraction.
[0053] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor to limit the scope of the present application. BRIEF DESCRIPTION OF DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the accompanying drawings needed to be used in the embodiments description will be briefly introduced. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0055] Figure 1 A flow chart of a pre-training model fine-tuning method provided for the first embodiment of the present application;
[0056] Figure 2 A flow chart of a pre-training model fine-tuning method provided for the second embodiment of the present application;
[0057] Figure 3 A flow chart of a scientific hypothesis information extraction method provided for the third embodiment of the present application;
[0058] Figure 4 A structural schematic diagram of a pre-training model fine-tuning device provided for the fourth embodiment of the present application;
[0059] Figure 5 A structural schematic diagram of a scientific hypothesis information extraction device provided for the fifth embodiment of the present application;
[0060] Figure 6 A structural schematic diagram of an electronic device implementing the pre-training model fine-tuning method and / or the scientific hypothesis information extraction method of the model of the embodiments of the present application. DETAILED DESCRIPTION
[0061] In order to make the technical personnel in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, but not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without any creative effort should belong to the scope of protection of the present application.
[0062] It is to be understood that the terminology "candidate", "target" and the like in the specification and claims of the application and the above-described drawings is used to distinguish between like objects and is not necessarily intended to describe a specific sequential or chronological order, unless otherwise expressly so limited. It is to be understood that the use of the term "or" in the including the variations thereof, such as "including, but not limited to," as well as "containing" and "having," is intended to encompass the inclusion of common or similar features, integers, steps, or elements, and any equivalents thereof, and does not exclude others not expressly listed. It is intended that the phrase "and / or" includes both the conjunctive and disjunctive sense, unless otherwise expressly stated.
[0063] Embodiment one:
[0064] Figure 1 A flowchart of a fine-tuning method of a pre-trained model is provided for the first embodiment of the application. The embodiment can be applied to obtain a scientific hypothesis extraction model through model fine-tuning, which is used to automatically identify and structure scientific hypothesis information containing core elements such as clinical problems, scientific problems, molecular targets, mechanisms of action, and interaction mechanisms from literature title information and / or literature abstract information. The method can be executed by a pre-trained model fine-tuning device, which can be realized in the form of hardware and / or software, such as a server. Figure 1 As shown in the figure, the method comprises:
[0065] S101, determining at least one to-be-trained linear layer from candidate linear layers contained in a target pre-trained model, and injecting a trainable adapter into the to-be-trained linear layer.
[0066] The target pre-trained model refers to a pre-trained model selected as the basis for fine-tuning. The target pre-trained model has been preliminarily trained on general data and has general semantic understanding or generation capabilities, including but not limited to open-source lightweight pre-trained models such as Qwen2.5-7B-instruct, which balance performance and computational efficiency.
[0067] The candidate linear layer refers to a linear layer contained in the model structure of the target pre-trained model. The linear layer is one of the basic components in a neural network, also known as a fully connected layer (Fully Connected Layer) or a dense layer (Dense Layer). Its main function is to map input data to a new feature space through linear transformation, achieving feature extraction and dimension conversion.
[0068] The to-be-trained linear layer refers to the candidate linear layer selected and added with the trainable adapter in the target pre-trained model, and only the trainable weight matrix of the trainable adapter part is trained, and the original weight matrix of the target pre-trained model itself is frozen and does not participate in training. The selection of the to-be-trained linear layer can be set according to actual business requirements, and preferably, all candidate linear layers are fine-tuned to improve the model capability.
[0069] The trainable adapter refers to a lightweight and independently trainable neural network module inserted in the to-be-trained linear layer of the pre-trained model, and the core function is to adapt to the downstream task through a small number of newly added parameters without changing the main structure of the target pre-trained model. For example, the trainable adapter can be a LoRA (Low-Rank Adaptation) adapter.
[0070] The trainable adapter has a trainable weight matrix, which refers to a small parameter matrix added to the to-be-trained linear layer through low-rank decomposition, which is only updated during downstream task training, while the original weight matrix of the target pre-trained model itself remains frozen. Taking the LoRA adapter as an example, the trainable weight matrix can be represented by B·A, where B represents an initialized all-zero matrix, and A represents an initialized Gaussian distribution.
[0071] In an embodiment, in response to the user's selection operation on each candidate linear layer contained in the target pre-trained model, the user-selected candidate linear layer is selected from each candidate linear layer as the to-be-trained linear layer. Further, the original weight matrix of each candidate linear layer is frozen so that it does not update the gradient during the training process, and then the trainable adapter is constructed, and the trainable adapter is injected in parallel beside the to-be-trained linear layer, that is, the trainable weight matrix A and B are added, A adopts a Gaussian distribution during initialization, and B is initialized as an all-zero matrix to ensure that there is no weight disturbance in the initial training stage.
[0072] S102, input the sample data set into the target pre-trained model, determine the sample prompt word information and sample text information in the sample data set through the target pre-trained model, and output the predicted scientific hypothesis information according to the sample prompt word information and the sample text information.
[0073] Among them, the sample data set refers to a structured data set used to train the trainable adapter, including sample prompt word information, sample text information and labeled scientific hypothesis information.
[0074] The sample prompt word information refers to instructions or context information for guiding the target pre-trained model to generate scientific hypothesis information. Its essence is structured artificial input for explicitly defining the task target, constraint condition or output direction of the model. The sample text information refers to the natural language text content contained in the sample data set itself, which is the core original data processed by the target pre-trained model.
[0075] The sample prompt word information is used to prompt the target pre-trained model to extract the target text elements contained in the sample text information, and generate predicted scientific hypothesis information according to the target text elements; the sample text information is the literature title information and / or literature abstract information of the sample literature; the target text elements include at least one of clinical problems, scientific problems, molecular targets, action mechanisms and interaction mechanisms.
[0076] For example, if the target text elements are clinical problems, scientific problems, molecular targets, action mechanisms and interaction mechanisms, the sample prompt word information can be selected as: you are a biomedical expert, please extract the following target text elements according to the literature title information and literature abstract information of the literature: clinical problems (such as "colorectal liver metastasis"), scientific problems (such as "epithelial mesenchymal transition"), molecular targets (such as "lncRNA ANCR"), action mechanisms (such as "Wnt-β Catenin signaling pathway"), variable roles (such as "upstream driving variable, main variable, interaction variable, effect variable"), and interaction mechanisms (for high-quality articles, such as "lncRNA ANCR binding protein TRIM21: RNA-protein binding"), and summarize into scientific hypothesis information.
[0077] The scientific hypothesis information refers to the verifiable conjectural conclusion about natural phenomena or scientific problems derived by the target pre-trained model based on the sample text information (literature title information and / or literature abstract information).
[0078] In an embodiment, the sample data set is input into the target pre-trained model, and the sample prompt word information and the sample text information in the sample data set are determined by the target pre-trained model. Further, the sample prompt word information and the sample text information are input into the target pre-trained model, and the original weight matrix of the target pre-trained model and the trainable weight matrix of the trainable adapter are used for forward propagation, and finally the predicted scientific hypothesis information is output by the target pre-trained model.
[0079] S103, determining the labeled scientific hypothesis information in the sample data set by the target pre-trained model, and training the trainable adapter according to the predicted scientific hypothesis information and the labeled scientific hypothesis information, for updating the trainable weight matrix to obtain an updated weight matrix.
[0080] The annotated scientific hypothesis information refers to standardized scientific hypothesis information manually annotated by professional annotators with business background in the sample data set.
[0081] In an embodiment, the annotated scientific hypothesis information in the sample data set is determined by the target pre-training model, the difference between the predicted scientific hypothesis information and the annotated scientific hypothesis information is quantified by a loss function, and a target loss value is obtained. Further, the original weight matrix of the target pre-training model is frozen, and only the trainable weight matrix of the trainable adapter is retained. According to the target loss value, back propagation is performed in the target pre-training model, the gradient is calculated by the optimizer, and the trainable weight matrix of the trainable adapter is updated until the training is completed, and the updated weight matrix of the trainable adapter is obtained.
[0082] S104, obtaining the scientific hypothesis extraction model of the target pre-training model after fine-tuning according to the updated weight matrix.
[0083] In an embodiment, a merged weight matrix is generated according to the updated weight matrix of the trainable adapter and the original weight matrix of the target pre-training model, and the scientific hypothesis extraction model of the target pre-training model after fine-tuning is obtained according to the merged weight matrix.
[0084] The present application fine-tunes the pre-training model to obtain a scientific hypothesis extraction model that can be used to extract scientific hypothesis information. The scientific hypothesis extraction model can automatically identify and structure scientific hypothesis information containing core elements such as clinical problems, scientific problems, molecular targets, mechanisms of action and interaction mechanisms from literature title information and / or literature abstract information, improving the accuracy and completeness of scientific hypothesis information extraction, and helping researchers quickly extract innovative points from literature and improving project design efficiency.
[0085] Optionally, the training set, validation set and test set are divided in the ratio of 8:1:1, which is the data division ratio confirmed by business background evaluators after testing multiple fine-tuned models, which can maximize the optimization of model learning scientific hypothesis extraction logic.
[0086] Optionally, the training configuration is: hardware requirement GPU: at least 1xRTX4090(24GB) or equivalent; memory: 32GB or more storage: 100GB or more SSD.
[0087] Hyperparameter settings: learning rate: 1E+5 to 5E+5; Batchsize: 4-16 (adjust according to GPU memory); Epochs: 3-5; sequence length: 2048 tokens.
[0088] Example two:
[0089] Figure 2A flowchart of a fine-tuning method of a pre-trained model provided for the second embodiment of the present application, the present embodiment further optimizes and extends the above-mentioned embodiments, and can be combined with the above-mentioned various optional embodiments. As shown in Figure 2 The method comprises the following steps:
[0090] S201. Determine at least one to-be-trained linear layer from the candidate linear layers contained in the target pre-trained model, and inject a trainable adapter into the to-be-trained linear layer, and input a sample data set into the target pre-trained model.
[0091] Among them, the sample data set contains at least one type of candidate structured data.
[0092] S202. Determine the structured role label corresponding to each candidate structured data respectively; obtain the candidate structured data with the structured role label as the system role as the sample prompt word information; and obtain the candidate structured data with the structured role label as the user role as the sample text information.
[0093] Among them, the structured role label refers to a classification mark in the structured data for identifying the function role to which the candidate structured data belongs, and its core function is to standardize the definition of the identity of the candidate structured data in system interaction. The structured role label includes a system role, a user role, and an assistant role. Among them, the system role, also known as the System role, refers to structured data automatically generated or preset by the system, mainly used to define process rules, provide context guidance, or control interaction logic. The user role, also known as the User role, refers to structured data directly input by the user or generated through interaction behavior, representing user intent, request, or feedback. The assistant role, also known as the Assistant role, refers to the response content generated by the model, which executes the tasks set by the system role and responds to user needs.
[0094] In one embodiment, the structured role label corresponding to each candidate structured data is determined, the candidate structured data with the structured role label as the system role in each candidate structured data is obtained as the sample prompt word information, and the candidate structured data with the structured role label as the user role in each candidate structured data is obtained as the sample text information.
[0095] By determining the structured role label corresponding to each candidate structured data respectively, obtaining the candidate structured data with the structured role label as the system role as the sample prompt word information, and obtaining the candidate structured data with the structured role label as the user role as the sample text information, the beneficial effects are as follows:
[0096] First, based on the structured role label, the sample text information (user data) and the sample prompt word information (system data) are automatically separated, feature confusion is eliminated, and the accuracy of scientific hypothesis information prediction is improved.
[0097] In a second aspect, the system role data (sample prompt information) has high reusability by nature, and can be directly cached and reused after extraction, reducing the cost of repeated cleaning. The user role data (sample text information) is dynamically loaded on demand, avoiding full data traversal and reducing resource consumption.
[0098] S203, according to the sample prompt information and the sample text information, output the predicted scientific hypothesis information, and obtain the candidate structured data with the structured role label of the helper role as the labeled scientific hypothesis information.
[0099] In an embodiment, the structured role label corresponding to each candidate structured data is determined, and the candidate structured data with the structured role label of the helper role in each candidate structured data is taken as the labeled scientific hypothesis information.
[0100] By obtaining the candidate structured data with the structured role label of the helper role as the labeled scientific hypothesis information, the beneficial effects are:
[0101] In the traditional model fine-tuning, both the system role data and the user role data are fitted by the model, which cannot focus on learning the answer content of the helper role data. By learning only the candidate structured data with the structured role label of the helper role during training, and not fitting the system role data or the user role data, the model can focus on learning "how to answer", thereby generating high-quality structured scientific hypothesis information.
[0102] S204, determine at least one candidate word element contained in the labeled scientific hypothesis information, and determine candidate mask information added to each candidate word element.
[0103] The candidate word element refers to all word elements contained in the labeled scientific hypothesis information. Token, also known as Token, refers to the smallest semantic unit that may carry the core semantics of the hypothesis in the labeled scientific hypothesis information, such as keywords, phrases, or variables. The candidate mask information refers to the binary mask information added to each candidate word element, which is used to make the target pre-training model aware of the "candidate word element" that needs to be learned, i.e. adding a type of candidate mask information to the "candidate word element" that the target pre-training model does not need to learn, such as adding 0 mask information; adding another type of candidate mask information to the "candidate word element" that the target pre-training model needs to learn, such as adding 1 mask information.
[0104] In an embodiment, at least one candidate word element contained in the labeled scientific hypothesis information is determined, and the candidate mask information added to each candidate word element is determined by traversing the token information of each candidate word element.
[0105] S205, obtain the candidate tokens of the candidate mask information as target masks as the to-be-learned tokens, and generate to-be-learned scientific hypothesis information according to each to-be-learned token; and train the trainable adapter according to the predicted scientific hypothesis information and the to-be-learned scientific hypothesis information.
[0106] The to-be-learned token refers to a candidate token selected from the candidate tokens, which has potential semantic value and needs to be cognitively associated through adapter training.
[0107] In an embodiment, the candidate tokens of the candidate mask information as target masks are determined as to-be-learned tokens, such as the candidate tokens of the candidate mask information as 1 mask information, and to-be-learned scientific hypothesis information is generated according to the combination results of each to-be-learned token. Further, the trainable adapter is trained according to the predicted scientific hypothesis information and the to-be-learned scientific hypothesis information.
[0108] By determining at least one candidate token contained in the labeled scientific hypothesis information, and determining the candidate mask information added to each candidate token, obtaining the candidate tokens of the candidate mask information as target masks as to-be-learned tokens, and generating to-be-learned scientific hypothesis information according to each to-be-learned token, and training the trainable adapter according to the predicted scientific hypothesis information and the to-be-learned scientific hypothesis information, the beneficial effects are:
[0109] The traditional model fine-tuning does not pay attention to tokens, which can easily mislead the training. By adding candidate mask information to each candidate token, the model can mark which candidate tokens are the objects that the model needs to pay attention to, usually the non-padding part, to help the model perform effective attention calculation.
[0110] Optionally, the training of the trainable adapter according to the predicted scientific hypothesis information and the to-be-learned scientific hypothesis information includes:
[0111] According to the predicted scientific hypothesis information and the to-be-learned scientific hypothesis information, loss value calculation is performed to determine a target loss value, and the trainable adapter is trained according to the target loss value.
[0112] In an embodiment, a preset loss function is used to perform loss value calculation according to the predicted scientific hypothesis information and the to-be-learned scientific hypothesis information to determine a target loss value. Further, the target loss value is back-propagated in the target pre-training model, the gradient is calculated by an optimizer, and the trainable weight matrix of the trainable adapter is updated.
[0113] The target loss value is determined by calculating the loss value according to the predicted scientific hypothesis information and the to-be-learned scientific hypothesis information; and the trainable adapter is trained according to the target loss value, which has the beneficial effect that: by simultaneously using the predicted scientific hypothesis information and the to-be-learned scientific hypothesis information to calculate the loss value, the model can more comprehensively capture the potential laws in the scientific hypothesis information, and improve the adaptability of the model to complex scientific hypothesis information.
[0114] In S206, information similarity between the predicted scientific hypothesis information and the labeled scientific hypothesis information is determined, and element completeness and element accuracy of text elements contained in the predicted scientific hypothesis information are determined; wherein the information similarity is text similarity and / or semantic similarity.
[0115] The information similarity is used to measure the surface similarity between the predicted scientific hypothesis information and the labeled scientific hypothesis information. The element completeness refers to whether the predicted scientific hypothesis information contains necessary text elements required by the scientific hypothesis, and the coverage comprehensiveness of these text elements. The element accuracy refers to the correctness of each text element in the predicted scientific hypothesis information, that is, whether it conforms to scientific facts, logical consistency or the original intention of the labeled hypothesis.
[0116] In an embodiment, a text similarity algorithm is used to determine the text similarity between the predicted scientific hypothesis information and the labeled scientific hypothesis information, and a semantic similarity algorithm is used to determine the semantic similarity between the predicted scientific hypothesis information and the labeled scientific hypothesis information. Moreover, the element completeness and the element accuracy of the text elements contained in the predicted scientific hypothesis information are determined.
[0117] In S207, in a case where the information similarity, the element completeness and the element accuracy all satisfy respective threshold values, a scientific hypothesis extraction model of the target pre-training model after fine-tuning is obtained according to the updated weight matrix.
[0118] In an embodiment, in a case where the information similarity satisfies a corresponding first threshold value, the element completeness satisfies a corresponding second threshold value, and the element accuracy satisfies a corresponding third threshold value, a scientific hypothesis extraction model of the target pre-training model after fine-tuning is obtained according to the updated weight matrix.
[0119] By determining the information similarity between the predicted scientific hypothesis information and the labeled scientific hypothesis information, and determining the element completeness and the element accuracy of the text elements contained in the predicted scientific hypothesis information, in a case where the information similarity, the element completeness and the element accuracy all satisfy respective threshold values, a scientific hypothesis extraction model of the target pre-training model after fine-tuning is obtained according to the updated weight matrix, which has the beneficial effect that: through the information similarity evaluation, the core logical structure of the scientific hypothesis information can be effectively captured; through the dual evaluation mechanism of the element completeness and the element accuracy, the integrity of the key components of the scientific hypothesis information is ensured.
[0120] Optionally, the scientific hypothesis extraction model after fine-tuning of the target pre-training model according to the updated weight matrix comprises:
[0121] The original weight matrix corresponding to the target pre-training model is obtained, and the weight matrix merging is performed according to the original weight matrix and the updated weight matrix to generate a merged weight matrix; and the scientific hypothesis extraction model after fine-tuning of the target pre-training model is obtained according to the merged weight matrix.
[0122] For example, the original weight matrix is W0, the updated weight matrix is W1, and the merged weight matrix is W0+W1. The scientific hypothesis extraction model after fine-tuning of the target pre-training model is obtained according to the merged weight matrix W0+W1.
[0123] By obtaining the original weight matrix corresponding to the target pre-training model, and performing weight matrix merging according to the original weight matrix and the updated weight matrix to generate a merged weight matrix, and obtaining the scientific hypothesis extraction model after fine-tuning of the target pre-training model according to the merged weight matrix, the beneficial effect is that the trainable parameter amount can be reduced in the training stage by adjusting only the low-rank incremental weight instead of the full amount of parameters.
[0124] Optionally, the sample data set is input into the target pre-training model, comprising:
[0125] The input data length threshold value corresponding to the target pre-training model is determined, and the sample data length corresponding to the sample data set is determined, and the input data length threshold value and the sample data length are compared; in the case that the sample data length is less than or equal to the input data length threshold value, the sample data set is input into the target pre-training model.
[0126] For example, the input data length threshold value is 100Tokens, and the sample data length is 990Tokens, and the sample data set is input into the target pre-training model.
[0127] By determining the input data length threshold value corresponding to the target pre-training model, and determining the sample data length corresponding to the sample data set, and comparing the input data length threshold value and the sample data length, and in the case that the sample data length is less than or equal to the input data length threshold value, the sample data set is input into the target pre-training model, which can control the training cost, prevent the memory from exploding, and also avoid the error caused by the truncated model input.
[0128] Embodiment three:
[0129] Figure 3A flowchart of a scientific hypothesis information extraction method provided for the third embodiment of the present application. The present embodiment can be applied to automatically identifying and structuring scientific hypothesis information containing core elements such as clinical problems, scientific problems, molecular targets, action mechanisms, and interaction mechanisms from literature title information and / or literature abstract information using a scientific hypothesis extraction model. The method can be executed by a scientific hypothesis information extraction device, which can be realized in the form of hardware and / or software. As shown in Figure 3 , the method comprises:
[0130] S301, obtaining current prompt word information and current text information.
[0131] S302, inputting the current prompt word information and the current text information into the scientific hypothesis extraction model, and outputting target scientific hypothesis information from the scientific hypothesis extraction model according to the current prompt word information and the current text information.
[0132] The scientific hypothesis extraction model is obtained by the fine-tuning method of the pre-training model disclosed in the present embodiment. The current prompt word information is used to prompt the scientific hypothesis extraction model to extract target text elements contained in the current text information, and to generate target scientific hypothesis information according to the target text elements. The current text information is the literature title information and / or the literature abstract information of the current literature. The target text elements include at least one of the clinical problems, the scientific problems, the molecular targets, the action mechanisms, and the interaction mechanisms.
[0133] The scientific hypothesis extraction model can automatically identify and structure scientific hypothesis information containing core elements such as clinical problems, scientific problems, molecular targets, action mechanisms, and interaction mechanisms from literature title information and / or literature abstract information, thereby improving the accuracy and completeness of scientific hypothesis information extraction, and helping researchers quickly extract innovative points from literature and improving the efficiency of project design.
[0134] Embodiment Four
[0135] Figure 4 A structural diagram of a pre-training model fine-tuning device provided for the fourth embodiment of the present application. The device can be applied to obtaining a scientific hypothesis extraction model by model fine-tuning, for automatically identifying and structuring scientific hypothesis information containing core elements such as clinical problems, scientific problems, molecular targets, action mechanisms, and interaction mechanisms from literature title information and / or literature abstract information, as shown in Figure 4 , the device comprises:
[0136] The adapter injection module 41 is configured to determine at least one to-be-trained linear layer from candidate linear layers contained in the target pre-training model, and inject a trainable adapter into the to-be-trained linear layer; wherein the trainable adapter has a trainable weight matrix.
[0137] The scientific hypothesis information prediction module 42 is configured to input a sample data set into the target pre-training model, determine sample prompt word information and sample text information in the sample data set by the target pre-training model, and output predicted scientific hypothesis information according to the sample prompt word information and the sample text information; wherein the sample prompt word information is used to prompt the target pre-training model to extract target text elements contained in the sample text information, and generate the predicted scientific hypothesis information according to the target text elements; the sample text information is literature title information and / or literature abstract information of a sample literature; and the target text elements include at least one of a clinical problem, a scientific problem, a molecular target, an action mechanism and an interaction mechanism.
[0138] The adapter training module 43 is configured to determine labeled scientific hypothesis information in the sample data set by the target pre-training model, and train the trainable adapter according to the predicted scientific hypothesis information and the labeled scientific hypothesis information, so as to update the trainable weight matrix to obtain an updated weight matrix.
[0139] The scientific hypothesis extraction model generation module 44 is configured to obtain a scientific hypothesis extraction model of the target pre-training model after fine-tuning according to the updated weight matrix.
[0140] Optionally, the sample data set contains at least one type of candidate structured data.
[0141] The scientific hypothesis information prediction module 42 is specifically configured to:
[0142] Determine structured role labels corresponding to each of the candidate structured data respectively;
[0143] Obtain the candidate structured data whose structured role label is a system role as the sample prompt word information;
[0144] Obtain the candidate structured data whose structured role label is a user role as the sample text information.
[0145] Optionally, the scientific hypothesis information prediction module 42 is specifically further configured to:
[0146] Obtain the candidate structured data whose structured role label is an assistant role as the labeled scientific hypothesis information.
[0147] Optionally, the adapter training module 43 is specifically configured to:
[0148] determine at least one candidate word element contained in the labeled scientific hypothesis information, and determine candidate mask information added to each of the candidate word elements;
[0149] obtain the candidate word element with the target mask as the candidate mask information as a to-be-learned word element, and generate to-be-learned scientific hypothesis information according to each of the to-be-learned word elements;
[0150] train the trainable adapter according to the predicted scientific hypothesis information and the to-be-learned scientific hypothesis information.
[0151] Optionally, the adapter training module 43 is specifically configured to:
[0152] perform loss value calculation according to the predicted scientific hypothesis information and the to-be-learned scientific hypothesis information to determine a target loss value;
[0153] train the trainable adapter according to the target loss value.
[0154] Optionally, the scientific hypothesis extraction model generation module 44 is specifically configured to:
[0155] determine an information similarity between the predicted scientific hypothesis information and the labeled scientific hypothesis information, and determine an element completeness and an element accuracy of a text element contained in the predicted scientific hypothesis information; wherein the information similarity is a text similarity and / or a semantic similarity;
[0156] in a case where the information similarity, the element completeness and the element accuracy all satisfy respective threshold values, obtain the scientific hypothesis extraction model after fine-tuning of the target pre-training model according to the updated weight matrix.
[0157] Optionally, the scientific hypothesis extraction model generation module 44 is specifically configured to:
[0158] obtain an original weight matrix corresponding to the target pre-training model, and perform weight matrix merging according to the original weight matrix and the updated weight matrix to generate a merged weight matrix;
[0159] obtain the scientific hypothesis extraction model after fine-tuning of the target pre-training model according to the merged weight matrix.
[0160] Optionally, the scientific hypothesis information prediction module 42 is specifically configured to:
[0161] determine an input data length threshold value corresponding to the target pre-training model, and determine a sample data length corresponding to the sample data set, and compare the input data length threshold value and the sample data length;
[0162] In a case where the sample data length is less than or equal to the input data length threshold value, input the sample data set into the target pre-training model.
[0163] The pre-training model fine-tuning device provided in Embodiment Four of the present application can perform the pre-training model fine-tuning method provided in the present application, and has the corresponding function modules and beneficial effects of the execution method.
[0164] Embodiment Five
[0165] Figure 5 A structural schematic diagram of a scientific hypothesis information extraction device provided in Embodiment Five of the present application, which can be applied to automatically identify and structure scientific hypothesis information containing core elements such as clinical problems, scientific problems, molecular targets, action mechanisms and interaction mechanisms from literature title information and / or literature abstract information by using a scientific hypothesis extraction model, as shown in the figure. Figure 5 The device comprises:
[0166] An information acquisition module 51 is configured to acquire current prompt word information and current text information.
[0167] A scientific hypothesis information extraction module 52 is configured to input the current prompt word information and the current text information into a scientific hypothesis extraction model, and output target scientific hypothesis information by the scientific hypothesis extraction model according to the current prompt word information and the current text information.
[0168] The scientific hypothesis extraction model is obtained by using the pre-training model fine-tuning method disclosed in the present application.
[0169] The current prompt word information is used to prompt the scientific hypothesis extraction model to extract target text elements contained in the current text information, and generate the target scientific hypothesis information according to the target text elements; the current text information is literature title information and / or literature abstract information of a current literature; and the target text elements include at least one of clinical problems, scientific problems, molecular targets, action mechanisms and interaction mechanisms.
[0170] The scientific hypothesis information extraction device provided in Embodiment Five of the present application can perform the scientific hypothesis information extraction method provided in the present application, and has the corresponding function modules and beneficial effects of the execution method.
[0171] According to embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0172] Embodiment six:
[0173] Figure 6 A structural schematic diagram of an electronic device 60 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present application described and / or claimed in this document.
[0174] As shown in Figure 6 The electronic device 60 includes at least one processor 61, and a memory, such as a read-only memory (ROM) 62, a random access memory (RAM) 63, etc., connected to the at least one processor 61 in communication, where the memory stores computer programs executable by the at least one processor. The processor 61 can perform various appropriate actions and processes according to the computer programs stored in the read-only memory (ROM) 62 or loaded into the random access memory (RAM) 63 from the storage unit 68. In the RAM 63, various programs and data required for the operation of the electronic device 60 can also be stored. The processor 61, the ROM 62, and the RAM 63 are connected to each other through a bus 64. An input / output (I / O) interface 65 is also connected to the bus 64.
[0175] A plurality of components in the electronic device 60 are connected to the I / O interface 65, including: an input unit 66, such as a keyboard, a mouse, etc.; an output unit 67, such as various types of displays, speakers, etc.; a storage unit 68, such as a magnetic disk, an optical disk, etc.; and a communication unit 69, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 69 allows the electronic device 60 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunications networks.
[0176] The processor 61 can be various general and / or special purpose processing components having processing and computing capabilities. Some examples of the processor 61 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 61 performs various methods and processes described above, such as the fine-tuning method of a pre-trained model and / or the extraction method of scientific hypothesis information.
[0177] In some embodiments, the fine-tuning method of a pre-trained model and / or the extraction method of scientific hypothesis information can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 68. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 60 via the ROM 62 and / or the communication unit 69. When the computer program is loaded onto the RAM 63 and executed by the processor 61, one or more steps of the fine-tuning method of a pre-trained model and / or the extraction method of scientific hypothesis information described above can be performed. Alternatively, in other embodiments, the processor 61 can be configured to perform the fine-tuning method of a pre-trained model and / or the extraction method of scientific hypothesis information by any other appropriate means, such as by means of firmware.
[0178] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0179] Computer programs for implementing the methods of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer program, when executed, enables the functions / acts specified in the flowcharts and / or block diagrams to be implemented. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a standalone software package and partially on a remote machine or entirely on a remote machine or server.
[0180] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0181] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0182] The systems and techniques described herein can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described herein, or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0183] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.
[0184] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be executed in parallel, executed in series, or executed in different orders, as long as the desired results of the technical solutions of the present disclosure are achieved, and the present disclosure is not limited herein.
[0185] The specific embodiments described above are not intended to limit the scope of the present disclosure. Those skilled in the art will understand that various modifications, combinations, sub-combinations, and alternatives can be made to the specific embodiments without departing from the spirit and principles of the present disclosure. Any further modifications, equivalents, and / or alternatives come within the scope of the present disclosure as set forth in the following claims.
Claims
1. A fine-tuning method of a pre-trained model, characterized in that, The method comprises: determining at least one linear layer to be trained from candidate linear layers included in a target pre-training model, and injecting a trainable adapter in the linear layer to be trained; wherein the trainable adapter has a trainable weight matrix; inputting a sample data set into the target pre-training model, determining sample prompt information and sample text information in the sample data set by the target pre-training model, and outputting predicted scientific hypothesis information according to the sample prompt information and the sample text information; wherein the sample prompt information is used to prompt the target pre-training model to extract target text elements included in the sample text information, and generate the predicted scientific hypothesis information according to the target text elements; the sample text information is literature title information and / or literature abstract information of a sample literature; the target text elements include at least one of clinical problems, scientific problems, molecular targets, mechanisms of action and interaction mechanisms; the predicted scientific hypothesis information refers to a verifiable speculative conclusion about a natural phenomenon or a scientific problem derived by the target pre-training model based on the sample text information; determining labeled scientific hypothesis information in the sample data set by the target pre-training model, and training the trainable adapter according to the predicted scientific hypothesis information and the labeled scientific hypothesis information, for updating the trainable weight matrix to obtain an updated weight matrix; obtaining a scientific hypothesis extraction model of the target pre-training model after fine-tuning according to the updated weight matrix.
2. The method of claim 1, wherein, The sample data set includes at least one type of candidate structured data; The determination of the sample prompt information and the sample text information in the sample data set by the target pre-training model comprises: determining structured role labels corresponding to each of the candidate structured data respectively; obtaining the candidate structured data whose structured role label is a system role as the sample prompt information; obtaining the candidate structured data whose structured role label is a user role as the sample text information.
3. The method of claim 2, wherein, The determination of the labeled scientific hypothesis information in the sample data set by the target pre-training model comprises: obtaining the candidate structured data whose structured role label is an assistant role as the labeled scientific hypothesis information.
4. The method according to claim 1, wherein The training of the trainable adapter according to the predicted scientific hypothesis information and the labeled scientific hypothesis information comprises: determining at least one candidate word element included in the labeled scientific hypothesis information, and determining candidate mask information added to each of the candidate word elements respectively; obtaining the candidate word element whose candidate mask information is a target mask as a to-be-learned word element, and generating to-be-learned scientific hypothesis information according to each of the to-be-learned word elements; training the trainable adapter according to the predicted scientific hypothesis information and the to-be-learned scientific hypothesis information.
5. The method of claim 4, wherein, The training of the trainable adapter according to the predicted scientific hypothesis information and the to-be-learned scientific hypothesis information comprises: The loss value is calculated according to the predicted scientific hypothesis information and the scientific hypothesis information to be learned, and a target loss value is determined; The trainable adapter is trained according to the target loss value.
6. The method of claim 1, wherein, The scientific hypothesis extraction model after fine-tuning of the target pre-training model according to the updated weight matrix comprises: The information similarity between the predicted scientific hypothesis information and the labeled scientific hypothesis information is determined, and the element completeness and element accuracy of the text elements contained in the predicted scientific hypothesis information are determined; wherein the information similarity is a text similarity and / or a semantic similarity; In the case that the information similarity, the element completeness and the element accuracy all satisfy their respective threshold values, the scientific hypothesis extraction model after fine-tuning of the target pre-training model according to the updated weight matrix is obtained.
7. The method of claim 6, wherein, The scientific hypothesis extraction model after fine-tuning of the target pre-training model according to the updated weight matrix comprises: An original weight matrix corresponding to the target pre-training model is obtained, and a weight matrix merging is performed according to the original weight matrix and the updated weight matrix to generate a merged weight matrix; The scientific hypothesis extraction model after fine-tuning of the target pre-training model according to the merged weight matrix is obtained.
8. The method of claim 1, wherein, The method comprises: An input data length threshold value corresponding to the target pre-training model is determined, and a sample data length corresponding to the sample data set is determined, and the input data length threshold value and the sample data length are compared; In the case that the sample data length is less than or equal to the input data length threshold value, the sample data set is input into the target pre-training model.
9. A method of extracting scientific hypothesis information, characterized by, The method comprises: Obtaining current prompt word information and current text information; The current prompt word information and the current text information are input into a scientific hypothesis extraction model, and the scientific hypothesis extraction model outputs target scientific hypothesis information according to the current prompt word information and the current text information; The scientific hypothesis extraction model is obtained by fine-tuning the pre-training model of any one of claims 1-8; The current prompt word information is used to prompt the scientific hypothesis extraction model to extract target text elements contained in the current text information, and generate the target scientific hypothesis information according to the target text elements; the current text information is literature title information and / or literature abstract information of a current literature; the target text elements include at least one of a clinical problem, a scientific problem, a molecular target, an action mechanism and an interaction mechanism.
10. A fine-tuning device for a pre-trained model, characterized in that: The device comprises: An adapter injection module for determining at least one trainable linear layer from candidate linear layers contained in a target pre-training model, and injecting a trainable adapter into the trainable linear layer; wherein the trainable adapter has a trainable weight matrix. A scientific hypothesis information prediction module is used to input a sample data set into the target pre-training model, determine the sample prompt word information and sample text information in the sample data set through the target pre-training model, and output predicted scientific hypothesis information based on the sample prompt word information and the sample text information; wherein the sample prompt word information is used to prompt the target pre-training model to extract the target text elements contained in the sample text information, and generate the predicted scientific hypothesis information based on the target text elements; the sample text information is the document title information and / or document abstract information of the sample document; the target text elements include at least one of clinical problems, scientific problems, molecular targets, mechanisms of action and interaction mechanisms; the predicted scientific hypothesis information refers to the verifiable speculative conclusions about natural phenomena or scientific problems derived by the target pre-training model based on the sample text information; an adapter training module, configured to determine the labeled scientific hypothesis information in the sample data set through the target pre-trained model, and to train the trainable adapter based on the predicted scientific hypothesis information and the labeled scientific hypothesis information to update the trainable weight matrix to obtain an updated weight matrix; A scientific hypothesis extraction model generation module is used to obtain a scientific hypothesis extraction model after fine-tuning the target pre-training model according to the updated weight matrix.
11. A scientific hypothesis information extraction device characterized by comprising: The device comprises: Information acquisition module, used to obtain current prompt word information and current text information; A scientific hypothesis information extraction module, configured to input the current prompt word information and the current text information into a scientific hypothesis extraction model, and output target scientific hypothesis information based on the current prompt word information and the current text information through the scientific hypothesis extraction model; Wherein, the scientific hypothesis extraction model is obtained by fine-tuning the pre-training model described in any one of claims 1 to 8; The current prompt word information is used to prompt the scientific hypothesis extraction model to extract the target text elements contained in the current text information, and generate the target scientific hypothesis information based on the target text elements; the current text information is the document title information and / or document abstract information of the current document; the target text elements include at least one of clinical problems, scientific problems, molecular targets, mechanisms of action and interaction mechanisms.
12. An electronic device, comprising: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to execute the method according to any one of claims 1 to 9.
14. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Large model-based medical text information governance method and system
CN118114718A
Medical knowledge relation extraction method and system based on large language model fine tuning and retrieval enhancement generation
CN118569263A