Text processing method, system and device based on knowledge graph and large language model
By constructing a compliance knowledge graph for human genetic resources and using large language models to process texts, it solves the problem that scientific researchers find it difficult to accurately judge the type of application when applying for human genetic resources, and achieves rapid and accurate application guidance, reducing the application risks and costs.
Patent Information
- Application Number
- CN202411978022.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
AI Technical Summary
Before applying for a biomedical research project, it is difficult for scientific researchers to accurately determine whether human genetic resources need to be applied for and what type of application is required, resulting in problems such as unreported and incorrect declaration.
Text processing methods based on knowledge graphs and large language models are used to construct human genetic resources compliance knowledge graphs and construct large language model prompt words based on the text input by users, and query pre-trained large language models to obtain professional answers to the types of human genetic resources declaration.
It has achieved rapid and accurate acquisition of professional answers to the types of human genetic resources application, reducing the application risks of scientific researchers and saving time and labor costs.
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computers, and in particular to a text processing method, system and device based on a knowledge graph and a large language model. Background Art
[0002] The stage before the application for a biomedical research project is one of the main risk stages in the management of the project. For example, when scientific research activities involve human genetic resources, due to the rapid update of relevant laws, regulations and rules and regulations, researchers find it difficult to accurately understand and grasp whether the scientific research activities to be carried out need to be reported for human genetic resources and what type of application to be made. This leads to problems such as "failure to report when required", "rejection of the wrong application type", and "spending a lot of manpower and material resources on preliminary preparations for application without application". However, the prior art does not disclose how to guide users to predict the type of project application before the start of scientific research activities. Since the content related to scientific research projects is mainly presented in text form, this application provides a corresponding text processing method based on knowledge graphs and large language models. Summary of the invention
[0003] The present invention provides a text processing method based on a knowledge graph and a large language model to solve the problem that scientific researchers are unable to efficiently and accurately determine whether activities involving human genetic resources need to be reported and what type of reporting is required. The present invention queries the human genetic resource compliance knowledge graph based on the pre-trained large language model according to the large language model prompt words, and uses the pre-trained large language model to quickly and accurately obtain professional answers to the types of human genetic resource reporting.
[0004] The present invention provides a text processing method based on a knowledge graph and a large language model, comprising: constructing a human genetic resource compliance knowledge graph based on an acquired human genetic resource corpus; constructing a large language model prompt word based on a user input text; querying the human genetic resource compliance knowledge graph based on a pre-trained large language model according to the large language model prompt word, and obtaining professional answers related to human genetic resource declaration; according to a text processing method based on a knowledge graph and a large language model provided by the present invention, the human genetic resource corpus includes a human genetic resource legal and regulatory set, a human genetic resource declaration form template, management regulations of different units on human genetic resources, human genetic resource-related documents and general corpus.
[0005] The human genetic resources compliance knowledge graph is a set of triples in the format of (p, r, l), where p, l∈ in is the entity vocabulary, It is a set of predefined relationships, where p represents the head entity, l represents the tail entity, and r represents the relationship between entities.
[0007] According to a text processing method based on a knowledge graph and a large language model provided by the present invention, before constructing the human genetic resource compliance knowledge graph, it also includes: annotating the human genetic resource corpus to obtain annotated human genetic resource corpus; cleaning and preprocessing the annotated human genetic resource corpus to obtain a usable human genetic resource corpus, so as to construct the human genetic resource compliance knowledge graph based on the usable human genetic resource corpus.
[0008] According to a text processing method based on a knowledge graph and a large language model provided by the present invention, it also includes: iterating the knowledge graph based on a preset graph evolution algorithm; the preset graph evolution algorithm is an algorithm that is a combination of one or more of verification, improvement, regular updating, maintenance and privacy protection.
[0009] The step of constructing a large language model prompt word based on the user input text includes:
[0010] Get the original text and use it as N 0 Level prompt words;
[0011] The prompt word model is used to optimize the original problem and the optimized N 1 Level prompt words;
[0012] Use dialogue to 1 The level prompt word is fed back to the user, and the user debugs and optimizes to obtain N 2 Level prompt words; continue to iterate until the user is satisfied with N final Prompt word, which is used as the prompt word of the large language model, where final is greater than or equal to 0.
[0013] Specifically, the original question comes from voice, text or picture content, and a plurality of original question templates may be pre-configured to provide guidance for the original question input by the user.
[0014] The N 1 Level to N final The prompt words of the level include fixed words and placeholders.
[0015] The prompt word model includes a prompt word template, which can be used to construct prompt words, including a fixed word part and a placeholder. The fixed word is used to guide the user to input the corresponding placeholder, and the placeholder is used to describe the prompt word activity type to be constructed and the content involved in each activity type. According to different types of activities, the placeholder in the prompt word module is variable.
[0016] Specifically, the placeholders include one or more of activity type, research type, participating units, implementation methods, human genetic resource types, and population types.
[0017] The prompt word template can be divided into five types according to different activity types. The number of placeholders in the five activity types is different, and each placeholder is assigned a value of 1. The five prompt word templates correspond to the five activity types respectively. The placeholder assignments in the five prompt word templates are added to obtain the set assignment sum of the prompt word template. The sum of the assignments of the five prompt word templates is a fixed value. The five activity types are scientific research, public data, transmission data, human material export and sample library construction, among which the placeholder assignment sum of the scientific research activity template is 5, the placeholder assignment sum of the public data activity template is 2, the placeholder assignment sum of the transmission data activity template is 3, the placeholder assignment sum of the human material export activity template is 2, and the placeholder assignment sum of the sample library construction activity template is 2.
[0018] The prompt word model also includes a text segmentation algorithm, which identifies keywords in the original input text and matches the keywords with placeholders. If there is a keyword corresponding to the placeholder, the keyword is assigned a value of 1, otherwise it is 0. If the original input text does not include a keyword that can correspond to the placeholder, the original input text is assigned a value of 0. The total assignment of keywords in the original input text is calculated; if the assignment sum is lower than the set value sum of the corresponding activity type prompt word template, the user is prompted to continue debugging the prompt word until the prompt word assignment is greater than or equal to the set assignment sum of the corresponding template.
[0019] According to a text processing method based on a knowledge graph and a large language model provided by the present invention, the method queries the human genetic resource compliance knowledge graph based on a pre-trained large language model according to the large language model prompt words to obtain an answer to the user input text, including: inputting the large language model prompt words into the pre-trained large language model; using the knowledge graph to perform relational reasoning and query on the user input text to obtain a professional answer to the human genetic resource declaration type of the user input text output by the pre-trained large language model.
[0020] According to a text processing method based on a knowledge graph and a large language model provided by the present invention, it also includes: iterating the pre-trained large language model based on a preset model evolution algorithm; the preset model evolution algorithm is an algorithm that is a combination of one or more of regular updates, security, privacy reviews, and compliance reviews.
[0021] The present invention also provides a text processing method recommendation system based on a knowledge graph and a large language model, comprising: a knowledge graph module, used to construct a human genetic resource compliant knowledge graph based on an acquired human genetic resource corpus; a prompt word module, used to construct a large language model prompt word based on a user input text; and a query module, used to query the human genetic resource compliant knowledge graph based on a pre-trained large language model according to the large language model prompt word to obtain an answer to the user input text.
[0022] Corresponding to the method, the present invention provides an electronic device, comprising: a memory and a processor; wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the above method.
[0023] The memory can store instructions for the processor to control the operation process. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0024] The processor may be an integrated circuit chip with data processing capabilities. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc. The methods, steps and logic block diagrams disclosed in the embodiments of the present invention may be implemented or executed. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0025] The processor may call instructions stored in the memory to control related operations. According to one embodiment, the memory stores instructions for the processor to control the following operations: construct a human genetic resource compliance knowledge graph based on the acquired human genetic resource corpus; ask the user to enter text, and construct a large language model prompt based on the user input text; according to the large language model prompt, query the human genetic resource compliance knowledge graph based on the pre-trained large language model to obtain professional answers related to human genetic resource declaration.
[0026] Beneficial effects of the present invention:
[0027] (1)Easy to use:
[0028] The declaration of human genetic resources projects is based on relevant laws, policies and management systems, such as the Regulations on the Administration of Human Genetic Resources (hereinafter referred to as the "Regulations"), the Implementation Rules of the Regulations on the Administration of Human Genetic Resources (hereinafter referred to as the "Rules"), and the management regulations of different units on human genetic resources. There is a certain correlation between the various conditions, and with the changes in laws, policies and management systems, the declaration conditions also change accordingly. According to the present invention, only the prompt text needs to be entered to obtain various information related to the human genetic resources project and the declaration, and guide the project to declare in compliance.
[0029] (2) Reduce costs:
[0030] The Regulations and Rules use legal language, which is often difficult for people engaged in biomedical research to accurately interpret, resulting in misreporting or failure to report, loss of time and labor costs, and possible risks of breaking the law. The present invention guides researchers on how to conduct scientific research activities involving human genetic resources in compliance with regulations, which can save a lot of time and labor costs. DETAILED DESCRIPTION
[0031] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be described clearly and completely below. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0032] The present invention provides a text processing method based on knowledge graph and large language model, including: collecting human genetic resource corpus as a data source, performing preprocessing operations based on the data source, and then extracting entities and relationships from the preprocessed data source to construct a human genetic resource compliance knowledge graph; generating large language model prompt words based on user input text; querying the knowledge graph to obtain professional answers about human genetic resource project declarations. The user data source is unstructured text.
[0033] 1. Construct a knowledge graph of human genetic resource compliance based on the acquired human genetic resource corpus.
[0034] The human genetic resources compliance knowledge graph is a set of triplets in the format of (p, r, l). in is the entity vocabulary, is a set of predefined relationships, p represents the head entity, l represents the tail entity, and r represents the relationship between entities. As a preferred embodiment, the steps of constructing a knowledge graph include:
[0035] (I) Node Generation
[0036] The unstructured input text is input into the extraction model and database to obtain pre-annotated data, which includes annotated data and model prediction data. The pre-annotated data is then manually annotated to obtain the required annotated data. The annotated data is cleaned and preprocessed to obtain usable data to build a knowledge graph.
[0037] Specifically, collect compliant human genetic resources corpus, and then use the annotation system to annotate the corpus, including marking the activity type, research type, participating units, implementation methods, human genetic resource type, population information, etc. in the corpus. Clean and preprocess the annotated data (text cleaning, word segmentation, tagging and annotation, etc.) to ensure the accuracy and consistency of the annotation. Perform annotation quality control, including verifying the work of the annotators, resolving disputes or errors in the annotation, and ensuring the availability and quality of the annotated data.
[0038] The pre-labeling process solves the problems of entity nesting and relationship overlap based on a multi-layer pointer network and a multi-head attention mechanism. First, the input text obtains a word embedding vector through the BERT pre-training model; secondly, the entity labeling method based on the multi-layer pointer network is used to identify entity boundaries and entity types. The multi-layer pointer network has 2M (M is the number of entity types) row annotation sequences, which respectively represent the start and end positions of entities belonging to a specific type, and the entities are obtained through the labels of the start and end positions; then, the output of the BERT model and the embedding vector of the entity label are concatenated to generate a multi-head relationship matrix. The relationship matrix is divided into M L*L two-dimensional matrices, where the rows represent the head entity and the columns represent the tail entity. The multi-head attention mechanism only labels the last tag of the entity; finally, the probability of entity pairs in different relationship categories is calculated to obtain the relationship of entity pairs.
[0039] As a preferred embodiment, the corpus includes a collection of laws and regulations on human genetic resources, management regulations of different units on human genetic resources, human genetic resources declaration form templates, human genetic resources related literature and general corpus; based on the corpus, a knowledge graph is constructed.
[0040] Node generation specifically includes using a pre-trained language model PLM, such as T5, to fine-tune on an entity (graph node) extraction task. In one embodiment, using PLM, node generation is formulated as a sequence-to-sequence problem, where the system is fine-tuned to convert text input into a node sequence separated by a special token, 〈PAD〉NODE 1 〈NODE_SEP〉NODE 2 ···〈 / S〉, where NODE i Represents one or more words.
[0041] In addition to entity recognition, the fine-tuning process also provides node features for the downstream edge generation task. Since each node can have multiple associated words, the generated string is greedily decoded and the node boundaries are demarcated using the separation marker 〈NODE_SEP〉, and the hidden state of the last layer of the decoder is mean pooled. The number of generated nodes is thus fixed in advance, and missing nodes are filled with a special 〈NO_NODE〉 token.
[0042] In one embodiment, the decoder receives as input a set of learnable node queries, represented as an embedding matrix. Causal masks are disabled to ensure that the Transformer processes all queries simultaneously. This is different from the traditional large language model encoder-decoder architecture. The output of the decoder in this embodiment can be directly read as d-dimensional node features of N (number of entities, also known as number of nodes) nodes. And passed to the prediction head (LSTM or GRU) to decode into node logits value Where S is the length of the generated node sequence, V is the vocabulary size, and n<=N.
[0043] (II) Edge Generation
[0044] The node feature set generated in the previous step is used for edge generation. Given a pair of node features, the prediction head decides whether there is an edge between their respective nodes. One option is to use a head similar to LSTM or GRU to generate edges in the form of token sequences. Another option is to use a classification head to predict edges.
[0045] Since it is very important to ensure the correct order (host-guest) between two nodes, a simple difference is used between the feature vectors when node i is the parent of node j.
[0046] When there is no edge between two nodes, it is indicated by a special tag 〈NO_EDGE〉. Since the number of actual edges is usually small and 〈NO_EDGE〉 is large, the generation and classification tasks are unbalanced for the 〈NO_EDGE〉 tag / class. To address the edge imbalance, the training setting is modified by sparse adjacency matrix to remove most of the 〈NO_EDGE〉 edges and rebalance the classes. The above modifications are only to improve training efficiency. During inference, the system still needs to output all edges.
[0047] Use a graph database to store the above nodes and edges to obtain a knowledge graph.
[0048] The knowledge graph describes the concepts and their interrelationships in the field of human genetic resources in a symbolic form, attempting to transform unstructured data into structured data based on the data itself, and connect various data together to form a graph model containing massive structured data. As time goes by, the knowledge and management regulations in the field of human genetic resources will be continuously updated. As a preferred embodiment, it also includes: iterating the knowledge graph based on a preset graph evolution algorithm; the preset graph evolution algorithm is an algorithm that is a combination of one or more of verification, improvement, regular update and maintenance, such as real-time response to external data updates, using incremental update technology, real-time streaming data collection based on events or based on timestamps, and integrating the changing relevant legal and regulatory data into the knowledge graph.
[0049] As a preferred embodiment, it also includes: using intuitive methods, such as Neo4j's bloom and / or Vue+D3v6 dynamic display, to realize the visualization of the human genetic resources compliance knowledge graph.
[0050] It should be noted that when constructing a knowledge graph, the privacy and security of the data also need to be considered. In particular, when sensitive information is involved, appropriate security measures need to be taken to protect the data. As a preferred embodiment, it also includes: using the FedE framework to allow multiple users to calculate entity embeddings locally, and then coordinate the aggregation of these embeddings through the server without involving the exchange of original knowledge graph data. As a preferred embodiment, it also includes: using the differential privacy knowledge graph embedding framework DPKGE to process knowledge graphs containing confidential and non-confidential statements.
[0051] 2. According to the prompt words of the large language model, the human genetic resources compliance knowledge graph is queried based on the pre-trained large language model to obtain the answer to the question entered by the user.
[0052] Specifically, the present invention regards the knowledge graph as a dynamic database accessible to the large language model, and uses the above knowledge graph to query to answer complex human genetic resources project application questions and provide intelligent application suggestions. The large language model is pre-trained with a set of human genetic resources laws and regulations, management regulations of different units on human genetic resources, human genetic resources application form templates, human genetic resources-related documents and general corpus on human genetic resources management regulations and general corpus. During training, the large language model will learn language and task-related patterns in the field of human genetic resources. Adjust the parameters of the large language model, including learning rate, batch size, number of iterations, etc., to optimize the performance of the large language model. Use the validation set to evaluate the large language model to check the performance and generalization ability of the large language model. Fine-tune and improve according to the validation results.
[0053] As a preferred embodiment, the present application can integrate the knowledge graph into the training objectives during the pre-training stage of the large language model, such as designing a word mask probability based on the knowledge graph structure, or introducing a text-entity correspondence training objective.
[0054] In order to answer the complex human genetic resource project declaration problem, that is, to query the correspondence between human genetic resource project declaration entities and regulatory entities based on the human genetic resource compliance knowledge graph, we can first fine-tune the text representation and then perform graph reasoning.
[0055] First, learn text representation. Use a pre-trained large language model, such as BERT, to implement an instance encoder. Use a neural network to encode the instance into a vector. Then, use an MLP layer to reduce the dimension and perform TransE fine-tuning.
[0056] Specifically, given an input project declaration text p and a regulatory text l, BERT is used to obtain the text representation as follows: p =BERT(p),B l =BERT(l)(1)
[0057] Among them, p and l are the original input text, and B p and B l is the output of BERT. Then, using the MLP layer, as follows: p =ReLU(W p *B p +b p ),f l =ReLU(W l *B l +b l )(2)
[0058] Among them, ReLU is the activation function, B p and B l As the input feature vector, W p and W l corresponds to B p and B l The weight matrix represents the connection strength between different nodes, f p and f l is the output text of MLP, which will be fed into GNN, and b is the bias parameter. Use TransE to fine-tune the text representation.
[0059] As shown below: score(p,r,l) transe =∥f p +f r -f l ∥ q (3)
[0060] where score(p,r,l) is the score of the triple (p,r,l), and f p , f l is the entity representation in (2), f r is a randomly initialized vector. The calculation formulas for DistMult and SimplE scores are:
[0061] score(p,r,l) dismult =∑(f p *f r *f l )(4)
[0062] score(p,r,l) SimplE =(∑(f p*f r *f l )+∑(f p *f rinv *f l )) / 2(5)
[0063] Here, inv represents the inverse function. P-tuning-v2 can be combined with BERT by adding a continuous prefix to the input layer of BERT and concatenating the Key and Value of each layer.
[0064] Second, graph reasoning. After obtaining the learned text representation, GNN is used to learn explicit relational knowledge. By assimilating the general message passing reasoning algorithm with the neural network counterpart, vertex embedding with legal reasoning is learned. Then, the residual connection in the text representation is used to obtain the final representation. Finally, TransE, DistMult and SimplE are used for scoring.
[0065] Specifically, let the vertex v i It is fed into the graph encoder to obtain a hidden vector, which explicitly models the graph structure of the legal knowledge graph. A GNN model following GAT and R-GCN is used. The GAT model encodes vertices using multiple graph attention layer connections, that is, the attention weights of adjacent nodes are calculated for aggregation. R-GCN utilizes a multi-layer relational graph convolution layer to represent nodes. The relationship of triples is considered during the convolution process, and different aggregation weights can be learned according to different relationships. Afterwards, the residual connection output by the MLP layer is added to the graph node representation, the residual connection y=F(x)+x, where F(x) is the nonlinear transformation of the network layer, x is the input, and the vertex v is converted into i After inputting the GNN model to obtain the residual connection, the graph node is represented as: i '=v i +GNN(v i ), where v i ' is the final entity representation of the entity utilizing text and graph reasoning.
[0066] Finally, the same TransE, DistMult, and SimplE as in the text representation learning phase are used for scoring. In the graph reasoning phase, f p With f l Combining text and graph features, while f r Initialized from the adjusted embeddings learned from text representations (Eqs. (3) and (5)).
[0067] For each legal and regulatory text, the tokenizer of the basic model encodes the text into a token sequence x = (x 0 ,x i ,···), using autoregressive method to analyze the basic model fθ (·) Conduct legal pre-training:
[0068] L P (θ,DP)=E x~DP [-∑ i logf θ (x i |x 0 , x 1 ,···,x i-1 )] (6)
[0069] where x 0 ,x 1 ,…,x i-1 represents the context tag, x i represents the target tag, θ is the basic model f θ (·) parameter. E represents epoch, and one epoch means that the data of the entire training set DP passes through the basic model of autoregression once, and x represents the number of epochs.
[0070] Furthermore, we construct a dataset DF, which consists of three subsets: (a) a dataset for predicting compliant types; (b) a dataset for compliant question-answering tasks; and (c) a dataset constructed by refining subsets (a) and (b) using a large language model, such as ChatGPT. We re-inject subset (c) using the following template 1, where instruction and question are injected with actual text and questions.
[0071] Template 1: I hope you will play the role of a compliance expert. I will give you a question-and-answer text related to human genetic resources project declaration and laws and regulations. Please polish it in a formal style. Requirements: 1. Correct grammatical errors and punctuation errors, remove special symbols, and make the sentences smooth. 2. Make the logic clearer and the format standardized, such as <answer>3. Do not write any explanatory statements. 4. <question>is the problem, <answer>is the answer. The conversation is:\n <question>:{instruction}\n <answer>:{output}\n\n, returns the result in JSON format.
[0072] In Template 2, the Stanford Alpaca template is used to wrap the instructions and outputs in the data set. Then the fine-tuning parameter θP is processed as follows: L F (θ,DF)=E x~DF [-∑ i∈{output} logf θP (x i |x 0 , x 1 ,···,x i-1 )] (7)
[0073] Where θ represents the optimized parameter, x = x 0 ,x 1 ,…,x i-1 represents the tokenized input sequence extracted from the dataset DF and wrapped by template 2, and represents the index set of the output tokens. That is, the optimization θP is θF.
[0074] Template 2: Below are instructions describing the task. Write an appropriate response based on the instructions. \n\n###Instruction:\n{instruction}\n\n###Response:\n{output}.
[0075] When applying the knowledge graph to downstream tasks, the text is labeled as x=x 0 ,x 1 ,…,x n Then, the labeled input sequence is fed into the fine-tuned model f θF (·), the response is generated in an autoregressive manner.
[0076] When fine-tuning data samples, the LoRA strategy can also be combined to inject the trainable rank decomposition matrix into each weight of the Transformer layer of the large language model.
[0077] In a preferred embodiment, based on the user input text, according to the large language model prompt words, based on the pre-trained large language model, the human genetic resource compliance knowledge map is queried to obtain a professional answer to the human genetic resource declaration type, including: based on the user input text, constructing the large language model prompt words:
[0078] Get the original text and use it as N 0 Level prompt words;
[0079] The prompt word model is used to optimize the original problem and the optimized N 1 Level prompt words;
[0080] Use dialogue to 1 The level prompt word is fed back to the user, and the user debugs and optimizes to obtain N 2 Level prompt words;
[0081] Keep iterating until you get N that satisfies the user. final The prompt word is used as the prompt word of the large prediction model, where final is greater than or equal to 0.
[0082] Specifically, the original question comes from voice, text or picture content, and a plurality of original question templates may be pre-configured to provide guidance for the original question input by the user.
[0083] The N 1 Level to N final The prompt words of the level include fixed words and placeholders.
[0084] The prompt word model includes a prompt word template, which can be used to construct prompt words, including a fixed word part and a placeholder. The fixed word is used to guide the user to input the corresponding placeholder, and the placeholder is used to describe the prompt word activity type to be constructed and the content involved in each activity type. According to different types of activities, the placeholder in the prompt word module is variable.
[0085] Specifically, the placeholders include one or more of activity type, research type, participating units, implementation methods, human genetic resource types, and population types.
[0086] The prompt word template can be divided into five types according to different activity types. The number of placeholders in the five activity types is different, and each placeholder is assigned a value of 1. The five prompt word templates correspond to the five activity types respectively. The placeholder assignments in the five prompt word templates are added to obtain the set assignment sum of the prompt word template. The sum of the assignments of the five prompt word templates is a fixed value. The five activity types are scientific research, public data, transmission data, human material export and sample library construction, among which the placeholder assignment sum of the scientific research activity template is 5, the placeholder assignment sum of the public data activity template is 2, the placeholder assignment sum of the transmission data activity template is 3, the placeholder assignment sum of the human material export activity template is 2, and the placeholder assignment sum of the sample library construction activity template is 2.
[0087] The prompt word model also includes a text segmentation algorithm, which identifies keywords in the original input text and matches the keywords with placeholders. If there is a keyword corresponding to the placeholder, the keyword is assigned a value of 1, otherwise it is 0. If the original input text does not include a keyword that can correspond to the placeholder, the original input text is assigned a value of 0. The total assignment of keywords in the original input text is calculated; if the assignment sum is lower than the set value sum of the corresponding activity type prompt word template, the user is prompted to continue debugging the prompt word until the prompt word assignment is greater than or equal to the set assignment sum of the corresponding template.
[0088] The placeholders include one or more of activity type, research type, participating units, implementation methods, human genetic resource type, and population information.
[0089] Specifically, the keywords and placeholders generally include activity types, entity types, and human genetic resource types; further, the activity types include scientific research, public information, information transmission, material export, and sample library construction; further, the research types include registration and listing research and exploratory research;
[0090] The participating units include: sponsors, clinical medical and health institutions, contract research organizations, third-party laboratories, information receiving units, and sample preservation units;
[0091] The implementation methods include whether a third-party laboratory is specified in the research protocol;
[0092] The types of human genetic resources include: human genetic resource materials and human genetic resource information; further, the human genetic resource materials include cells, whole blood, tissues or tissue sections, semen, cerebrospinal fluid, pleural / peritoneal effusion, blood / bone marrow smears, and hair with hair follicles; the human genetic resource information includes: genetic information involving accounting information such as human genes, genomes, transcriptomes, and epigenomes;
[0093] The population information includes: population in specific areas, important genetic family lines, and population numbers exceeding 3,000.
[0094] For example, the prompt word template is in the form of "carry out scientific research activities, the research type is ${var}, use human genetic resources ${var}, participating units ${var}, implementation methods ${var}, population information ${var}", where the placeholder ${var} is filled with the text entered by the user as a variable, so that the prompt word template can obtain a prompt word in combination with the information entered by the user. The number of placeholders in the prompt word template is 3, and the set assignment value of the prompt word template is 3;
[0095] The prompt words can continuously improve the large language model prompt words according to the detail of the input question until the answer to the question related to human genetic resources can be obtained based on the prompt words.
[0096] The large language model involved in the present invention can be any open source ChatGLM-6B, ChatGPT series, StableVicuna, PaLM, Galactica or LLaMA series model.
[0097] In addition, the present invention systematically analyzes the feedback provided by users to identify repetitive problems, misunderstandings or deficiencies, and actively collects feedback from professionals and users in the biomedical and legal fields to understand their views on the performance of large language models, including the accuracy, interpretability and practicality of large language models. Based on user feedback, relevant data is collected and annotated to improve the training and fine-tuning of large language models.
[0098] The present invention also provides a recommendation system for assisting declaration of human genetic resources based on a large language model, including: a prompt word module, used to construct a large language model prompt word based on user input text; a knowledge graph module, used to construct a human genetic resource compliance knowledge graph based on the acquired corpus; a query module, used to query the human genetic resource compliance knowledge graph based on a pre-trained large language model according to the large language model prompt word, and obtain an answer to the user input text.
[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.< / answer> < / question> < / answer> < / question> < / answer>
Claims
1. A text processing method based on knowledge graph and large language model, characterized in that: Construct a human genetic resources compliance knowledge graph, which is a set of triples in the format of (p, r, l), where p, in is the entity vocabulary, It is a set of predefined relationships, p represents the head entity, l represents the tail entity, and r represents the relationship between entities; the steps of constructing the knowledge graph include node generation and edge generation: first, node generation, obtaining pre-labeled data from unstructured input text, using a pre-trained large language model, and fine-tuning on the graph node extraction task; second, edge generation, given a pair of node features, using the prediction head LSTM or GRU to determine whether there is an edge between the nodes; use a graph database to store the above nodes and edges to obtain a compliant knowledge graph of human genetic resources.
2. The method according to claim 1, characterized in that The pre-labeled data is obtained based on a multi-layer pointer network and a multi-head attention mechanism.
3. The method according to claim 1, characterized in that The ChatGLM-6B model fine-tuning system is used in the node generation process to convert text input into node sequences. Prior to this, the ChatGLM-6B model is fine-tuned using the P-tuning-v2 fine-tuning method.
4. The method according to claim 1, characterized in that During the edge generation process, when node i is the parent node of node j, the difference between the feature vectors is used.
5. A method for querying answers to questions input by users using the method of claim 1, characterized in that: Based on the user input text, build a large language model prompt word, including: Get the original text and use it as the N0 level prompt word; The original question is optimized using the prompt word model to obtain the optimized N1-level prompt words; The N1 prompt words are fed back to the user in a dialogue mode, and the user debugs and optimizes to obtain the N2 prompt words. The iteration is continued until the N2 prompt words that the user is satisfied with are obtained. final Prompt word, which is used as the prompt word of the large language model, where final is greater than or equal to 0.
Citation Information
Patent Citations
Chinese entity and relation joint extraction method and system based on RoBERTa and pointer network
CN116663539A
Legal assistance method and system based on large language model
CN117743590A
Knowledge graph construction method and device based on pre-trained large language model
CN117851610A
Cited By
Complaint text generation method, system and equipment based on large language model and medium
CN120493899A
Entity label generation method and device and product
CN121599080A