A method and system for interpreting medical literature based on large language models
By constructing a medical word library relationship network and association path optimization, the language model-based method solves the problems of low interpretation efficiency and insufficient semantic understanding of existing literature, and achieves efficient and accurate interpretation of medical literature.
Patent Information
- Application Number
- CN202411644099.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-18
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-11-18
AI Technical Summary
The existing literature interpretation methods are inefficient and have insufficient semantic understanding, making it difficult to automatically generate coherent interpretation content.
构建基于语言大模型的医学文献解读方法,通过分词处理、医学单词库关系网、词向量匹配、上下文嵌入和关联路径优化,生成连贯解读内容。
It realizes accurate interpretation of medical literature, improves the accuracy and efficiency of literature analysis, and provides intelligent auxiliary tools for scientific researchers and doctors.
Smart Images

Figure CN119740575B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of language analysis, and specifically to a method and system for interpreting medical literature based on a large language model. Background Art
[0002] With the advent of the information age, the computing requirements to be processed have also shown exponential growth. Facing the rapid development of the biomedical field and the publication of a large number of scientific research results, the number of medical literature has increased exponentially. Researchers and medical workers are facing the challenge of information overload. How to efficiently obtain and understand the latest medical information has become a key issue in scientific research and clinical practice. Traditional literature reading and analysis methods are difficult to cope with the massive data and are prone to information omission.
[0003] In recent years, the progress of artificial intelligence technology, especially in the fields of natural language processing (NLP) and deep learning, has provided a new way to solve this problem. Large-scale pre-trained language models (such as BERT, GPT, etc.) perform excellently in tasks such as semantic understanding, information extraction, and content generation, showing great potential in medical literature processing. These models can automatically identify information such as medical terms, symptom descriptions, and pathological relationships in the literature through semantic analysis and context understanding, and can extract the core points from complex texts.
[0004] In addition, language models developed specifically for the medical field (such as BioBERT, ClinicalBERT, etc.) combine the pre-training of medical-specific corpora and can more accurately identify the semantic relationships between medical terms. The application of these technologies not only helps to accelerate the dissemination of medical knowledge but also provides more intelligent literature assistance tools for researchers and doctors, improving the efficiency of knowledge acquisition and the scientific nature of decision-making. Summary of the Invention
[0005] In view of the above existing problems, the present invention is proposed.
[0006] Therefore, the technical problem solved by the present invention is the optimization problem of the existing literature interpretation method, which has low efficiency, insufficient semantic understanding, and difficulty in automatically generating coherent interpretation content.
[0007] To solve the above technical problem, the present invention provides the following technical solution: A method for interpreting medical literature based on a large language model, including:
[0008] Obtain the medical literature to be interpreted and perform word segmentation processing;
[0009] Construct a relationship network between words for the medical word library to obtain the first relationship network;
[0010] Match the words in the text with the words in the medical word library to obtain the matching results of the text words; according to the first relational network, determine the first association relationship between each text word and other words in the text, and assign weights to the first association relationship according to the number of associations.
[0011] Construct a relational network for the words in the text according to the text content to obtain a second relational network; determine the second association relationship between each text word and other words in the text, and assign weights to the second association relationship.
[0012] Optimize the association path according to the first relational network and the weights assigned to the first relational network, and obtain the interpretation result through the interpretation of the content on the path.
[0013] As a preferred solution of the medical literature interpretation method based on a large language model according to the present invention, wherein: the word segmentation process includes using the BioBERT model to segment the content in the text.
[0014] The medical word library includes the standard expressions and synonyms of all medical specialized terms entered, and assimilates the standard expression of each term and its synonyms as a word that can be matched.
[0015] As a preferred solution of the medical literature interpretation method based on a large language model according to the present invention, wherein: the first relational network includes entering the relationship between each word while entering the medical word library.
[0016] Use the words in the medical word library as nodes, connect the words with specific semantic or logical relationships, and assign a weight of 1 to each connection line.
[0017] As a preferred solution of the medical literature interpretation method based on a large language model according to the present invention, wherein: matching the words in the text with the words in the medical word library includes using a word embedding model to convert the words in the text and the words in the medical word library into vector representations.
[0018] Calculate the i-th word vector in the text and the j-th word vector in the medical word library the similarity between them;
[0019] The similarity is expressed as:
[0020]
[0021] When When the match is successful, each word in the text is successively matched one by one with all the words in the medical word library; during the matching process, calculate one by one the similarity between the corresponding standard expression and its synonyms and and take the minimum value as the similarity final result;
[0022] For the i-th word vector the set of all word vectors matched = ;
[0023] Among them, represents the L-th successfully matched word vector in the medical word library;
[0024] The first association relationship includes, after completing the word matching, obtaining all the word vectors in the text whose matching set is not an empty set; arbitrarily taking two word vectors and , for each element in the set matched with , analyze and the connection relationship existing between the elements in the set matched with, and assign a weight to the first association relationship by accumulating the connection times;
[0025]
[0026] Among them, represents the i-th word vector in the text and the v-th word vector in the text association weight; a represents the element index in the set matched with the i-th word vector in the text; b represents the element index in the set matched with the v-th word vector in the text; A represents the number of elements in the set matched with the i-th word vector in the text; B represents the number of elements in the set matched with the v-th word vector in the text; represents the i-th word vector in the text in the matched set, the a-th element; represents the v-th word vector in the text in the matched set, the b-th element; represents the discrimination function of whether there is an association. In the first relational network, if there is a connection between them, record it as 1; if there is no connection between them, record it as 0.
[0027] As a preferred solution of the medical literature interpretation method based on the large language model of the present invention, wherein: the second relational network includes, for the input text , using the pre-trained large language model to perform context encoding on each word;
[0028] After encoding, the context embedding vector of each word is obtained:
[0029]
[0030] Wherein, represents the context embedding of the word , a context embedding vector containing the semantic information of the word in a specific context obtained through the output of the large model; represents the d-th word, and d represents the number of words; represents the specific position output by the index model;
[0031] After obtaining the context embedding vector of each word, a neural network module is used to calculate the semantic relevance weight between each pair of words;
[0032] The formula for calculating the semantic relevance weight is:
[0033]
[0034] Wherein, represents the semantic relevance weight of the word pair ; and represent trainable weight matrices; is the bias vector, and σ is the activation function; represents concatenation of, the word context embedding vector; , 1 means the words and have the strongest relevance, and 0 means no relevance;
[0035] After completing the semantic relevance calculation of each pair of words, multiplying by the adjustable coefficient E, the actual association weight between every two words in the text is obtained.
[0036] As a preferred solution of the medical literature interpretation method based on the large language model of the present invention, wherein: the optimization of the association path includes setting the starting word as the first word in the text whose similarity to any word in the medical word library is greater than , denoted as ;
[0037] From Analyze the sum of weights with other words in the text, select the word with the largest sum of weights as the next word in the path, and connect them in sequence until the end;
[0038] The sum of weights is expressed as:
[0039]
[0040] Wherein, Represents the sum of weights between word and ; Represents the semantic relevance weight between word and ; Represents a selection function. If and both have a similarity greater than with any word in the medical word library, then the output is ; Represents the first association relationship weight between word and ; If and do not satisfy the condition that the similarity with any word in the medical word library is greater than , then the output is 0.
[0041] As a preferred solution of the medical literature interpretation method based on the language large model of the present invention, wherein: the content interpretation includes using the word string in the path as input, evaluating the integrity of the word string through a trained bidirectional LSTM network, and outputting an integrity score T;
[0042] If T is greater than the set threshold , it is determined that the input word string is sufficient to form a coherent content; otherwise, it is determined that it cannot be recognized; when the output is a non-recognizable judgment result, E starts from the initial value, increases by 1 each time, and after the increase ends, re-optimize the path, and re-interpret the optimization result until the content of the path is interpretable;
[0043] Use a pre-trained Seq2Seq generation model, and by calling the generation function, input the word string into the model to generate a coherent interpretation text:
[0044]
[0045] Wherein, Is the generated interpretation content, Represents the BART model used to generate coherent interpretation content, Represents the input word string.
[0046] A medical literature interpretation system based on a large language model using the method described in the present invention, wherein:
[0047] An acquisition unit that obtains the medical literature to be interpreted and performs word segmentation processing;
[0048] A thesaurus unit that constructs a relationship network between words for the medical word library to obtain a first relationship network;
[0049] A first analysis unit that matches the words in the text with the words in the medical word library to obtain a matching result of the text words; according to the first relationship network, determines the first association relationship between each text word and other words in the text, and assigns weights to the first association relationship according to the number of associations;
[0050] A second analysis unit that constructs a relationship network for the words in the text according to the text content to obtain a second relationship network; determines the second association relationship between each text word and other words in the text, and assigns weights to the second association relationship;
[0051] An interpretation unit that optimizes the search for association paths according to the first relationship network and the weights assigned to the first relationship network, and obtains an interpretation result through the interpretation of the content on the path.
[0052] A computer device, comprising: a memory and a processor; the memory stores a computer program, wherein: when the processor executes the computer program, the steps of the method described in any one of the present invention are implemented.
[0053] A computer-readable storage medium, on which a computer program is stored, wherein: when the computer program is executed by a processor, the steps of the method described in any one of the present invention are implemented.
[0054] Advantages of the present invention: The medical literature interpretation method based on a large language model provided by the present invention realizes the accurate interpretation of medical literature by constructing a multi-layer relationship network of a medical word library and utilizing the context analysis ability of the large language model. By optimizing the path generation method, the present invention can effectively extract the core medical information in the literature and automatically generate coherent interpretation content, thereby overcoming the deficiencies of traditional calculation methods in processing efficiency and semantic understanding. Compared with the prior art, the present invention greatly improves the accuracy and efficiency of literature analysis, provides an intelligent auxiliary tool for medical researchers and clinicians, and greatly improves the convenience of knowledge acquisition. Description of the Drawings
[0055] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0056] Figure 1 This is the overall flowchart of a method for interpreting medical literature based on a large language model provided by the first embodiment of the present invention. Specific embodiments
[0057] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the specific embodiments of the present invention in detail with reference to the accompanying drawings of the specification. Obviously, the described embodiments are some, rather than all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0058] Embodiment 1, referring to Figure 1 , which is an embodiment of the present invention, provides a method for interpreting medical literature based on a large language model, including:
[0059] S1: Obtain the medical literature to be interpreted and perform word segmentation processing.
[0060] The word segmentation processing includes using the BioBERT model to perform word segmentation on the content in the text.
[0061] Furthermore, BioBERT is a language model pre-trained based on medical corpora, which can identify professional medical vocabulary and phrases in the literature (such as drug names, diseases, treatment methods, etc.), so as to more accurately separate the key information units in the medical literature.
[0062] Word segmentation is the basis of semantic analysis. After dividing the literature content into words, it can provide basic data for constructing the semantic relationship network between words, enabling the model to analyze the context semantic associations between words. The word segmentation processing can remove irrelevant or auxiliary words in the literature, retain meaningful medical terms and keywords, reduce the complexity of semantic analysis, and improve the accuracy of interpretation.
[0063] S2: For the medical word library, construct a relationship network between words to obtain the first relationship network.
[0064] The medical word library includes the standard expressions and synonyms of all medical proprietary vocabulary. The standard expression of each vocabulary and its synonyms are assimilated as a matchable word.
[0065] The first relational network includes, while inputting the medical word library, inputting the relationships between each word. Taking the words in the medical word library as nodes, connecting the words with specific semantic or logical relationships, and assigning a weight of 1 to each connection line.
[0066] It should be noted that by node-ifying the vocabulary in the medical word library and constructing the semantic relationship connections between words, a semantic network can be formed. Each connection line represents a specific relationship between two medical vocabularies (such as the treatment relationship between drugs and diseases, the manifestation relationship between diseases and symptoms, etc.), which helps to analyze the logical or functional associations between vocabularies. There are a large number of synonyms and different standard expressions in medical vocabulary. By assimilating and processing, they are mapped to a standardized word node, enabling the system to recognize these synonymous expressions and reduce misunderstandings, ensuring the accuracy of the lexical relationships.
[0067] By assigning weights (such as 1) to the relationships between vocabularies, a standardized semantic structure is formed. Subsequent analysis can utilize this weight information to calculate the lexical association paths, facilitating path optimization and semantic depth analysis.
[0068] It should also be known that in medical literature, many medical vocabularies do not exist independently, but appear in pairs or groups. For example, certain drugs have a treatment relationship with specific diseases, and the descriptions of certain diseases contain specific symptoms. The first relational network can reflect these combined relationships and provide necessary context support for the interpretation process. In medical literature, terms such as diseases, symptoms, drugs, treatment plans, etc. are related to each other. The first relational network provides a semantic network, enabling the system to perform reasoning and select association paths based on the relationships between terms, and mine deeper medical information. The first relational network not only provides a direct mapping of lexical relationships, but also helps to determine the association strength between vocabularies by assigning weights. Through the weight information, the paths can be preferentially selected according to the strength of the association in subsequent path optimization, ensuring that the generated content is more coherent and logically clear semantically.
[0069] S3: Match the words in the text with the words in the medical word library to obtain the matching results of the text words; according to the first relational network, determine the first association relationships between each text word and other words in the text, and assign weights to the first association relationships according to the number of associations.
[0070] Specifically, use a word embedding model to convert the words in the text and the words in the medical word library into vector representations. Calculate the similarity between the i-th word vector in the text and the j-th word vector
[0071] in the medical word library. The similarity
[0072]
[0073] when When , the match is successful; each word in the text is matched one by one with all the words in the medical word library in turn; during the matching process, the The corresponding standard expressions and their synonyms are The minimum value is taken as the similarity The final result.
[0074] For the i-th word vector The set of all matched word vectors = .
[0075] in, Represents the Lth successfully matched word vector in the medical word library.
[0076] It should be noted that words in a text may be expressed in different ways, especially in the medical field, where the same concept may have multiple synonyms or near-synonyms. By matching each word with the standard expression and its synonyms in the medical word library one by one, it is ensured that the same term in different expressions can be identified, and the minimum similarity value is selected as the final matching standard to obtain higher matching accuracy. · Medical terms are often expressed in a variety of ways, and different words may express similar or related medical concepts. The word embedding model is used to convert words in the text and words in the medical word library into vector representations, which can quantify the semantic similarity between words. Through similarity calculation, these semantically associated words can be effectively captured, helping the system to still identify the same medical concept in the case of diverse word expressions. For each word in the text, the set of all matched word vectors (where each element represents a successfully matched word vector) is used for subsequent association relationship analysis. By analyzing these successfully matched word vector sets, the relationship between words can be better captured, and potential medical topics and concept associations in the literature can be identified, thereby providing data support when constructing the first relationship network.
[0077] Assume that the word in the text is "heart disease", and the medical word library contains "heart disease" (standard expression) and its synonyms "coronary heart disease" and "cardiovascular disease". The similarity between "heart disease" and each synonym is calculated through the word embedding model. If a synonym has the lowest similarity with "heart disease" and reaches the matching threshold, then the synonym will be considered a successful match and enter the set. This ensures that different term expressions point to standard medical concepts and ensures the accuracy of the analysis.
[0078] The first association relationship includes, after completing the word matching, obtaining all word vectors in the text whose matching set is not an empty set; arbitrarily taking two word vectors and , for each element in the set matched with , analyze the connection relationship existing between the elements in the set matched with . By accumulating the connection times, assign a weight to the first association relationship.
[0079]
[0080] Among them, represents the association weight between the i-th word vector in the text and the v-th word vector in the text ; a represents the element index in the set matched with the i-th word vector in the text ; b represents the element index in the set matched with the v-th word vector in the text ; A represents the number of elements in the set matched with the i-th word vector in the text ; B represents the number of elements in the set matched with the v-th word vector in the text ; represents the a-th element in the set matched with the i-th word vector in the text ; represents the b-th element in the set matched with the v-th word vector in the text ; represents the discriminant function of whether there is an association. In the first relational network, if there is a connection between them, record it as 1; if there is no connection between them, record it as 0.
[0081] It should be noted that there are close logical or semantic relationships among certain medical terms in the text, such as the relationships between diseases and symptoms, and between drugs and therapies. By analyzing whether there are connections between the elements in each pair of word vector sets and accumulating them, the relevance of these terms can be accurately identified, thus constructing the first association relationship that reflects the actual knowledge structure in medical literature. The first association relationship not only needs to identify whether there is an association between terms, but also distinguish the strength of the association. By accumulating the number of connections of the elements in the matching sets to assign weights, the strength of the association relationship can be reflected. The higher the weight, the more frequently the two terms are mentioned together in the literature, indicating a higher degree of association. This design helps the model to preferentially select high-weight paths in subsequent analysis, making the content interpretation more accurate. In medical literature, certain terms often appear in fixed combinations. Through this step, these co-occurrence patterns can be captured, providing data support for analyzing the relationships between medical concepts. For example, cardiovascular disease and hypertension may co-occur frequently. By accumulating the weights of the co-occurrence relationships, these fixed combinations can be identified, laying the foundation for the concepts in the medical knowledge graph.
[0082] The weighted first association relationship is crucial in the subsequent path optimization and content generation processes. By preferentially selecting the term paths with higher weights, the system can generate more accurate interpretations based on real semantics, improving the coherence and logic of the interpreted content. This provides reliable support for the final output of the model, ensuring that the interpretation results can accurately reflect the medical information in the literature.
[0083] S4: According to the text content, construct a relationship network for the words in the text to obtain the second relationship network; determine the second association relationship between each text word and other words in the text, and assign weights to the second association relationship.
[0084] Furthermore, the second relationship network includes, for the input text , use a pre-trained large language model to perform context encoding on each word.
[0085] After encoding, obtain the context embedding vector of each word:
[0086]
[0087] Among them, represents the context embedding of the word , which is obtained from the output of the large model and is a context embedding vector containing the semantic information of the word in a specific context; represents the d-th word, where d represents the number of words; represents the specific position output by the index model.
[0088] After obtaining the context embedding vectors of each word, a neural network module is used to calculate the semantic relevance weights between each pair of words. The formula for calculating the semantic relevance weights is as follows:
[0089]
[0090] where represents the semantic relevance weight of the word pair ; and represent trainable weight matrices; is the bias vector, and σ is the activation function (such as the Sigmoid function); represents the concatenation of ; is the context embedding vector of the word ; , 1 indicates that the words and have the strongest relevance, and 0 indicates no relevance.
[0091] After completing the calculation of the semantic relevance of each pair of words, multiplying by the adjustable coefficient E gives the actual association weights between every two words in the text.
[0092] It should be noted that the concatenated vector contains the context information of the two words, so it is more helpful to capture their semantic relevance. This concatenation method enables the model to simultaneously focus on the semantic relationship between the two words in the context during learning. Through concatenation, the model can perform non-linear operations based on the combined information of the two words, such as weighted sums of weight matrices and transformations of activation functions. This allows the model to perform complex interaction calculations on the context vectors of the two words and extract more advanced semantic association features. After inputting the concatenated context vectors, the neural network can learn more complex relationships through multiple levels of weights, rather than just the features of individual words. This makes the model have stronger discriminative ability when facing diverse contexts and complex semantic relationships. Using the same concatenated input format for all word pairs helps to maintain the simplicity and consistency of the neural network calculation, avoiding inconsistencies or information loss when calculating word vectors one by one.
[0093] It should also be noted that the second relational network is a relational network generated based on the context of the text, which is different from the static relationship of the first relational network based on the medical word library. Through the second relational network, the system can construct a lexical relational network that is closer to the actual context of the text, reflecting the real associations of the words in the literature in the context. This dynamic relationship helps to capture the subtle semantic associations in the literature, thereby enabling a more accurate understanding of the connections between medical concepts. The semantic relevance weights in the second relational network are calculated through a neural network. By using the context embedding vectors of the words and concatenating and inputting them into the neural network module, the semantic relevance weights of each pair of words can be obtained. The relevance score ranges from 0 to 1, indicating the degree from no association to the strongest association. In this way, the system can quantitatively analyze the relationships between words in detail, making the generated relational network more precise and expressive. The dynamic semantic relationship of the second relational network provides a richer context basis during content generation. A high semantic relevance weight means a strong contextual association between two words in the literature, which is very important for content generation. By preferentially selecting high-weight paths, the system can generate more coherent and semantically logical interpretation content, thereby ensuring the accuracy and natural fluency of the output content. Suppose in a medical literature, there are two words, "heart disease" and "blood pressure". Through context encoding and the relevance calculation of the neural network, their context embedding vectors are obtained, and a high semantic relevance weight (such as 0.8) is calculated. This score indicates that in the specific context of the current literature, "heart disease" and "blood pressure" have a strong semantic association. This high association can prompt the system to preferentially connect these two concepts when interpreting the content, thereby generating explanatory content containing the relationship between the two, such as "Patients with heart disease usually have blood pressure problems."
[0094] S5: Optimize the association path according to the first relational network and the weights assigned by the first relational network, and obtain the interpretation result through the content interpretation on the path.
[0095] Set the starting word as the first word in the text whose similarity to any word in the medical word library is greater than and set it as . Starting from analyze the sum of weights with other words in the text, select the word with the largest sum of weights as the next word in the path, and connect them in sequence until the end.
[0096] The sum of weights is expressed as:
[0097]
[0098] where represents the sum of weights between word and , represents word and The semantic relevance weight; Indicates a selection function. If and both have a similarity greater than with any word in the medical word library, the output is , Indicates the first association relationship weight of words and ; If and do not satisfy the condition that the similarity with any word in the medical word library is greater than , the output is 0.
[0099] In the literature, the relationship network between medical terms and concepts may be relatively complex. Through path optimization, the most semantically relevant path can be identified. Starting from the first word with a relatively high similarity to the medical word library, by selecting the next word with the largest sum of weights, it is ensured that the system can move forward along the path with the closest semantic association during interpretation. This enables the interpreted content to focus on the core semantics and avoid interference from irrelevant information. During the path optimization process, by accumulating the sum of weights to select the order of the most relevant vocabulary, the coherence of the interpreted content is guaranteed. At each step, the best next word is selected according to the sum of weights, enabling the system to generate an interpreted content that conforms to the logical order and avoiding situations where the interpretation results are fragmented or have semantic jumps.
[0100] By introducing the calculation of the sum of weights for semantic relevance weights, path optimization not only considers the direct association between two words but also refers to the strength of their semantic relationship when selecting the next word. The selection function is used to judge the relevance of each pair of words and determine their weights, ensuring that in the path optimization process, vocabulary combinations with relatively high similarity to the terms in the medical word library are preferentially selected, thereby improving the accuracy and semantic consistency of the interpretation.
[0101] Taking the word string in the path as the input, the trained bidirectional LSTM network evaluates the integrity of the word string and outputs an integrity score T.
[0102] If T is greater than the set threshold , it is determined that the input word string is sufficient to form coherent content; otherwise, it is determined that it cannot be recognized. When the output is a non-recognizable judgment result, E starts from the initial value, increases by 1 each time, and after the increase ends, path optimization is performed again, and the optimized result is interpreted again until the content of the path is interpretable.
[0103] It should be noted that when the score T does not reach the set coherence threshold, it indicates that the interpreted content may lack sufficient information. By gradually increasing the weight coefficient E, more attention can be paid to the core semantic relationships during the path optimization process, increasing the proportion of semantics, thereby strengthening the semantic relevance of the interpreted content. The continuously increasing semantic weight makes the system more inclined to select highly relevant word combinations during path optimization to ensure the accuracy and depth of the final interpreted content.
[0104] In the case where the score T does not meet the threshold requirements, the gradual increase of the weight coefficient E will guide the system to re-optimize the path and add more keywords. By introducing more related words, the system can capture richer semantic information, making the interpreted content more hierarchical and coherent. This design ensures that the interpreted content contains sufficient key information while making the content more complete and thorough. Directly outputting when the interpreted content is not coherent enough will affect the interpretation quality. Therefore, this step avoids generating meaningless or incomplete interpretations through integrity assessment. When the content is determined to be "unrecognizable", the system will trigger the semantic enhancement mechanism and introduce more related information during the optimization process to ensure that the output content meets the set coherence standard, thereby improving the quality of the interpretation result. Through repeatedly increasing the semantic weight and re-optimizing the path, an adaptive optimization mechanism for the interpreted content is formed. This design enables the system to adaptively adjust the degree of semantic attention according to the interpretation score, continuously optimizing the path selection until the content that meets the coherence requirements is generated. The iterative mechanism ensures that the system can adapt to text content of different complexities, making the interpretation result more accurate and reliable.
[0105] Suppose the word string "coronary heart disease, myocardial ischemia, surgery" obtained through path optimization is scored by a bidirectional LSTM network, and the result TTT does not reach the threshold. This indicates that this path may lack a certain semantic coherence. The system will start the process of increasing the weight coefficient EEE, re-optimize the path, and may introduce related words such as "coronary artery" or "interventional therapy" to further enhance the medical semantic information in the interpreted content. Repeat this iteration until an interpreted content with sufficient semantic information and clear logic is generated, such as "Coronary heart disease may cause myocardial ischemia and usually requires interventional therapy or surgery".
[0106] Using a pre-trained Seq2Seq generation model, by calling the generation function, input the said word string into the model to generate a coherent interpreted text:
[0107]
[0108] Among them, is the generated interpreted content, represents the BART model used to generate coherent interpreted content, represents the input word string.
[0109] You know, as a type of Seq2Seq architecture, the BART model, after being pre-trained on a large corpus, has powerful natural language understanding and generation capabilities. It can generate coherent and smooth interpretation texts based on the context after receiving the input key word string. This generation ability is especially suitable for inferring and completing the complete medical interpretation content from incomplete term strings, making the output more readable and logical. Seq2Seq models like BART have the characteristic of handling incomplete inputs. They can understand the potential semantics from the incomplete information in the form of word strings and generate coherent texts. Compared with traditional rule matching or template filling, BART can accurately infer reasonable interpretation content through context-related understanding in the case of incomplete or fragmented inputs, thus filling the semantic gaps in the word strings. The bidirectional encoder part of the BART model can capture the context semantics of the input word string, while its decoder part can generate coherent texts with long-range dependencies. For medical literature interpretation, the context dependencies and logical relationships between words are crucial. BART's bidirectional encoding ability ensures that the input word strings are reasonably connected at the semantic level, thus outputting interpretation content that conforms to medical logic. Seq2Seq models such as BART can combine pre-trained knowledge during the generation process, bringing more details and expansions to the word strings. For example, when the input contains keywords such as "heart disease", "treatment", and "drug", BART can combine the medical knowledge it learned during pre-training to automatically generate explanations about the treatment plan or drug selection for heart disease. Such a design greatly enriches the level and information volume of the interpretation content, making the output more professional and practical.
[0110] Assume the input word string is "coronary heart disease", "interventional treatment", "drug". BART can generate complete interpretation content based on this input, such as "Coronary heart disease is a common cardiovascular disease, usually treated by interventional methods and controlled with drugs." This interpretation is not just a simple combination of word strings, but provides complete medical backgrounds, treatment methods, and measures, which is exactly the value of the BART generation model design.
[0111] On the other hand, this embodiment also provides a medical literature interpretation system based on a large language model, which includes:
[0112] A collection unit that acquires the medical literature to be interpreted and performs word segmentation processing.
[0113] A thesaurus unit that constructs a relationship network between words for the medical word library to obtain the first relationship network.
[0114] The first analysis unit matches the words in the text with the words in the medical word library to obtain the matching results of the text words; according to the first relational network, determines the first association relationship between each text word and other words in the text, and assigns weights to the first association relationship according to the number of associations.
[0115] The second analysis unit constructs a relational network for the words in the text according to the text content to obtain a second relational network; determines the second association relationship between each text word and other words in the text, and assigns weights to the second association relationship.
[0116] The interpretation unit optimizes the association paths according to the first relational network and the weights assigned to the first relational network, and obtains the interpretation results through the interpretation of the content on the paths.
[0117] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0118] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0119] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or, if necessary, other suitable processing, and then stored in a computer memory.
[0120] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0121] Example 2, an embodiment of the present invention, provides a method for interpreting medical literature based on a large language model. To verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.
[0122] In this experiment, the performance of "Intelligent Interpretation System A" (based on the present invention) and "Standard Interpretation System B" (control system) was evaluated. The accuracy, time consumption, stability of multiple identifications, content coherence, and user satisfaction of the systems were investigated under different experimental conditions such as the amount of literature, text complexity, and the number of keywords, as shown in Table 1.
[0123] The experimental environment included three groups of test literature sets, each containing 50, 100, and 150 documents respectively. The literature covered complex medical topics such as cardiovascular diseases, cancer, diabetes, etc., with complex content structures and rich terminology. The number of key terms in each document was between 20 and 100, covering different diseases, treatment methods, drugs, etc. During the experiment, both System A and System B identified key terms from each group of literature, performed semantic associations, and generated interpretation content.
[0124] Process of the intelligent interpretation system A: System A uses the BioBERT model for word segmentation, identifies medical terms, and matches them with the medical word library. By constructing the first relational network, the standardized expressions of the terms and their synonyms are mapped to the same node, and initial weights are assigned according to the semantic relationships and logical associations between the terms. Then, System A uses the BART generation model to construct the second relational network, generates the semantic relevance weights of word pairs through context embedding, and uses the path optimization algorithm to select the most relevant path in the literature to generate coherent content. Finally, System A uses a bidirectional LSTM network to score the coherence of the generated path and adjusts the weights to optimize the path when the score is insufficient.
[0125] Process of the standard interpretation system B: System B uses the traditional keyword matching method, without relational network and path optimization, only identifies keywords in the text and simply combines and outputs them. System B lacks semantic association during content generation, generates content only based on surface matching, lacks coherence and association optimization, and has limited interpretation effect.
[0126] Table 1 Experimental data table
[0127] Test environment Number of documents (articles) Average accuracy rate (%) Average time consumption (seconds) Stability of multiple recognition results (%) Degree of content coherence (full score: 10) User satisfaction (full score: 10) System A (invention) 50 96.5 8 94 9.2 9.3 System B (control) 50 85.0 5 85 7.0 6.8 System A (invention) 100 96.0 12 93 9.4 9.2 System B (control) 100 82.0 7 82 6.9 6.5 System A (invention) 150 95.7 15 92 9.5 9.1 System B (control) 150 81.0 8 80 6.8 6.3
[0128] As can be seen from the above table data, for System A in different numbers of literatures, the accuracy of term recognition always remains above 95%, while that of System B is between 80% - 85%. The high accuracy of System A comes from the medical term recognition ability of the BioBERT model, as well as the standardization and synonym mapping processing through the medical word library and the first relational network. This enables System A to accurately identify the key terms in medical literature and reduce misidentifications. In contrast, System B has limited recognition ability for synonyms and medical terms due to the lack of lexical standardization processing, resulting in a lower accuracy rate and being unable to meet the high-precision requirements for medical literature interpretation.
[0129] The time consumption of System A is slightly higher than that of System B. As the number of literatures increases, the growth of the processing time of System A is more obvious. System A involves the construction of relational networks, context encoding, and path optimization, resulting in increased time consumption. However, these steps significantly improve the quality and logic of the interpreted content. In contrast, System B has less time consumption, but the interpreted content is fragmented and unable to generate a complete semantic structure. Although System A takes more time, for the medical literature analysis scenario that requires high-quality interpretation, the increased time investment is reasonable and necessary.
[0130] System A performed well in terms of multiple recognition stability, with stability maintained between 92%-94%, significantly higher than System B's 80%-85%. System A ensures that relatively consistent interpretation content can be generated during each interpretation process through multi-layer relationship network construction and weight dynamic adjustment mechanism. System B lacks semantic network and path optimization mechanism, and the results are highly volatile, and the stability of multiple recognition cannot be guaranteed. For the interpretation of medical literature, stability is a key factor in ensuring content coherence and information accuracy. System A's advantages in this regard further prove its reliability in practical applications. (The stability of multiple recognition results refers to the consistency of the recognized content during the test. In this experiment, the minimum value of the content consistency is taken as the "multiple recognition result")
[0131] System A always maintains a content coherence score of more than 9.2, while System B's score is between 6.8 and 7.0. The path optimization algorithm used by System A can give priority to paths with higher semantic weights, forming semantically coherent interpretation content when generating content. However, System B only extracts keywords, and the generated interpretation content lacks logic and contextual association, so the content coherence is poor. The design of System A is particularly suitable for the coherent interpretation of complex content in medical literature, and can provide more logical content output.
[0132] System A has a higher user satisfaction score than System B, with a stable score between 9.1 and 9.3, while System B's score is only between 6.3 and 6.8. System A performs well in terms of accuracy, coherence, and stability of multiple recognitions, making users more satisfied with its interpretation results. Especially in the application of medical literature interpretation, the semantic association and optimization mechanism of System A make it more suitable for professional needs. System B's interpretation content is fragmented and difficult to meet user needs, and its satisfaction score is low.
[0133] It can be seen from the experimental data that system A is significantly better than system B in terms of accuracy, multiple recognition stability and content coherence. Although system A is slightly more time-consuming, it has demonstrated outstanding innovation and reliability in overall performance by optimizing the interpretation path, improving semantic associations and ensuring content coherence. The design of system A not only improves the quality of interpretation, but also meets the high standards of medical literature interpretation for accurate information, logical coherence and consistent content. These experimental data fully demonstrate the advantages of the present invention, especially in the interpretation of complex medical content, which has extremely high application value.
[0134] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A method for interpreting medical literature based on a large language model, characterized in that, Including: Obtain the medical literature to be interpreted and perform word segmentation processing; For the medical word library, construct a relationship network between words to obtain the first relationship network; Match the words in the text with the words in the medical word library to obtain the matching results of the text words; according to the first relationship network, determine the first association relationship between each text word and other words in the text, and assign weights to the first association relationship according to the number of associations; According to the text content, construct a relationship network for the words in the text to obtain the second relationship network; determine the second association relationship between each text word and other words in the text, and assign weights to the second association relationship; According to the first relationship network and the weights assigned by the first relationship network, optimize the association path, and obtain the interpretation result through the interpretation of the content on the path; The first relationship network includes, when inputting the medical word library, inputting the relationship between each word; Use the words in the medical word library as nodes, connect the words with specific semantic or logical relationships, and assign a weight of 1 to each connection line; The second relationship network includes, for the input text, using a pre-trained large language model to perform context encoding on each word; After obtaining the context embedding vectors of each word, use a neural network module to calculate the semantic relevance weights between each pair of words; The optimization of the associated path includes setting the starting word as the first word in the text whose similarity to any word in the medical word library is greater than to be ; From Analyze the sum of weights with other words in the text, select the word with the largest sum of weights as the next word in the path, and connect them in sequence until the end; The weight sum is expressed as: Among them, represents the weighted sum between and represents the semantic relevance weight and ; represents a selection function. If and both have a similarity greater than with any word in the medical word library, the output is , represents the first association relationship weight between and If and there is no satisfaction that the similarity with any word in the medical word library is greater than , the output is 0; E represents an adjustable coefficient.
2. The method for interpreting medical literature based on a large language model according to claim 1, wherein: The word segmentation processing includes using the BioBERT model to segment the content in the text; The medical word library includes inputting the standard expressions and synonyms of all medical specialized terms, and assimilating the standard expression and its synonyms of each term as a matchable word.
3. The method for interpreting medical literature based on a large language model according to claim 2, wherein: Matching the words in the text with the words in the medical word library includes using a word embedding model to convert the words in the text and the words in the medical word library into vector representations; Calculate the i-th word vector in the text and the j-th word vector in the medical word library to obtain the similarity The similarity is expressed as: When a match is successful; each word in the text is sequentially matched one by one with all the words in the medical word library; during the matching process, the similarity between the corresponding standard expression and its synonyms and is calculated one by one, and the minimum value is taken as the final result of the similarity ; For the i-th word vector Set of all word vectors matched = ; Among them, represents the L-th successfully matched word vector in the medical word library; The first association relationship includes, after completing the word matching, obtaining all word vectors in the text whose matching set is not an empty set; arbitrarily taking two word vectors and , for each element in the set matched with , analyze the connection relationship existing between the elements in the set matched with , and assign a weight to the first association relationship by accumulating the connection times Among them, represents the i-th word vector in the text and the association weight with the v-th word vector in the text ; a represents the index of the element in the set matched by the i-th word vector in the text ; b represents the index of the element in the set matched by the v-th word vector in the text ; A represents the number of elements in the set matched by the i-th word vector in the text ; B represents the number of elements in the set matched by the v-th word vector in the text ; represents the i-th word vector in the text the a-th element in the matched set; represents the v-th word vector in the text the b-th element in the matched set; represents the discriminant function for whether there is an association. In the first relational network, if there is a connection between them, it is recorded as 1; if there is no connection between them, it is recorded as 0.
4. The method for interpreting medical literature based on a large language model according to claim 3, wherein: For the input text , after encoding, the context embedding vectors of each word are obtained: Among them, represents the context embedding of the word, which is obtained from the output of the large model and contains the context embedding vector of the semantic information of the word in a specific context; represents the d-th word, where d represents the number of words; represents the specific position of the output of the index model; The formula for calculating the semantic relevance weight is: Among them, represents the semantic relevance weight of the word pair ; and represent a trainable weight matrix; is a bias vector, and σ is an activation function; represents the concatenation of the word and the context embedding vector of the word ; 1 indicates that the words and have the strongest relevance, and 0 indicates no relevance.
5. The method for interpreting medical literature based on a large language model according to claim 4, wherein: The content interpretation includes using the word string in the path as input, evaluating the integrity of the word string through a trained bidirectional LSTM network, and outputting an integrity score T; If T is greater than the set threshold value , it is determined that the input word string is sufficient to form coherent content; otherwise, it is determined that it cannot be recognized. When the determination result of unrecognizable is output, E starts from the initial value and increases by 1 each time. After the increase ends, path optimization is performed again, and the optimization result is interpreted again until the content of the path can be interpreted; Use a pre-trained Seq2Seq generation model, by calling the generation function, input the word string into the model to generate a coherent interpretation text: Among them, is the generated interpretation content, represents the BART model used to generate coherent interpretation content, represents the input word string.
6. A medical literature interpretation system based on a large language model using the method according to any one of claims 1-5, characterized in that: A collection unit that obtains the medical literature to be interpreted and performs word segmentation processing; A word library unit that constructs a relationship network between words for the medical word library to obtain the first relationship network; A first analysis unit that matches the words in the text with the words in the medical word library to obtain the matching results of the text words; according to the first relationship network, determines the first association relationship between each text word and other words in the text, and assigns weights to the first association relationship according to the number of associations; A second analysis unit that constructs a relationship network for the words in the text according to the text content to obtain the second relationship network; Determine the second association relationship between each text word and other words in the text, and assign weights to the second association relationship; An interpretation unit, according to the first relationship network and the weights given by the first relationship network, performs optimization of the association path, and obtains an interpretation result through the interpretation of the content on the path.
7. A computer device, comprising: A memory and a processor; The memory stores a computer program, characterized in that: when the processor executes the computer program, the steps of the method according to any one of claims 1-5 are implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, the steps of the method according to any one of claims 1-5 are implemented.
Citation Information
Patent Citations
Deep reading method and system based on knowledge graph
CN118193748A
System for automated analysis of clinical text for pharmacovigilance
US20160048655A1