Biomedical literature abstract auxiliary writing method and system based on semantic analysis
Through semantic analysis-based methods, including format recognition, deep semantic understanding and Transformer architecture generation abstracts, the problem of difficult-to-handle differences in literature formats and content in the prior art is solved, efficient and diverse literature abstract generation is achieved, and writing efficiency and quality is significantly improved.
Patent Information
- Application Number
- CN202411763861.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art is difficult to effectively analyze and process the content and format differences of different biomedical literatures, and it is impossible to generate abstracts of multiple styles according to user needs, and the system performance is low, resulting in insufficient writing efficiency and quality of document abstracts.
Using a semantic analysis method, efficient analysis and abstract generation of biomedical literature is achieved through the process of format recognition and content extraction, deep semantic understanding, Transformer architecture generation abstracts, user interaction and self-learning optimization models.
It improves the writing efficiency and quality of literature abstracts, can effectively process multiple document formats, generates various styles of abstracts according to user needs, and optimizes system performance through self-learning, which significantly improves the accuracy and user satisfaction of the abstract.
Smart Images

Figure CN120030998A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and natural language processing, and in particular relates to a method and system for assisting writing of biomedical literature abstracts based on semantic analysis. Background Art
[0002] In the biomedical field, researchers need to read and write a large number of English literature abstracts. This process not only consumes a lot of time and energy, but also requires researchers to have a high level of professional knowledge and language expression skills. The quality of literature abstracts directly affects the dissemination of literature and the display of scientific research results. Therefore, how to write literature abstracts efficiently and accurately has become an important issue facing researchers.
[0003] Existing automatic summary generation technologies mainly rely on rules and statistical methods, which are difficult to handle the professional terms and complex sentences in biomedical literature. These technologies usually cannot fully understand the deep semantics of the content of the literature, and the generated summaries often lack accuracy and coherence, and cannot effectively convey the core content of the literature. In addition, when generating summaries, existing technologies are difficult to adapt to the needs of different styles and formats, resulting in insufficient versatility and practicality of the summaries.
[0004] With the development of natural language processing technology, especially the application of deep learning and pre-trained language models (such as BioBERT), new solutions are provided for the generation of biomedical literature summaries. By using these advanced language models, we can understand the content of the literature more deeply, extract key information and core ideas, and generate more accurate and coherent literature summaries.
[0005] However, although these technologies have improved the effect of abstract generation to a certain extent, there are still many challenges in practical applications. For example, the content and format of different documents vary greatly. How to effectively parse and process these documents; how to generate abstracts of various styles according to user needs; how to continuously optimize and improve system performance through user feedback, etc. Therefore, there is an urgent need for a method and system for assisting the writing of biomedical literature abstracts based on semantic analysis to solve the above problems and improve the writing efficiency and quality of literature abstracts. Summary of the invention
[0006] The purpose of the present invention is to provide a method and system for assisting in writing biomedical literature abstracts based on semantic analysis, so as to improve the writing efficiency and quality of literature abstracts; to solve the technical problems in the prior art that the contents and formats of different documents vary greatly, the documents cannot be effectively parsed and processed, and abstracts of various styles cannot be generated according to user needs, and the performance of the summary generation system is low.
[0007] In order to solve the above technical problems, the specific technical solutions of the present invention are as follows:
[0008] A method for assisting writing biomedical literature abstracts based on semantic analysis comprises the following steps:
[0009] Step S1: Format recognition and content extraction of input biomedical literature;
[0010] Step S2: Perform deep semantic understanding on the parsed document content;
[0011] Step S3: Process the document content based on the Transformer architecture to generate a concise and coherent summary;
[0012] Step S4: providing a friendly interface to allow users to input literature, view and edit the generated abstract;
[0013] Step S5: Self-learning, responsible for collecting and analyzing user feedback information and optimizing the model.
[0014] Furthermore, step S1 includes the following steps:
[0015] Step S11: Identify the file format through the file header information or the file extension, and then use the corresponding parsing tool to parse and extract the text content;
[0016] Step S12: using natural language processing technology and rule engine to extract title, author, keywords, introduction, method, result and discussion from the extracted text content;
[0017] Step S13: using a deep learning model to perform classification and recognition processing;
[0018] Step S14: Perform terminology standardization on the extracted information to ensure uniform terminology and format.
[0019] Further, step S2 includes the following steps:
[0020] Step S21: Use the SpaCy tool to perform word segmentation, part-of-speech tagging, and named entity recognition preprocessing operations on the document content;
[0021] Step S22: using a pre-trained biomedical language model to perform deep semantic analysis on the document content;
[0022] Step S23: Construct a topic model, perform topic classification and cluster analysis on the documents, and extract the main topics and related contents.
[0023] Further, step S3 includes the following steps:
[0024] Step S31: Select a suitable summary template according to user needs;
[0025] Step S32: Use the encoder of the Transformer model to perform semantic encoding, generate hidden state representation, and generate a summary through the decoder.
[0026] Step S33: proofread the summary, proofread the generated summary in terms of grammar and logic to ensure that the summary is accurate and coherent.
[0027] Further, step S32 includes the following steps:
[0028] Step S321: Transformer encoder, using the embedding provided by the pre-trained language model, inputs into the Transformer encoder for semantic encoding.
[0029] The Transformer encoder consists of a multi-head self-attention mechanism and a feed-forward neural network:
[0030] H = TransformerEncoder(E)
[0031] Among them, E is the embedded representation of the document content generated by the pre-trained language model, TransformerEncoder represents the Transformer encoder, and H is the hidden state output by the Transformer encoder. The calculation of the Transformer encoder mainly involves the calculation formula of multi-head self-attention, as shown below:
[0032]
[0033] MultiHead(Q,K,V)=Concat(head 1 , ..., head i ,...,head h )W o
[0034] Among them, Attention(Q, K, V) represents the core operation of the attention mechanism, Q represents the query matrix, K represents the key matrix, V represents the value matrix, and d k represents the dimension of the key vector, softmax represents the normalized exponential function, T represents transpose, MultiHead(Q, K, V) represents multi-head attention, Concat represents concatenation, head i represents the i-th attention matrix, h represents the number of attention heads, and W O represents the learnable parameter matrix.
[0035]
[0036] Among them, W i Q is the learnable parameter matrix.
[0037] The encoder layer of the present invention can be expressed as:
[0038] EncLayer(H)=LayerNorm(H+MultiHead(H,H,H))
[0039] EncLayer(H)=LayerNorm(H+FFN(H))
[0040] Among them, FFN(H)=max(0,HW 1 +b 1 )W 2 +b 2 is a feedforward neural network. max represents the ReLU activation function, W 1 , W 2 represents the weight matrix, b 1 , b 2 represents the bias, LayerNorm represents the normalization operation of the normalization layer, and EncLayer(H) represents the encoding operation of the encoding layer.
[0041] Step S322: Transformer decoder, passing the output of the encoder to the Transformer decoder to generate a summary. The Transformer decoder also includes a multi-head self-attention mechanism, an encoder-decoder attention mechanism, and a feedforward neural network:
[0042] S = TransformerDecoder(H,T)
[0043] Where H is the hidden state of the encoder, T is the summary template, S is the generated summary, and TransformerDecoder represents the Transformer decoder.
[0044] Similar to the encoder layer, the decoder layer of the present invention can be expressed as:
[0045] DecLayer(H,E)=LayerNorm(H+MultiHead(H,H,H))
[0046] DecLayer(H,E)=LayerNorm(H+Attention(H,E,E))
[0047] DecLayer(H,E)=LayerNorm(H+FFN(H))
[0048] Among them, H comes from the decoder, E comes from the encoder, and DecLayer(H, E) represents the decoding operation of the decoding layer.
[0049] Further, step S4 includes the following steps:
[0050] Step S41: Document input. Users can input document files by dragging and dropping or uploading. The system supports multiple file formats.
[0051] Step S42: Summary editing, the user can view, edit and modify the generated summary, and the system provides real-time feedback and suggestions.
[0052] Step S43: Summary preview, the user can preview summary effects of different styles and select the most suitable summary version.
[0053] The present invention also provides a biomedical literature abstract auxiliary writing system based on semantic analysis, which is used to implement the method described in any one of claims 1-6, including a literature parsing module, a semantic analysis module, a summary generation module, a user interaction module, and a self-learning module.
[0054] Furthermore, the document parsing module is used to perform format recognition and content extraction on the input biomedical documents;
[0055] The semantic analysis module is used to conduct in-depth semantic understanding of the parsed document content;
[0056] The summary generation module processes the document content based on the Transformer architecture to generate concise and coherent summaries;
[0057] The user interaction module provides a friendly interface, allowing users to input literature, view and edit the generated abstracts;
[0058] The self-learning module is responsible for collecting and analyzing user feedback information and optimizing the model.
[0059] Compared with the prior art, the present invention has the following beneficial technical effects: the method and system for assisting writing biomedical literature abstracts based on semantic analysis provided by the present invention can deeply understand and analyze the content of the literature when generating literature abstracts by using pre-trained biomedical language models and deep learning technology, and obtain high-quality abstract data. These abstract data are then used for processing in the abstract generation module. Since these abstract data are close to the final required literature abstract content, the abstract generation process can be accelerated when generating literature abstracts, thereby greatly reducing the time and energy of manually writing abstracts, reducing resource consumption and shortening generation time.
[0060] (1) The document parsing module in the present invention can efficiently identify and process a variety of document formats (such as PDF, DOCX, TXT) and extract structured information from them. This process uses natural language processing technology and rule engines to ensure that the key information extracted from the document (such as title, author, keywords, etc.) is accurate. By using a biomedical terminology dictionary for term standardization, the consistency and accuracy of the terminology is guaranteed, thereby improving the quality of subsequent semantic analysis and summary generation.
[0061] (2) The semantic analysis module in the present invention uses deep learning and pre-trained biomedical language models (such as BioBERT) to conduct in-depth semantic understanding of the content of the document. Through pre-processing operations such as word segmentation, part-of-speech tagging, and named entity recognition, sentence vector representations are generated, and core ideas and key information are identified. A topic model (such as LDA) is constructed for topic classification and cluster analysis to ensure that the extracted main topics and related content are accurate. This module significantly improves the depth and accuracy of understanding of the content of the document, providing a solid foundation for generating high-quality summaries.
[0062] (3) The summary generation module in the present invention is based on the Transformer architecture. It can select appropriate summary templates (such as concise, detailed, and professional) according to user needs, and generate concise, coherent, and accurate summaries through encoders and decoders. The generated summaries are proofread for grammar and logic to ensure the coherence and accuracy of the content, and can effectively convey the core ideas of the document. This module significantly improves the quality and readability of the summaries.
[0063] (4) The user interaction module in the present invention provides a friendly interface, allowing users to input document files by dragging or uploading files, and supports multiple formats. Users can preview and edit the generated abstracts in real time and choose abstract versions of different styles. The system provides real-time feedback and suggestions to help users generate the most satisfactory abstracts. This module enhances the user experience and improves the practicality and flexibility of the system.
[0064] (5) The self-learning module continuously optimizes the model by collecting and analyzing user feedback information to improve the accuracy of summary generation and user satisfaction. The system records user editing operations and feedback and uses this data to optimize the model through supervised learning methods. The system regularly upgrades the model and methods based on the latest technological developments and user feedback to ensure the advancement and practicality of processing biomedical literature abstract generation. This module enables the system to self-improve and maintain long-term effectiveness and advancement.
[0065] In summary, the present invention significantly improves the writing efficiency and quality of literature abstracts through the synergistic effect of five modules, solves many difficult problems in the prior art, and has broad application prospects and important practical significance. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.
[0067] Figure 1 The figure is a flow chart of the method for assisting writing biomedical literature abstracts based on semantic analysis of the present invention.
[0068] Figure 2 It is a schematic diagram of the document parsing process of the present invention.
[0069] Figure 3 It is a schematic diagram of the semantic analysis process of the present invention.
[0070] Figure 4 A schematic diagram of the process of generating an abstract for the present invention.
[0071] Figure 5 It is a schematic diagram of the user interaction flow of the present invention.
[0072] Figure 6 It is a self-learning flow chart of the present invention.
[0073] Figure 7 It is a schematic diagram of the structure of the biomedical literature abstract auxiliary writing system based on semantic analysis of the present invention.
[0074] Figure 8 It is a schematic diagram of a UI interface in an embodiment. DETAILED DESCRIPTION
[0075] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0076] In order to improve the writing efficiency and quality of literature abstracts, the present invention proposes a method for assisting writing biomedical literature abstracts based on semantic analysis. Figure 1 As shown, the following steps are included:
[0077] Step S1: Format recognition and content extraction of input biomedical literature.
[0078] First, identify the document format (such as PDF, DOCX, TXT) through the file header information or file extension, and use the corresponding parsing tools to extract the text content. Secondly, use natural language processing technology to extract structured information from the document, such as title, author, keywords, introduction, method, results and discussion. Finally, the extracted information is standardized and mapped and standardized using biomedical terminology dictionaries (such as UMLS, MeSH) to ensure the consistency of terminology and expression, such as Figure 2 shown.
[0079] Step S11: Identify the file format through the file header information or the file extension, and then use the corresponding parsing tool to perform parsing and extract the text content.
[0080] The file extension is the suffix in the file name that identifies the file type, such as .pdf, .docx, .txt, etc. By looking at the file extension, the system can quickly determine the format of the file and select the appropriate parser or tool to process the file. For example: .pdf indicates the PDF document format, and the PDF parsing tool can be used to extract the document content. .docx indicates the Microsoft Word document format, and the corresponding processing tool can be used for parsing. .txt indicates a plain text file, and the text content can be read directly.
[0081] Sometimes the file extension may be inaccurate or modified. In this case, the file header information (also called file signature or magic number) can be used to more reliably identify the file format. The file header information is a specific identifier in the first few bytes of the file when it is stored, which is used to indicate the actual type of the file. Different types of files usually have unique header identifiers. For example, the header of a PDF file usually starts with %PDF. The header of a ZIP format file usually starts with PK (.docx, .xlsx are also based on the ZIP format). The header of a JPEG image file starts with FFD8.
[0082] Use the corresponding parsing tool to parse and extract the text content in the following way:
[0083]
[0084] Among them, F is the input file, format(F) indicates the return file format, parsePDF indicates parsing PDF files, parseDOCX(F) indicates parsing DOCX files, parseTXT indicates parsing text files, and f format (F) represents the text content extracted after parsing the input document.
[0085] Step S12: Using natural language processing technology and rule engines, extract the title, author, keywords, introduction, method, results and discussion from the extracted text content.
[0086] Initial extraction can be performed through regular expression-based methods, followed by classification and recognition using deep learning models.
[0087] Specifically, the processing steps for preliminary extraction based on the regular expression method are as follows:
[0088] Step S121: Title extraction. The title is usually located at the top of the document and has a specific format (such as large font, bold). You can use a regular expression to match the title format:
[0089] Title=match(TitlePatterm,Document)
[0090] Wherein, Title represents the extracted title, match represents the matching function, TitlePattern is the regular expression used to match the title, and Document represents the text content, i.e., the text content extracted after parsing the input document in step S11. format (F).
[0091] Step S122: Author extraction. The author list usually follows the title, and the format may include a comma-separated list of names. Regular expressions can be used for matching and extraction:
[0092] Authors=match(AuthorPattern,Doccmemt)
[0093] Among them, Authors represents the extracted authors, and AuthorPattern is the regular expression used for authors.
[0094] Step S123: keyword extraction. Keywords are usually before the abstract or introduction, marked as "keywords" or "Keywords". Match the keyword part through regular expressions:
[0095] Keywords=match(KeywordsPattern,Document)
[0096] Among them, Keywords represents the extracted keywords, and KeywordsPattern is a regular expression used to match the keyword part.
[0097] Step S124: Extract the introduction. The introduction is usually marked with "Introduction" or "Introduction" and ends before the "Method" section. The introduction can be extracted by token matching and text segmentation:
[0098] Introduction=extract(ntroductionmPattern,Document)
[0099] Among them, Introduction represents the extracted introduction, extract represents the extraction function, and IntroductionPattern is a marker used to match the introduction part.
[0100] Step S125: Method extraction. The method part is usually identified by "Method" or "Methods" and ends before the "Result" part. The method can be extracted by token matching and text segmentation:
[0101] Methods=extract(MethodsPattern,Documemt)
[0102] In the method, Methods represents the extracted method, and MethodsPattern is a marker used to match the method part.
[0103] Step S126: Extract results. The results section is usually marked with "results" or "Results" and ends before the "discussion" section. Results can be extracted through tag matching and text segmentation:
[0104] Results=extract(ResultsPattern,Document)
[0105] Among them, Results represents the extracted results, and ResultsPattern is a tag used to match the result part.
[0106] Step S127: Extract discussion. The discussion part is usually marked with "Discussion" or "Discussion" and ends at the end of the document. Discussion can be extracted through tag matching and text segmentation:
[0107] Discussion=etract(DiscussionPattern,Document)
[0108] Among them, Discussion represents the extracted discussion, and DiscussionPattern is a tag used to match the discussion part.
[0109] Step S13: Use the deep learning model to perform classification and recognition processing, the steps are as follows:
[0110] Step S131: Text paragraph classification: Based on the preliminary extraction, the pre-trained language model is used to classify and identify the document paragraphs.
[0111] The deep learning model can use pre-trained language models such as BERT, BioBERT, etc. These models can understand the semantics of the paragraph and identify the structured part to which it belongs.
[0112] The pre-trained language model used in this embodiment is the BERT model. Assume that the content of the input document is D and each paragraph in the document is P. i . Each paragraph P is transformed into i Convert to embedding vector E i :
[0113] E i =BERT(P i )
[0114] Among them, BERT represents the BERT pre-trained language model.
[0115] Step S132: Classification model, using a classification model (such as a bidirectional long short-term memory network (BiLSTM), a convolutional neural network (CNN), etc.) to classify the embedding vector E i Classify and determine the part to which the paragraph belongs (such as title, author, keywords, introduction, method, results and discussion, etc.). The output of the classification model is the label L of the part to which the paragraph belongs. i :
[0116] L i =Classify(E i )
[0117] Among them, Classify is a classification model.
[0118] Step S133: Classification model training. The classification model is trained through supervised learning. The training data set includes document paragraphs and their corresponding labels. The training goal is to minimize the classification loss function The classification loss function is expressed as follows:
[0119]
[0120] Among them, N is the number of training samples, C is the number of label categories, and y ij is the true label, is the predicted probability.
[0121] Step S134: Recognition results: According to the output of the classification model, the paragraphs in the document are classified and annotated to form structured information:
[0122] StructuredDgta={(P i L i )|i=1,2,...,n}
[0123] Among them, StructuredData represents structured information, P i It is a paragraph, L i is the label of the part to which the paragraph belongs, and n represents the total number of label categories.
[0124] Step S14: Standardize the extracted information to ensure uniform terminology and format. Use biomedical terminology dictionaries (such as UMLS, MeSH, etc.) to map and standardize the terms:
[0125] stamdardize(Term)=map(Term,Dictionary)
[0126] Among them, standardize is the term standardization function, map is the mapping operation, Term is the extracted term, and Dictionary is the term dictionary.
[0127] Step S2: Perform deep semantic understanding on the parsed document content.
[0128] First, the document content is preprocessed by word segmentation, part-of-speech tagging, and named entity recognition. Then, a pre-trained biomedical language model (such as BioBERT) is used to generate vector representations of each sentence in the document and identify core ideas and key information. Finally, a topic model (such as Latent Dirichlet Allocation, LDA) is constructed to perform topic classification and cluster analysis on the document content, extracting the main topics and related content, such as Figure 3 shown.
[0129] Step S21: Use the SpaCy tool to perform word segmentation, part-of-speech tagging, and named entity recognition preprocessing operations on the document content. The processing steps are as follows:
[0130] Step S211: Word segmentation is to split the input document content D into word sequences T. In SpaCy, word segmentation is based on predefined language models and rules:
[0131] T = SpaCy.tokenize(D)
[0132] Among them, SpaCy.tokenize represents the word segmentation operation of SpaCy, and the result T after word segmentation is:
[0133] T=[t 1 , t 2 ,...ti ,...,t n ]
[0134] Among them, t i represents the i-th word in the document, and n represents the total number of words.
[0135] Step S212: Part-of-speech tagging is to assign appropriate part-of-speech tags to each word. Part-of-speech tags can be nouns, verbs, adjectives, etc. Part-of-speech tagging is performed using SpaCy's pre-trained model:
[0136] POS(t i ) = SpaCy.pos t ag(t i )
[0137] Among them, SpaCy.postag represents the part-of-speech tagging operation of SpaCy, POS(t i ) represents the part-of-speech tagging result of the i-th word.
[0138] Step S213: Named entity recognition is to identify entities with specific meanings in the text and assign corresponding labels to them. Use SpaCy's NER model for named entity recognition:
[0139] NER(t i ) = SpaCy.ner(t i )
[0140] Among them, SpaCy.ner represents SpaCy's named entity recognition operation, NER(t i ) represents the named entity recognition result of the i-th word. The results of named entity recognition of the document are as follows:
[0141] T NER =[(t 1 .NER(t 1 )),...,(t i .NER(t i )),...,(t n .NER(t n ))]
[0142] Step S214: synthesize the preprocessing results, and combine the results of word segmentation, part-of-speech tagging, and named entity recognition to form complete information for each word:
[0143] PreprocessedData
[0144] =[(t 1 .POS(t 1 ),NER(t 1 )),...,(ti .POS(t i ), NER(t i )),...,(t n .POS(t n ),NER(t n ))]
[0145] The comprehensive preprocessing results can be represented as a structured dataset containing each word, its part-of-speech tag, and named entity tag:
[0146] PreprocessedData = {(t i .POS(t i ),NER(t i ))|i=1,2,...,n}
[0147] This dataset provides detailed semantic information for subsequent semantic analysis and summary generation. By using SpaCy for word segmentation, part-of-speech tagging, and named entity recognition, key information in the document can be effectively extracted, improving the accuracy and quality of summary generation.
[0148] Step S22: Use a pre-trained biomedical language model (such as BioBERT) to perform deep semantic analysis on the document content. The model calculates the vector representation of each sentence in the document and identifies the core ideas and key information. The deep semantic analysis of the document content is performed as follows:
[0149] Step S221: semantic embedding, using a pre-trained language model (such as BERT) to convert each sentence in the document into a semantic embedding vector that captures the semantic information of the sentence.
[0150] Assume that the document content is processed into sentence sequence S = {s 1 , ..., s i , ..., s m}, m represents the number of sentences processed, and each sentence s i Convert to sentence embedding vector E i :
[0151] E i = BioBERT(s i )
[0152] Among them, E i It is a sentence i The semantic embedding vector of , BioBERT represents the BioBERT pre-trained biomedical language model.
[0153] Step S222: Sentence relationship identification, by analyzing the relationship between semantic embedding vectors, identifying the logical and semantic relationship between sentences. This can be achieved by calculating the similarity or distance between sentence embedding vectors. i and j The similarity between sim(E i , E j ) can be calculated using cosine similarity:
[0154]
[0155] Among them, E i 、E j represents the semantic embedding vector, ||·|| represents the norm, and based on the similarity, a sentence relationship graph can be constructed to represent the correlation between sentences.
[0156] Step S223: Keyword and topic extraction, using a deep learning model to identify keywords and topics in the document. This can be achieved by analyzing the clustering characteristics of the semantic embedding vector. Assume that all sentence embedding vectors in the document {E 1 ,...,E i ,...,E m}Cluster into several topics, where m represents the number of sentences and each topic represents a core content. Topic extraction can be performed using clustering algorithms (such as K-means):
[0157] C 1 ,...,C i ,...,C k =K-means(E 1 , ...E i , ..., E m , k)
[0158] Among them, C i represents the i-th topic, K-means represents the K-means clustering algorithm, and k is the number of topics.
[0159] Step S224: Dependency Parsing, using dependency parsing to identify the grammatical structure and dependency relationship within the sentence, and further understand the semantics of the sentence. SpaCy provides an efficient dependency parsing tool. Assume that the sentence s i The dependency tree of i , each node v and edge e represents the words of the sentence and the grammatical relationship between them:
[0160] T i = SpaCy.dependency parse(s i )
[0161] Among them, SpaCy.dependency parse represents the dependency syntactic analysis operation of SpaCy. Through dependency syntactic analysis, the subject-verb-object structure and modification relationship of the sentence can be identified.
[0162] Step S225: Comprehensive semantic analysis results, combining the results of word segmentation, part-of-speech tagging, named entity recognition, semantic embedding, sentence relationship recognition, keyword and topic extraction, and dependency syntax analysis to form a deep semantic analysis result of the document. The comprehensive semantic analysis result is expressed as:
[0163] SemamticAnalysis = {(s i , E i ,sim(E i , E j ), C i , T i )|i=1,2,...,m}
[0164] Among them, s i It is a sentence, E i is the sentence embedding vector, sim(E i , E j ) is the sentence similarity, C i It is the theme, T i It is a dependency tree.
[0165] Deep semantic analysis comprehensively understands the semantic information of the document through multi-level and multi-angle analysis, providing strong support for subsequent abstract generation.
[0166] Step S23: Construct a topic model (such as Latent Dirichlet Allocation, LDA), perform topic classification and cluster analysis on the documents, and extract the main topics and related content:
[0167]
[0168] Among them, θ is the document topic distribution, α is the prior of topic distribution, and z i is the theme, i is a word, β is the word-topic distribution, P(θ|α) represents the probability of θ appearing under the condition α, and P(z i |θ) means that z i The probability of occurrence, P(w i |z i , β) represents the i , w under β condition i Probability of occurrence.
[0169] Step S3: Process the document content based on the Transformer architecture to generate a concise and coherent summary.
[0170] First, select an appropriate summary template (such as concise, detailed, or professional) based on user needs. Then, use the Transformer model's encoder to perform semantic encoding, generate hidden state representations, and generate summaries through the decoder. The generated summaries are proofread for grammar and logic to ensure that the content is accurate. Figure 4 shown.
[0171] Step S31: Select a suitable summary template according to user needs. The summary templates include concise, detailed, and professional:
[0172]
[0173] Among them, u represents the template type selected by the user, Template(u) represents the summary template selected according to user needs, concise represents concise type, detailed represents detailed type, and professional represents professional type. Concise represents concise summary template, Detailed represents detailed summary template, and Professional represents professional summary template.
[0174] Step S32: Use the encoder of the Transformer model to perform semantic encoding, generate hidden state representation, and generate a summary through the decoder.
[0175] The Transformer model for generating a summary includes an encoder and a decoder. The encoder and the decoder of the present invention are described below respectively:
[0176] Step S321: Transformer encoder, using the embedding provided by the pre-trained language model, inputs into the Transformer encoder for semantic encoding.
[0177] The Transformer encoder consists of a multi-head self-attention mechanism and a feed-forward neural network:
[0178] H = TransformerEncoder(E)
[0179] Among them, E is the embedded representation of the document content generated by the pre-trained language model, TransformerEncoder represents the Transformer encoder, and H is the hidden state output by the Transformer encoder. The calculation of the Transformer encoder mainly involves the calculation formula of multi-head self-attention, as shown below:
[0180]
[0181] MultiHead(Q,K,V)=Concat(head 1 ,...,head i ,...,head h )W o
[0182] Among them, Attention(Q, K, V) represents the core operation of the attention mechanism, Q represents the query matrix, K represents the key matrix, V represents the value matrix, and d k represents the dimension of the key vector, softmax represents the normalized exponential function, T represents transpose, MultiHead(Q, K, V) represents multi-head attention, Concat represents concatenation, head i represents the i-th attention matrix, h represents the number of attention heads, and W O represents the learnable parameter matrix.
[0183]
[0184] Among them, W i Q , is the learnable parameter matrix.
[0185] The encoder layer of the present invention can be expressed as:
[0186] EncLayer(H)=LayerNorm(H+MultiHead(H,H,H))
[0187] EncLayer(H)=LayerNorm(H+FFN(H))
[0188] Among them, FFN(H)=max(0,HW 1 +b 1 )W 2 +b 2 is a feedforward neural network. max represents the ReLU activation function, W 1 , W 2 represents the weight matrix, b 1 , b 2 represents the bias, LayerNorm represents the normalization operation of the normalization layer, and EncLayer(H) represents the encoding operation of the encoding layer.
[0189] Step S322: Transformer decoder, passing the output of the encoder to the Transformer decoder to generate a summary. The Transformer decoder also includes a multi-head self-attention mechanism, an encoder-decoder attention mechanism, and a feedforward neural network:
[0190] S = TransformerDecoder(H,Te)
[0191] Among them, H is the hidden state of the encoder, Te is the summary template, S is the generated summary, and TransformerDecoder represents the Transformer decoder.
[0192] Similar to the encoder layer, the decoder layer of the present invention can be expressed as:
[0193] DecLayer(H,E)=LayerNorm(H+MultiHead(H,H,H))
[0194] DecLayer(H,E)=LayerNorm(H+Attention(H,E,E))
[0195] DecLayer(H,E)=LayerNorm(H+FFN(H))
[0196] Among them, H comes from the decoder, E comes from the encoder, and DecLayer(H, E) represents the decoding operation of the decoding layer.
[0197] Step S33: Summary proofreading: Proofread the generated summary for grammar and logic (such as Grammarly, LanguageTool) to ensure that the summary content is accurate and coherent.
[0198] The present invention uses a grammar checking tool and a logical consistency checking algorithm:
[0199] Proofread(S)=GrammarCheck(S)+LogicCheck(S)
[0200] Among them, Proofread means proofreading the text, S is the generated summary, GrammarCheck is a grammar check, and LogicCheck is a logic check.
[0201] Abstract proofreading is an important step in generating high-quality literature abstracts. It ensures that the generated abstracts are accurate, coherent, and meet readers' expectations through grammar checking, logical consistency checking, and information accuracy verification.
[0202] Step S4: Provide a friendly interface to allow users to input documents, view and edit generated abstracts. Users can input document files by dragging and dropping or uploading files, and the system supports multiple formats. Users can preview and edit generated abstracts in real time and choose different styles of abstract versions. The system provides real-time feedback and suggestions to help users generate the most satisfactory abstracts and enhance user experience, such as Figure 5 shown.
[0203] Step S41: Document input. Users can input document files by dragging, dropping, uploading, etc. The system supports multiple file formats.
[0204] Step S42: Summary editing, the user can view, edit and modify the generated summary, and the system provides real-time feedback and suggestions.
[0205] Step S43: Summary preview, the user can preview summary effects of different styles and select the most suitable summary version.
[0206] Step S5: Self-learning, responsible for collecting and analyzing user feedback information and optimizing the model. The system records the user's editing operations and feedback, and uses these data to optimize the model through supervised learning methods to improve the accuracy of abstract generation and user satisfaction. The system regularly upgrades models and methods based on the latest technological developments and user feedback to ensure advancement and practicality in processing biomedical literature abstract generation, such as Figure 6 shown.
[0207] Step S51: user feedback collection, recording the user's editing operations and feedback information, including modified content, satisfaction evaluation, etc.
[0208] Step S52: Model optimization: Based on the collected user feedback, the model is optimized using a supervised learning method to improve the accuracy of summary generation and user satisfaction.
[0209] Step S53: Method upgrade: according to the optimized model and the development of new technologies, the method is regularly upgraded to ensure the advancement and practicality of the method.
[0210] In addition, the present invention also proposes a biomedical literature abstract auxiliary writing system based on semantic analysis, including a literature parsing module, a semantic analysis module, an abstract generation module, a user interaction module, and a self-learning module. Figure 7 shown.
[0211] The document parsing module is used to identify the format and extract the content of the input biomedical documents.
[0212] Specifically, the document format (such as PDF, DOCX, TXT) is identified through the file header information or file extension, and the text content is extracted using the corresponding parsing tools. Secondly, natural language processing technology is used to extract structured information from the document, such as title, author, keywords, introduction, method, results and discussion. Finally, the extracted information is subjected to terminology standardization, and biomedical terminology dictionaries (such as UMLS, MeSH) are used for mapping and standardization to ensure the consistency of terminology and expression.
[0213] The semantic analysis module is used to conduct in-depth semantic understanding of the parsed document content.
[0214] Specifically, the semantic analysis module performs preprocessing operations such as word segmentation, part-of-speech tagging, and named entity recognition on the document content. Then, a pre-trained biomedical language model (such as BioBERT) is used to generate vector representations of each sentence in the document and identify core ideas and key information. Finally, a topic model (such as Latent Dirichlet Allocation, LDA) is constructed to perform topic classification and cluster analysis on the document content to extract the main topics and related content.
[0215] The summary generation module processes the document content based on the Transformer architecture to generate concise and coherent summaries.
[0216] Specifically, the summary generation module selects appropriate summary templates (such as concise, detailed, and professional) according to user needs. Then, the encoder of the Transformer model is used for semantic encoding to generate hidden state representations, and the decoder generates a summary. The generated summary is proofread for grammar and logic to ensure the accuracy of the content.
[0217] The user interaction module provides a friendly interface that allows users to input literature, view and edit the generated abstracts.
[0218] like Figure 8 The figure shows a UI interface schematic diagram of the present invention. Users can input document files by dragging or uploading files through the user interaction module. The system supports multiple formats. Users can preview and edit the generated abstracts in real time and choose abstract versions of different styles. The system provides real-time feedback and suggestions to help users generate the most satisfactory abstracts and enhance user experience.
[0219] The self-learning module is responsible for collecting and analyzing user feedback information and optimizing the model.
[0220] The system records the user's editing operations and feedback, and uses this data to optimize the model through supervised learning methods to improve the accuracy of summary generation and user satisfaction. The system regularly upgrades the model and method based on the latest technological developments and user feedback to ensure the advancement and practicality of processing biomedical literature abstract generation.
[0221] It is to be understood that the present invention is described by some embodiments, and it is known to those skilled in the art that various changes or equivalent substitutions may be made to these features and embodiments without departing from the spirit and scope of the present invention. In addition, under the teachings of the present invention, these features and embodiments may be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the scope of protection of the present invention.
Claims
1. A method for assisting writing biomedical literature abstracts based on semantic analysis, characterized in that: The steps include: Step S1: Format recognition and content extraction of input biomedical literature; Step S2: Perform deep semantic understanding on the parsed document content; Step S3: Process the document content based on the Transformer architecture to generate a concise and coherent summary; Step S4: providing a friendly interface to allow users to input literature, view and edit the generated abstract; Step S5: Self-learning, responsible for collecting and analyzing user feedback information and optimizing the model.
2. The method for assisting writing biomedical literature abstracts based on semantic analysis according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: Identify the file format through the file header information or the file extension, and then use the corresponding parsing tool to parse and extract the text content; Step S12: using natural language processing technology and rule engine to extract title, author, keywords, introduction, method, result and discussion from the extracted text content; Step S13: using a deep learning model to perform classification and recognition processing; Step S14: Perform terminology standardization on the extracted information to ensure uniform terminology and format.
3. The method for assisting writing biomedical literature abstracts based on semantic analysis according to claim 2, characterized in that: Step S2 includes the following steps: Step S21: Use the SpaCy tool to perform word segmentation, part-of-speech tagging, and named entity recognition preprocessing operations on the document content; Step S22: using a pre-trained biomedical language model to perform deep semantic analysis on the document content; Step S23: Construct a topic model, perform topic classification and cluster analysis on the documents, and extract the main topics and related contents.
4. The method for assisting writing biomedical literature abstracts based on semantic analysis according to claim 3, characterized in that: Step S3 includes the following steps: Step S31: Select a suitable summary template according to user needs; Step S32: Use the encoder of the Transformer model to perform semantic encoding, generate hidden state representation, and generate a summary through the decoder; Step S33: proofread the summary, proofread the generated summary in terms of grammar and logic to ensure that the summary is accurate and coherent.
5. The method for assisting writing biomedical literature abstracts based on semantic analysis according to claim 4, characterized in that: Step S32 includes the following steps: Step S321: Transformer encoder, using the embedding provided by the pre-trained language model, inputs into the Transformer encoder for semantic encoding; The Transformer encoder consists of a multi-head self-attention mechanism and a feed-forward neural network: H = TransformerEncoder(E) Among them, E is the embedded representation of the document content generated by the pre-trained language model, TransformerEncoder represents the Transformer encoder, and H is the hidden state output by the Transformer encoder; the calculation of the Transformer encoder involves the calculation formula of multi-head self-attention, as shown below: MultiHead(Q,K,V)=Concat(head1,...,head i ,...,head h )W O Among them, Attention(Q, K, V) represents the core operation of the attention mechanism, Q represents the query matrix, K represents the key matrix, V represents the value matrix, and d k represents the dimension of the key vector, softmax represents the normalized exponential function, T represents transpose, MultiHead(Q, K, V) represents multi-head attention, Concat represents concatenation, head i represents the i-th attention matrix, h represents the number of attention heads, and W O represents the learnable parameter matrix; Among them, W i Q , is a learnable parameter matrix; The encoder layer is represented as: EncLayer(H)=LayerNorm(H+MultiHead(H, H, H)) EncLayer(H)=LayerNorm(H+FFN(H)) Among them, FFN(H)=max(0,HW1+b1)W2+b2 is a feedforward neural network; max represents the ReLU activation function, W1 and W2 represent weight matrices, b1 and b2 represent biases, LayerNorm represents the normalization operation of the normalization layer, and EncLayer(H) represents the encoding operation of the encoding layer; Step S322: Transformer decoder, passing the output of the encoder to the Transformer decoder to generate a summary; the Transformer decoder also includes a multi-head self-attention mechanism, an encoder-decoder attention mechanism, and a feedforward neural network: S = TransformerDecoder(H, T) Where H is the hidden state of the encoder, T is the summary template, S is the generated summary, and TransformerDecoder represents the Transformer decoder; Similar to the encoder layer, the decoder layer is represented as: DecLayer(H,E)=LayerNorm(H+MultiHead(H,H,H)) DecLayer(H,E)=LayerNorm(H+Attention(H,E,E)) DecLayer(H,E)=LayerNorm(H+FFN(H)) Among them, H comes from the decoder, E comes from the encoder, and DecLayer(H, E) represents the decoding operation of the decoding layer.
6. The method for assisting writing biomedical literature abstracts based on semantic analysis according to claim 1, characterized in that: Step S4 includes the following steps: Step S41: Document input, the user can input document files by dragging and dropping or uploading, and the system supports multiple file formats; Step S42: Summary editing, the user can view, edit and modify the generated summary, and the system provides real-time feedback and suggestions; Step S43: Summary preview, the user can preview summary effects of different styles and select the most suitable summary version.
7. A biomedical literature abstract writing system based on semantic analysis, characterized in that: Used to implement the method described in any one of claims 1-6, including a document parsing module, a semantic analysis module, a summary generation module, a user interaction module, and a self-learning module.
8. The biomedical literature abstract assisting writing system based on semantic analysis according to claim 7 is characterized in that: The document parsing module is used to identify the format and extract the content of the input biomedical documents; The semantic analysis module is used to conduct in-depth semantic understanding of the parsed document content; The summary generation module processes the document content based on the Transformer architecture to generate concise and coherent summaries; The user interaction module provides a friendly interface, allowing users to input literature, view and edit the generated abstracts; The self-learning module is responsible for collecting and analyzing user feedback information and optimizing the model.
9. The biomedical literature abstract assisting writing system based on semantic analysis according to claim 8, characterized in that: The semantic analysis module performs preprocessing operations on the document content, including word segmentation, part-of-speech tagging, and named entity recognition. Then, it uses a pre-trained biomedical language model to generate vector representations of each sentence in the document and identify core ideas and key information. Finally, it constructs a topic model to perform topic classification and cluster analysis on the document content to extract the main topics and related content.
10. The biomedical literature abstract assisting writing system based on semantic analysis according to claim 8, characterized in that: The summary generation module selects a suitable summary template according to user needs, then uses the encoder of the Transformer model for semantic encoding to generate hidden state representation, and generates a summary through the decoder. The generated summary is grammatically and logically proofread.
Citation Information
Patent Citations
Method for generating guided text abstract based on Transformer
CN111897949A
Chinese patent abstract rewriting method
CN112417853A
GENERATIVE AUTOMATIC abstracting METHOD BASED ON BERT AND EXTERNAL KNOWLES
CN114398478A
Multi-document scientific abstract generation method based on topic knowledge graph joint enhancement
CN116821371A
Method and device for generating medical text abstract
CN118333038A
Cited By
Multi-agent collaborative clinical scientific research implementation method and device and computer equipment
CN122154948A
Multi-agent collaborative clinical research implementation method and device and computer equipment
CN122154948B