Outline content generation method and system combined with natural language processing
By combining natural language processing technology with semantic analysis and literature database retrieval, personalized answers to scientific research questions are generated, which solves the problems of inaccurate and untimely answers in existing technologies and improves the quality of scientific research question answers.
Patent Information
- Application Number
- CN202511259826.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-09-04
AI Technical Summary
Existing technologies for answering biomedical research questions suffer from poor timeliness, inaccurate answers, and an inability to provide personalized solutions.
By combining natural language processing technology, semantic parsing is performed to obtain a set of question features. Then, a biomedical literature database is accessed to extract structured knowledge units and assign weights to them, generating personalized answer text.
It enables accurate, comprehensive, and personalized answers to scientific research questions, improving the quality of text-based question answers and user experience in the scientific research field.
Smart Images

Figure CN120763320B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of generative artificial intelligence, in particular to a natural language processing combined outline content generation method and system. BACKGROUND
[0002] In the field of biomedical scientific research communication, the scientific research academic community provides a platform for communication and sharing for scientific research professionals. Users often raise various basic scientific research questions in the community to seek professional answers. However, the existing technical solutions for answering these scientific research questions have many shortcomings.
[0003] On the one hand, traditional manual answering methods rely on the knowledge and experience of individual experts. Due to the limited time and energy of experts, they cannot answer a large number of user questions in a timely manner, resulting in poor timeliness of question answering. On the other hand, some automatic answering systems based on simple keyword retrieval can only match relevant literature fragments in a biomedical literature database according to the keywords input by the user, but cannot accurately understand the semantics of the user's question, cannot provide accurate answers to the core points of the question, and the returned content may contain a large amount of irrelevant information, requiring the user to spend a lot of time screening useful information. In addition, existing answering methods do not fully consider the text feature quantization of different types of information for users, and cannot provide personalized answering schemes for users. SUMMARY
[0004] Therefore, the purpose of the present application is to provide a natural language processing combined outline content generation method and system.
[0005] According to a first aspect of the present application, a natural language processing combined outline content generation method is provided, which comprises:
[0006] Receiving a scientific research field text, performing a basic scientific research based semantic analysis process on the scientific research field text to obtain a problem feature set corresponding to the scientific research field text, the problem feature set including relevant gene / protein names, scientific research question type classification and text feature quantization parameters;
[0007] Based on the problem feature set, calling a biomedical literature database interface to perform a basic scientific research literature knowledge retrieval process, obtaining public scientific research literature resources matched with the problem feature set, the public scientific research literature resources including core viewpoint abstracts, basic experimental method descriptions and molecular mechanism result fragments of multiple basic scientific research literatures;
[0008] Performing a knowledge unit extraction process on the public scientific research literature resources to extract scientific research knowledge units with independent semantics from the core viewpoint abstracts, basic experimental method descriptions and molecular mechanism result fragments of the literatures to generate a structured knowledge unit set;
[0009] The structured knowledge unit set is associated and matched with the problem feature set, weight distribution is performed on the scientific research knowledge units in the structured knowledge unit set according to the scientific research problem type classification and the text feature quantization parameter, and a weighted knowledge integration result is generated;
[0010] A basic scientific research problem answer text is generated according to the weighted knowledge integration result, and the basic scientific research problem answer text is fed back to a scientific research field user interaction interface for display.
[0011] According to a second aspect of the present application, a natural language processing combined outline content generation system is provided, which comprises a machine readable storage medium and a processor, the machine readable storage medium stores machine executable instructions, and the processor executes the machine executable instructions, and the natural language processing combined outline content generation system realizes the natural language processing combined outline content generation method described above.
[0012] According to a third aspect of the present application, a computer readable storage medium is provided, which stores computer executable instructions, and when the computer executable instructions are executed, the natural language processing combined outline content generation method described above is realized.
[0013] According to any one of the above aspects, the technical effect of the present application is that:
[0014] By combining natural language processing technology, deep analysis of scientific research field user input scientific research text is realized. First, the scientific research field text is subjected to semantic analysis based on basic scientific research, and the problem feature set is accurately obtained, covering related gene / protein names, scientific research problem type classification and text feature quantization parameters. Then, based on the problem feature set, the biomedical literature database interface is called for knowledge retrieval, which can quickly locate the basic scientific research literature resources highly matched with the problem, avoid the interference of irrelevant literature, and improve the retrieval efficiency. Then, the knowledge unit extraction is performed on the public scientific research literature resources to generate the structured knowledge unit set, so that the scientific research knowledge is more standardized and easy to process. The structured knowledge unit set is associated and matched with the problem feature set, and weight distribution is performed according to the scientific research problem type and the text feature quantization parameter to generate a weighted knowledge integration result. The differences in different information needs of users are fully considered, individualized knowledge integration is realized, and finally, a basic scientific research problem answer text is generated according to the weighted knowledge integration result and is fed back for display, which provides accurate, comprehensive and personalized answers for users, greatly improves the problem solving quality of scientific research field text and user experience. BRIEF DESCRIPTION OF DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some of the embodiments of the present application, and therefore should not be regarded as limiting the scope. Those skilled in the art can obtain other related drawings according to these drawings without any creative effort.
[0016] Figure 1 The flowchart of the outline content generation method provided by the embodiments of the present application is shown.
[0017] Figure 2 The component structure diagram of the outline content generation system for implementing the above-mentioned outline content generation method provided by the embodiments of the present application is shown. DETAILED DESCRIPTION
[0018] The embodiments of the present application will be described below in conjunction with the drawings in the present application. It should be understood that the embodiments described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not limit the technical solutions of the embodiments of the present application.
[0019] Those skilled in the art can understand that, unless specifically stated, the singular forms "a", "an" and "the" used herein include plural forms. It should be further understood that the terms "include" and "contain" used by the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude other features, information, data, steps, operations, elements, components and / or their combinations supported by the present technical field. It should be understood that when it is said that one element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein can include wireless connection or wireless coupling. The term "and / or" used herein means that at least one of the items defined by the term, for example, "A and / or B" can be implemented as "A", or as "B", or as "A and B".
[0020] In order to make the purposes, technical solutions and advantages of the present application more clear, the embodiments of the present application will be further described in detail below in conjunction with the drawings. The technical solutions of the embodiments of the present application and the technical effects generated by the technical solutions of the present application will be described below through the description of several exemplary embodiments. It should be pointed out that the following embodiments can be mutually referenced, borrowed or combined, and the same terms, similar features and similar implementation steps in different embodiments will not be described repeatedly.
[0021] Figure 1 The flowchart of the outline content generation method combined with natural language processing provided by the embodiments of the present application is shown. It should be understood that in other embodiments, the order of some steps of the outline content generation method combined with natural language processing of the present embodiment can be shared according to actual needs, or some steps can be omitted or maintained. The detailed steps of the outline content generation method combined with natural language processing include:
[0022] The present embodiment takes the basic research problem description of “molecular mechanism of type 2 diabetes and hypertension and its basic research on the impact on kidney function” proposed in the field of scientific research as an example, and elaborates the specific implementation process of the outline content generation method combined with natural language processing.
[0023] Step S110: receiving the basic research problem description input in the field of scientific research text, performing semantic analysis processing on the problem description, obtaining a corresponding problem feature set, and the problem feature set contains relevant gene / protein names, scientific research problem type classification and text feature quantization parameters.
[0024] In the field of scientific research text, after the user inputs the above problem description through the interactive interface, the relevant processing module receives the text, and then starts the semantic analysis processing flow to obtain a feature set that can accurately reflect the core of the problem.
[0025] Step S111: receiving the basic research problem description input in the field of scientific research text, performing format standardization processing on the text, removing special symbols and irrelevant characters in the text, and retaining valid text content containing basic research terms.
[0026] After receiving the basic research problem description “molecular mechanism of type 2 diabetes and hypertension and its basic research on the impact on kidney function” input by the user, the first text format standardization processing is performed. In this process, the system will scan the entire text and identify possible special symbols such as “#”, “@”, “$” and other irrelevant characters such as extra spaces, line breaks or meaningless placeholders.
[0027] Then, these identified special symbols and irrelevant characters are removed one by one. For example, if the text contains “2 type diabetes # and accompanied by hypertension, its molecular mechanism and related research on the impact on kidney function @”, after processing, it will become “2 type diabetes and accompanied by hypertension, its molecular mechanism and related research on the impact on kidney function”.
[0028] During the removal process, special attention should be paid to not mistakenly deleting characters related to basic research terms to ensure that the remaining text content is all valid parts containing basic research terms.
[0029] Step S112: Perform professional word segmentation processing on the effective text content, and call a professional word segmentation tool to split the effective text content into a plurality of scientific research word units, the scientific research word units including relevant gene / protein names, verbs, and adjectives.
[0030] After completing the text format standardization processing, the effective text content is obtained, and then professional word segmentation processing is performed. A special professional word segmentation tool is called, which internally stores a large number of scientific research professional vocabularies and word segmentation rules, and can accurately segment scientific research texts.
[0031] For the effective text "Basic research on the molecular mechanism of type 2 diabetes and hypertension and its impact on kidney function", the word segmentation tool will split it according to the language habits and vocabulary characteristics of the scientific research field. The scientific research word units obtained after splitting include "type 2 diabetes", "and", "hypertension", "the", "molecular mechanism", "and", "it", "on", "kidney function", "impact", "the", "basic research", etc.
[0032] Among them, scientific terms such as "type 2 diabetes", "hypertension", "kidney function", "molecular mechanism", "basic research", verbs such as "impact", and adjectives such as "related" are the basic elements for subsequent processing.
[0033] Step S113: Perform entity recognition processing on the plurality of scientific research word units, identify entity words belonging to the categories of relevant gene / protein names, cell types, molecular markers, and experimental methods in the plurality of scientific research word units through a pre-trained entity recognition model, and form an entity name set.
[0034] After obtaining the plurality of scientific research word units, a pre-trained entity recognition model is used to accurately locate various scientific research entities.
[0035] Step S1131: Input the plurality of scientific research word units into the pre-trained entity recognition model, the entity recognition model including an embedding layer, a bidirectional recurrent neural network layer, and a conditional random field layer.
[0036] The plurality of scientific research word units obtained by word segmentation, such as "type 2 diabetes", "and", "hypertension", "the", "molecular mechanism", "and", "it", "on", "kidney function", "impact", "the", "basic research", etc. are input into the pre-trained entity recognition model.
[0037] The structure of the entity recognition model is composed of an embedding layer, a bidirectional recurrent neural network layer, and a conditional random field layer connected in turn, and the information flow and processing between the layers are realized through a specific parameter transmission mechanism, which together completes the entity recognition task.
[0038] Step S1132: converting each scientific term unit into a low-dimensional dense term vector through the embedding layer, so as to retain the semantic information of the scientific term.
[0039] After the embedding layer receives the input scientific term unit, it can convert each term unit into a low-dimensional dense term vector according to the internal mapping rules and the trained parameters. For example, “type 2 diabetes” will be converted into a vector of a specific dimension, and the value of each element in the vector is determined through training, which can represent the semantic information of “type 2 diabetes” in a low-dimensional space, so that the vector of “type 2 diabetes” maintains a reasonable distance relationship with the vectors of other related terms such as “hypertension” in the space, so as to reflect the semantic association between them.
[0040] Similarly, scientific term units such as “hypertension” and “kidney function” will also be converted into corresponding low-dimensional dense term vectors, and each vector carries the specific semantics of the corresponding term in the scientific field.
[0041] Step S1133: performing context feature extraction processing on the term vector through the bidirectional recurrent neural network layer to capture the before-and-after semantic association information of the scientific term unit in the text sequence.
[0042] The term vector obtained through the embedding layer processing is input into the bidirectional recurrent neural network layer. The layer is composed of a forward recurrent neural network and a backward recurrent neural network. The forward recurrent neural network starts from the beginning of the text sequence and processes each term vector in turn to capture the context semantic association of the term in the sequence. The backward recurrent neural network starts from the end of the text sequence and processes each term vector in reverse to capture the context semantic association of the term in the sequence.
[0043] Taking the sequence “molecular mechanism of kidney function impact” as an example, the forward recurrent neural network can combine the information carried by the term vectors of “hypertension” and other terms in front of “kidney function” when processing the term vector corresponding to “kidney function”. The backward recurrent neural network can combine the information carried by the term vectors of “molecular mechanism” and other terms behind “kidney function” when processing the term vector corresponding to “kidney function”. Through the above bidirectional processing mode, the context semantic association information of each scientific term unit in the entire text sequence can be fully captured, and the extracted features are more comprehensive and accurate.
[0044] Step S1134: performing sequence labeling processing on the feature vector output by the bidirectional recurrent neural network layer through the conditional random field layer, and assigning an entity class label to each scientific term unit, wherein the entity class label includes a related gene / protein label, a cell type label, a molecular marker label, and an experimental method label.
[0045] The feature vector output by the bidirectional recurrent neural network layer is sent to the conditional random field layer, which assigns an entity class label to each scientific term unit according to the features of each word in the sequence and the dependency between the words.
[0046] In this example, "type 2 diabetes" and "hypertension" are labeled as related disease model labels; "kidney" belongs to the cell type related label, so the part of the entity associated with "kidney" in "kidney function" is labeled as a cell type label, and the "molecular mechanism of kidney function impact" as a whole is also labeled as a molecular mechanism related label. For word units such as "basic research", since they do not belong to the above four types of entity categories, they will not be assigned corresponding entity class labels.
[0047] Step S1135: Perform entity boundary merging processing on the labeled scientific term units, and merge consecutive scientific term units with the same entity class label into a complete entity name.
[0048] After labeling, the scientific term units need to be processed for entity boundary merging. For example, if there are consecutive word units that are all labeled as related gene / protein labels, such as "type 2" and "diabetes", although they are two units in the previous segmentation, they actually constitute a complete entity name "type 2 diabetes", so they will be merged into a complete entity name "type 2 diabetes".
[0049] Similarly, for similar cases that may occur, such as "hypertension" being split into "high" and "blood pressure" and both being labeled as related gene / protein labels, they will also be merged into the complete entity name "hypertension" to ensure the integrity of the entity name.
[0050] Step S1136: Remove duplicate entities in the merged entity names, keep unique entity names, and classify all unique entity names according to entity class labels to form an entity name set.
[0051] After the entity boundary merging processing, it can be checked whether there are duplicate entities in the obtained entity names. In this example, the entity names obtained after merging may include "type 2 diabetes", "hypertension", "kidney", "molecular mechanism of kidney function impact", etc. If there are no duplicate entities, these entity names are directly retained.
[0052] Subsequently, the effective text content is classified according to the entity category labels, and the entity names corresponding to the relevant gene / protein labels are "type 2 diabetes", "hypertension"; the entity name corresponding to the cell type label is "kidney"; the entity name corresponding to the molecular marker label is "molecular mechanism affecting kidney function"; and the entity name corresponding to the experimental method label in this example is null. The classified entity names are combined to form an entity name set.
[0053] Step S114: The effective text content is classified by a scientific research question type classification process, and based on the semantic relationship between the entity name set and the scientific research word unit, the basic scientific research question description is classified into mechanism exploration, function verification, biomarker screening, or phenotype research, and a question type identifier is generated.
[0054] After obtaining the entity name set, the effective text content is classified by a question type classification process in combination with the semantic relationship of the scientific research word unit. The basic scientific research question description in the example is "basic research on the molecular mechanism of type 2 diabetes and hypertension and its effect on kidney function". The "molecular mechanism" reflects the inquiry of mechanism exploration related content, and the "kidney function impact" clearly points to the content of function verification.
[0055] Through analysis of the key semantics in the text, it is determined that the question involves both mechanism exploration and function verification. However, according to the preset classification rules, when the question involves multiple types, the main focus is given priority, and in this example, the focus of the two types is balanced, so a composite question type identifier such as "mechanism exploration + function verification" can be generated to accurately reflect the type of the question.
[0056] Step S115: Calculate the text feature quantization parameter in the effective text content. The number of determiners, interrogative words, and entity names contained in the basic scientific research question description is calculated to obtain the text feature quantization parameter, which is positively correlated with the number of determiners and entity names.
[0057] Next, the text feature quantization parameter is calculated. First, the determiners in the basic scientific research question description are identified, such as "and", "its", etc. These words limit the research content; interrogative words are implied inquiries; and the entity name set contains "type 2 diabetes", "hypertension", "kidney", "molecular mechanism affecting kidney function", etc., with a large number.
[0058] According to the preset calculation mode, the calculation of the text feature quantization parameter considers the number of determiners, the use of interrogative words, and the number of entity names. Since the parameter is positively correlated with the number of determiners and the number of entity names, in this example, the number of determiners and entity names is relatively large, and therefore the calculated text feature quantization parameter is at a high level, reflecting the user's detailed information requirement for the question.
[0059] Step S116: Integrating the entity name set, the question type identifier, and the text feature quantization parameter to generate a question feature set containing field-specific semantic information.
[0060] Finally, the entity name set, the question type identifier "mechanism discussion + function verification", and the text feature quantization parameter obtained through the above steps are integrated. During the integration process, these information can be organized according to a certain structure to ensure that the association between them is clear and identifiable, forming a question feature set that can fully reflect the core features of the scientific research question.
[0061] Step S120: Based on the question feature set, call the biomedical literature database interface to perform basic scientific literature knowledge retrieval processing, and obtain a basic scientific literature knowledge set matched with the question feature set, which contains the core viewpoint abstract, basic experimental method description, and molecular mechanism description of multiple basic scientific literature.
[0062] After obtaining the question feature set, the processing module will call the biomedical literature database interface based on the set to start the retrieval operation of the basic scientific literature knowledge. By matching the key information in the question feature set with the literature in the database, relevant literature is filtered out, and the core content thereof is extracted to form a basic scientific literature knowledge set.
[0063] Step S121: Analyzing the entity name set in the question feature set, converting each entity name into a standard search term supported by the biomedical literature database to generate a list of standardized search terms.
[0064] The entity name set in the question feature set is analyzed to obtain entity names such as "type 2 diabetes", "hypertension", "kidney", and "molecular mechanism of kidney function impact". Then, each entity name is converted into a corresponding standard search term by referring to the standard vocabulary system supported by the biomedical literature database.
[0065] For example, the standard search term corresponding to "type 2 diabetes" in the database may be "type 2 diabetes" itself, the standard search term corresponding to "hypertension" is "hypertension", the standard search term corresponding to "kidney" is "kidney", and the standard search term corresponding to "molecular mechanism affecting kidney function" may be "diabetic kidney molecular mechanism", etc. These converted standard search terms are sorted together to form a list of standardized search terms.
[0066] Step S122: According to the problem type identification in the problem feature set, select the corresponding literature retrieval strategy from the preset retrieval strategy library, and the literature retrieval strategy includes retrieval field combination method, logical operator configuration and literature publication time range limit.
[0067] According to the problem type identification "mechanism exploration + function verification", the corresponding literature retrieval strategy is selected from the preset retrieval strategy library. For mechanism exploration type problems, the retrieval field combination method may focus on "abstract", "keyword", "introduction" and other fields to obtain content related to molecular mechanism; for function verification type problems, the retrieval field combination method may pay more attention to "method", "result" and other fields.
[0068] In terms of logical operator configuration, "AND" can be used to connect different search terms to ensure that the search results contain multiple key information at the same time, such as "type 2 diabetes" AND "hypertension" AND "kidney function" AND "molecular mechanism", etc. The literature publication time range limit may be set to the past ten years to ensure that the obtained literature has certain timeliness and reference value.
[0069] Step S123: Combine the standardized search term list with the literature retrieval strategy to generate a structured search expression, which includes multiple search terms and their logical relationships.
[0070] Combine the search terms "type 2 diabetes", "hypertension", "kidney", "molecular mechanism affecting kidney function" and others in the standardized search term list with the selected literature retrieval strategy to generate a structured search expression. For example, the expression may be "(type 2 diabetes AND hypertension AND kidney function) AND (molecular mechanism OR function verification)", which includes multiple search terms and their logical relationships such as "AND" and "OR" to accurately locate the required literature.
[0071] At the same time, combined with the retrieval field combination method, the corresponding retrieval field of each search term is determined, such as "type 2 diabetes [keyword] AND hypertension [keyword]", etc., so that the search expression is more accurate.
[0072] Step S124: Call the biomedical literature database interface, send the structured search expression to the database, trigger the literature search operation, obtain the preliminary search result set, which contains the literature titles, author information and abstract fragments of multiple basic research literatures.
[0073] The interface of the biomedical literature database is called, and the generated structured search expression is sent to the database. After the database receives the expression, it starts the literature search operation and performs matching search in the database according to the search terms, logical relationships and search fields in the expression.
[0074] After the search is completed, the preliminary search result set is returned, which contains the relevant information of multiple basic research literatures that meet the conditions, such as literature titles such as "Molecular Mechanism Research on the Influence of Type 2 Diabetes Combined with Hypertension on Kidney Cell Model Function" and "Progress of In Vitro Verification of Molecular Mechanism of Diabetic Kidney", as well as author information and abstract fragments of each literature.
[0075] Step S125: Perform relevance sorting processing on the preliminary search result set, based on the matching degree of the text feature quantization parameters and entity name set in the problem feature set with the literature content, sort the basic research literatures in the preliminary search result set, and select the top ranked preset number of literatures to form the target literature set.
[0076] After obtaining the preliminary search result set, the literatures in it need to be sorted for relevance to filter out the most relevant literatures to the question. This process needs to consider the matching degree of the text feature quantization parameters and the literature content with the entity name set to ensure the rationality of the sorting result.
[0077] Step S1251: Extract the literature title and abstract fragment of each basic research literature in the preliminary search result set to form the literature content text.
[0078] Extract the literature title and abstract fragment of each basic research literature from the preliminary search result set, and combine them to form the literature content text. For example, for the literature "Molecular Mechanism Research on the Influence of Type 2 Diabetes Combined with Hypertension on Kidney Cell Model Function", combine the title with the content in the abstract fragment to form the literature content text corresponding to this literature.
[0079] Step S1252: Match the literature content text with the entity name set in the problem feature set, calculate the number of entity names contained in each literature and the matching frequency, and generate an entity matching score.
[0080] The content text of each document is matched with entity names in the entity name set, such as "type 2 diabetes", "hypertension", "kidney", "molecular mechanism", and the like. The number of entity names contained in each content text is counted, and the frequency of each entity name is counted.
[0081] According to a preset calculation rule, the number and frequency are converted into entity matching scores. The more entity names contained and the higher the matching frequency, the higher the entity matching score. For example, if a document contains all four entity names and has a high frequency, the entity matching score of the document will be relatively high.
[0082] Step S1253: Perform semantic similarity calculation processing on the content text of the document, convert the content text of the document and the basic scientific research problem description into semantic vectors, and calculate the cosine similarity between the semantic vectors as a semantic correlation score.
[0083] The content text of the document and the basic scientific research problem description are subjected to semantic analysis, and are converted into semantic vectors respectively. These semantic vectors can represent the overall meaning of the text in the semantic space. Then, the cosine similarity between the two semantic vectors is calculated. The closer the value of the cosine similarity is to 1, the more similar the semantics of the two texts are.
[0084] The calculated cosine similarity is taken as the semantic correlation score. For example, if the content of a document is highly semantically related to the scientific research problem description, the semantic correlation score of the document will be high.
[0085] Step S1254: Based on the text feature quantization parameter in the problem feature set, the entity matching score and the semantic correlation score are weighted and summed to generate a document comprehensive correlation score, wherein the higher the text feature quantization parameter, the greater the weight of the entity matching score.
[0086] After obtaining the entity matching score and the semantic correlation score, the text feature quantization parameter in the problem feature set is combined for weighted summation. First, the entity matching score and the semantic correlation score are respectively set with a basic weight ratio. Assuming that when the text feature quantization parameter is at a medium level, the basic weight of the entity matching score is the same as the basic weight of the semantic correlation score.
[0087] Since the text feature quantization parameter in this example is high, the weight of the entity matching score needs to be increased according to the rule. For example, the ratio of the weight of the entity matching score and the weight of the semantic correlation score is one to one before adjustment, and after adjustment, the weight of the entity matching score may be increased and the weight of the semantic correlation score may be decreased.
[0088] Then, the entity matching score is multiplied by the adjusted entity weight, the semantic correlation score is multiplied by the adjusted semantic weight, and the two results are added to obtain the literature comprehensive correlation score of each literature. Through the above method, when the text feature quantization parameter is high, the influence of the entity matching degree on the literature comprehensive correlation is greater, which is more in line with the user's demand for accurate entity information.
[0089] Step S1255: The basic scientific literatures in the preliminary search result set are sorted in order of the literature comprehensive correlation score from high to low.
[0090] After calculating the literature comprehensive correlation score of each literature, all the basic scientific literatures in the preliminary search result set are arranged in order of the score from high to low. The literature with the highest score is ranked first, and the rest are arranged in sequence to form an ordered literature sequence.
[0091] During the sorting process, if the literature comprehensive correlation scores of two literatures are the same, the publication time of the literature can be further referred to, and the literature with a more recent publication time is arranged in front to preferentially display the updated research results.
[0092] Step S1256: According to the preset number of literatures, the top ranked literatures in the sorted basic scientific literatures are selected to form a target literature set.
[0093] After sorting, according to the preset number of literatures, the corresponding number of literatures ranked in the front are selected from the sorted basic scientific literature sequence. For example, if the preset number is twenty, the top twenty literatures are selected.
[0094] The selected literatures are integrated to form a target literature set. The literatures in the target literature set are the most relevant literatures to the basic scientific problem description proposed by the user.
[0095] Step S126: The core viewpoint abstract, basic experimental method description, and molecular mechanism description of each basic scientific literature are extracted from the target literature set and integrated into a basic scientific literature knowledge set.
[0096] For each basic scientific literature in the target literature set, the content is extracted one by one. The core viewpoint abstract part mainly extracts the summary description of the research core argument and main findings in the literature; the basic experimental method description part extracts the experimental design, cell model construction, data acquisition and analysis method, etc. introduced in the literature; and the molecular mechanism description extracts the specific mechanism, function results, etc. fragment information based on the research results in the literature.
[0097] The core idea summary, basic experimental method description, and molecular mechanism description of each extracted literature are collected and preliminarily sorted according to the order of the literature or the theme to form a basic scientific research literature knowledge collection. For example, for the research on the molecular mechanism of type 2 diabetes combined with hypertension affecting the function of kidney cell models, the core ideas, cell model experimental methods, and related mechanism conclusions about the molecular mechanism are extracted to form a basic scientific research literature knowledge collection together with the corresponding contents of other literatures.
[0098] Step S130: Knowledge unit extraction processing is performed on the basic scientific research literature knowledge collection to extract scientific research knowledge units with independent semantics from the core idea summary, basic experimental method description, and molecular mechanism description, and a structured knowledge unit collection is generated.
[0099] After obtaining the basic scientific research literature knowledge collection, scientific research knowledge units with independent semantics need to be extracted from it. These knowledge units are the basic units that constitute the subsequent answer content. Through careful processing of the literature knowledge collection, complex literature content is decomposed into independent and meaningful knowledge units, which are then structured and sorted.
[0100] Step S131: The core idea summary, basic experimental method description, and molecular mechanism description in the basic scientific research literature knowledge collection are traversed, and the corresponding text content is divided into multiple text paragraph units.
[0101] Each item of content in the basic scientific research literature knowledge collection, including the core idea summary, basic experimental method description, and molecular mechanism description, is traversed. For each part of the text content, it is divided into multiple text paragraph units according to natural paragraph division or according to the logical hierarchy of the content.
[0102] For example, the core idea summary of a literature may contain three natural paragraphs, each of which elaborates the core idea from a different angle, so it is divided into three text paragraph units. In the basic experimental method description, the content about cell model construction and experimental operation may each form a paragraph, and is accordingly divided into two text paragraph units. By dividing long texts into relatively independent paragraph units in the above manner, theme recognition and knowledge unit extraction are facilitated in the subsequent steps.
[0103] Step S132: Each text paragraph unit is subjected to semantic theme recognition processing, and the core theme of each text paragraph unit is determined by a pre-trained theme model to generate a theme label.
[0104] Each text paragraph unit is input into a pre-trained theme model, which has been trained on a large amount of basic scientific research literature corpus and can identify the core theme of the text paragraph.
[0105] The topic model analyzes the text paragraph unit, extracts key terms, high-frequency words and semantic associations, and determines the core topic of the paragraph. For example, a text paragraph unit mainly discusses the oxidative stress molecular mechanism of type 2 diabetes and hypertension on the function of kidney cell models, and the topic model identifies this core topic and generates a corresponding topic label, such as "type 2 diabetes hypertension kidney cell model oxidative stress molecular mechanism".
[0106] Each text paragraph unit corresponds to a generated topic label, which can succinctly reflect the core content of the paragraph.
[0107] Step S133: Based on the topic label, the text paragraph unit is clustered and classified into the same topic cluster, and each topic cluster corresponds to a basic scientific research knowledge topic.
[0108] Collect the topic labels of all text paragraph units, compare and analyze the similarity of the topic labels. The text paragraph units with the same topic label are directly classified into the same class; for similar topic labels, such as "oxidative stress mechanism of type 2 diabetes and hypertension on kidney function" and "oxidative stress path of type 2 diabetes combined with hypertension on kidney function", after similarity judgment, they are also classified into the same class.
[0109] Each class of text paragraph units forms a topic cluster, and each topic cluster corresponds to a clear basic scientific research knowledge topic. For example, the two similar topic labels correspond to the text paragraph units that form a topic cluster, and the corresponding basic scientific research knowledge topic is the oxidative stress related molecular mechanism of type 2 diabetes combined with hypertension on kidney function.
[0110] Through clustering, the scattered text paragraph units are collected according to the topic, which facilitates the extraction and integration of knowledge units for the same topic in the subsequent process.
[0111] Step S134: Perform knowledge unit boundary recognition processing on the text paragraph units in each topic cluster, and according to the semantic rules and punctuation symbol characteristics of the basic scientific research field, divide the text paragraph units into multiple basic scientific research knowledge units with independent semantics, including statement type knowledge units, method type knowledge units and conclusion type knowledge units.
[0112] Step S1341: Perform punctuation symbol recognition processing on the text paragraph units in each topic cluster, and mark the punctuation symbol positions in the text paragraph units, including period symbol positions, semicolon symbol positions, colon symbol positions and dash symbol positions.
[0113] Scan the content of each text paragraph unit one by one, identify and mark the specific positions of punctuation marks such as periods, semicolons, colons, and dashes in them. For example, in a certain text paragraph unit, "In a model of type 2 diabetes, hyperglycemic state can accelerate kidney cell damage. Studies have shown that hypertension model can further aggravate the above damage; when the two work together, the related phenotype of abnormal kidney function is significantly increased." Here, the position of the period after "damage" and the semicolon after "damage" are marked.
[0114] Step S1342: Based on the semantic rules of the basic scientific research field, identify the conjunctions representing causal relationships, transitional relationships, and progressive relationships in the text paragraph unit, and determine the positions of the conjunctions in the text paragraph unit.
[0115] Referring to the semantic rules of the basic scientific research field, identify the conjunctions representing different logical relationships in the text paragraph unit. Conjunctions representing causal relationships include "because", "therefore", "resulting in", etc.; conjunctions representing transitional relationships include "however", "but", "on the contrary", etc.; conjunctions representing progressive relationships include "in addition", "further", "and", etc.
[0116] Locate the positions of these conjunctions in the text paragraph unit, for example, "Because there is insulin resistance in the model of type 2 diabetes, the kidney's ability to regulate hemodynamics is reduced, and the hypertension model exacerbates the situation." The positions of "because", "therefore", and "and" are determined. The positions of the conjunctions help determine the logical separation of the semantics and assist in determining the boundaries of the knowledge units.
[0117] Step S1343: Combine the punctuation mark positions and conjunction positions to preliminarily segment the text paragraph unit into multiple semantic segments.
[0118] Consider the position information of punctuation marks and conjunctions to preliminarily segment the text paragraph unit. For example, in a text containing periods and conjunctions, segmentation is usually performed at the period, and the new semantic part guided by the conjunction may also be considered as the beginning of a new semantic segment.
[0119] For example, "Type 2 diabetes combined with hypertension increases the burden on the kidneys. However, early intervention can effectively delay the progression of related pathologies." Combine the positions of the period and "however" to segment it into two semantic segments: "Type 2 diabetes combined with hypertension increases the burden on the kidneys" and "Early intervention can effectively delay the progression of related pathologies."
[0120] Through the above preliminary segmentation, the text paragraph unit is decomposed into relatively independent semantic segments, laying the foundation for further processing.
[0121] Step S1344: Scientific terminology integrity check is performed on each semantic segment to ensure that the segmented semantic segment contains complete relevant gene / protein name and molecular concept.
[0122] Each semantic segment obtained by preliminary segmentation is checked to ensure that the relevant gene / protein name and molecular mechanism concept contained therein are complete. For example, if a semantic segment is segmented as "renal cells in a model of type 2 diabetes", where "model of type 2 diabetes" and "renal cells" are both complete basic scientific objects, the semantic segment passes the check; if a semantic segment is "hypertension combined with type 2", where "type 2" is incomplete, lacking the key part "model of diabetes", it needs to be adjusted and combined with the subsequent related semantic part to form a complete semantic segment such as "hypertension combined with model of type 2 diabetes".
[0123] Through the scientific terminology integrity check, it is ensured that each semantic segment has complete scientific semantics, avoiding semantic loss due to improper segmentation.
[0124] Step S1345: Secondary segmentation is performed on the semantic segment containing multiple short sentences, and the short sentences are segmented into independent basic scientific knowledge units according to the logical relationship between the short sentences.
[0125] For the semantic segment that has passed the integrity check, if it contains multiple short sentences and these short sentences have relatively independent semantics, secondary segmentation is needed. For example, "urine microalbumin detection is a commonly used indicator, and this detection has sensitivity in the early manifestations of renal function impairment", the semantic segment contains two short sentences, which respectively introduce the detection indicator and its function, and are relatively independent in logic, so it is secondary segmented into two independent parts "urine microalbumin detection is a commonly used indicator" and "this detection has sensitivity in the early manifestations of renal function impairment".
[0126] Secondary segmentation is performed according to the logical relationship such as parallelism, explanation, and supplement between short sentences, to ensure that each part after segmentation can become an independent basic scientific knowledge unit.
[0127] Step S1346: A type label is added to each basic scientific knowledge unit, and the knowledge unit is marked as a statement type knowledge unit, a method type knowledge unit, or a conclusion type knowledge unit according to the content of the knowledge unit.
[0128] According to the specific content of the knowledge unit, a corresponding type label is added to each unit. The statement type knowledge unit is mainly the description of scientific research facts, phenomena, molecular mechanisms, etc., such as "type 2 diabetes combined with hypertension can cause changes in renal hemodynamics"; the method type knowledge unit involves experimental methods, detection means, research processes, etc., such as "enzyme-linked immunosorbent assay is used to detect the expression level of urinary microalbumin"; the conclusion type knowledge unit is the conclusion or inference obtained from the experiment, such as "early control of blood pressure can slow the occurrence of renal dysfunction".
[0129] By adding type labels, the nature of each basic scientific knowledge unit is clear, which facilitates subsequent weight allocation and integration processing.
[0130] Step S135: Add source literature identification and theme cluster identification to each basic scientific knowledge unit, and establish the association between knowledge units and original basic scientific literature and theme clusters.
[0131] A source literature identification is added to each extracted basic scientific knowledge unit. The identification can be a unique identifier of the original basic scientific literature, such as the DOI number or internal database number of the literature, etc. Through the identification, it can be traced back to which literature the knowledge unit comes from.
[0132] At the same time, a theme cluster identification is added to each knowledge unit, indicating which theme cluster the knowledge unit belongs to. For example, a knowledge unit belongs to the "type 2 diabetes hypertension renal function influence oxidative stress mechanism" theme cluster, and the corresponding theme cluster identification is added.
[0133] By adding these two identifications, the clear association between knowledge units and original basic scientific literature and theme clusters is established, which facilitates subsequent review, verification and integration.
[0134] Step S136: Group all basic scientific knowledge units according to theme cluster identification, and generate a structured knowledge unit set containing theme classification information and source association information.
[0135] Group all basic scientific knowledge units that have added type labels, source literature identification and theme cluster identification according to theme cluster identification. Knowledge units belonging to the same theme cluster identification are grouped together to form multiple theme knowledge unit groups.
[0136] Within each theme knowledge unit group, the knowledge units can be arranged according to their type or extraction order. These theme knowledge unit groups are integrated together to form a structured knowledge unit set, which not only contains the content of each knowledge unit, but also contains theme classification information (reflected by theme cluster identification) and source association information (reflected by source literature identification), making the organization of knowledge units more orderly and standardized.
[0137] Step S140: The structured knowledge unit set is associated and matched with the scientific research problem feature set. The knowledge units in the structured knowledge unit set are weighted and distributed according to the scientific research problem type identifier and the information demand intensity parameter, and a weighted knowledge integration result is generated.
[0138] The structured knowledge unit set and the scientific research problem feature set are two key elements for association and matching. By analyzing the degree of association between the two, combining the problem type and the information demand intensity, a corresponding weight is assigned to each knowledge unit, and finally a weighted knowledge integration result is formed, so as to generate a targeted answer text subsequently.
[0139] Step S141: Extract the text content of each knowledge unit in the structured knowledge unit set, match the text content of the knowledge unit with the related gene / protein name set in the scientific research problem feature set, and calculate the number of related gene / protein names contained in each knowledge unit and the matching degree.
[0140] The text content of each knowledge unit is extracted from the structured knowledge unit set, and then the text content is matched with the related gene / protein name set in the problem feature set. Each knowledge unit is checked one by one to see which related gene / protein name is contained in the text content, and the number of related gene / protein names is counted.
[0141] At the same time, the matching degree is calculated, and the calculation of the matching degree considers the frequency and importance of the related gene / protein name in the knowledge unit. For example, a knowledge unit contains three related scientific research objects "type 2 diabetes model", "hypertension model", and "kidney function abnormality", and the frequency is high, so the matching degree is relatively high; while a knowledge unit contains only one related scientific research object and the frequency is low, the matching degree is low.
[0142] Through the above matching processing, the close degree of each knowledge unit to the core scientific research entity of the problem can be known.
[0143] Step S142: According to the problem type identifier in the scientific research problem feature set, determine the knowledge unit type weight corresponding to the problem type identifier, and assign a basic weight value to knowledge units of different types.
[0144] Referring to the problem type identifier "disease mechanism exploration + function verification class", the basic weight values of different types of knowledge units are determined. For the disease mechanism exploration part, the statement type knowledge unit related to the molecular mechanism may have a higher basic weight; for the function verification part, the basic weight of the method type knowledge unit may be relatively high.
[0145] For example, the basic weight value of the statement-type knowledge unit (related to the mechanism), the basic weight value of the method-type knowledge unit (related to the functional verification), and the basic weight value of the conclusion-type knowledge unit are set, wherein the basic weight values of the former two are higher than that of the conclusion-type knowledge unit, so as to highlight the importance of the knowledge units related to the problem type.
[0146] Step S143: Adjust the basic weight value based on the information demand intensity parameter, and normalize the weight value of each knowledge unit.
[0147] According to the adjustment of the basic weight value based on the information demand intensity parameter, since the information demand intensity parameter is high in this example, the weight value of the knowledge unit with high matching degree of the related gene / protein name can be appropriately increased. For example, for the knowledge unit with high matching degree, an adjustment coefficient greater than one is multiplied on the basis weight value thereof; for the knowledge unit with low matching degree, the adjustment coefficient can be less than one.
[0148] After the adjustment, the weight values of all the knowledge units are normalized. The purpose of the normalization is to make the weight values of all the knowledge units in the same numerical range, so as to facilitate the comparison and subsequent sorting processing. The specific method is to divide the weight value of each knowledge unit by the sum of the weight values of all the knowledge units, so that the sum of all the normalized weight values is one.
[0149] Step S144: Sort the knowledge units in the structured knowledge unit set according to the normalized weight value, group and integrate the sorted knowledge units according to the theme cluster identifier, and generate a weighted knowledge integration result containing weight information and theme classification.
[0150] For example, step S1441: sort the knowledge units in the structured knowledge unit set according to the normalized weight value from high to low, and generate a knowledge unit sorting sequence.
[0151] All the knowledge units are traversed, and sorted according to the size of the normalized weight value. The knowledge unit with the highest normalized weight value is ranked first, and the same applies to the others, so as to form a knowledge unit sorting sequence.
[0152] In the sorting process, if the normalized weight values of two knowledge units are the same, the order of them can be determined by referring to the matching degree or the degree of relevance to the theme.
[0153] Step S1442: Traverse the knowledge unit sorting sequence, and extract the theme cluster identifier of each knowledge unit.
[0154] Each knowledge unit in the knowledge unit ordering sequence is checked one by one, and the respective topic cluster identifier thereof is extracted. For example, the topic cluster identifier of the first knowledge unit is "2 type diabetes hypertension kidney function influence oxidative stress mechanism", and the topic cluster identifier of the second knowledge unit is "abnormal kidney function related molecular marker", and so on.
[0155] The topic cluster identifiers are associated with the corresponding knowledge units, in preparation for subsequent grouping and integration.
[0156] Step S1443: Knowledge units with the same topic cluster identifier are classified into the same topic group, forming a plurality of topic knowledge groups.
[0157] According to the extracted topic cluster identifier, the knowledge units in the knowledge unit ordering sequence with the same topic cluster identifier are classified into a category to form a topic knowledge group. For example, all knowledge units with the topic cluster identifier "2 type diabetes hypertension kidney function influence oxidative stress mechanism" are integrated into a topic knowledge group, and all knowledge units with the topic cluster identifier "abnormal kidney function related molecular marker" are integrated into another topic knowledge group.
[0158] Through the above classification, content-related knowledge units are concentrated together, facilitating knowledge presentation by topic.
[0159] Step S1444: After the knowledge units in each topic knowledge group are sorted according to the normalized weight value, a topic description information is added to each topic knowledge group, which is generated based on the core topic corresponding to the topic cluster identifier.
[0160] After the knowledge units are classified into various topic knowledge groups, the knowledge units in each topic knowledge group need to be sorted again. The basis for sorting is the normalized weight value obtained earlier, and the knowledge units in each topic knowledge group are arranged in order of normalized weight value from high to low. The purpose of this is to let the knowledge units inside each topic knowledge group also present according to the importance, so as to facilitate the subsequent generation of answer text, and preferentially display more critical content.
[0161] After the secondary sorting, a topic description information is added to each topic knowledge group. The generation of the topic description information is based on the core topic pointed by the topic cluster identifier corresponding to the topic knowledge group. For example, for the topic knowledge group with the topic cluster identifier of "oxidative stress mechanism of the influence of type 2 diabetes combined with hypertension on kidney function", the core topic is the oxidative stress related molecular mechanism of the influence of type 2 diabetes combined with hypertension on kidney function, and the generated topic description information can be "analysis of the oxidative stress related molecular mechanism of the influence of type 2 diabetes combined with hypertension on kidney function". For the topic knowledge group with the topic cluster identifier of "molecular markers related to abnormal kidney function", the topic description information can be "screening of molecular markers related to abnormal kidney function of type 2 diabetes combined with hypertension". These topic description information can briefly summarize the core content of each topic knowledge group, so that the user can quickly know the main information involved in the topic knowledge group.
[0162] Step S1445: Sort all the topic knowledge groups according to the average normalized weight values of the knowledge units contained in the topic knowledge groups, and generate a topic knowledge group sequence.
[0163] After adding the topic description information to each topic knowledge group, the overall sorting of all the topic knowledge groups is needed. The basis for sorting is the average normalized weight value of the knowledge units contained in each topic knowledge group. The way to calculate the average normalized weight value is to add the normalized weight values of all the knowledge units in the topic knowledge group, and then divide the result by the number of knowledge units in the topic knowledge group, and the result is the average normalized weight value of the topic knowledge group.
[0164] For example, a topic knowledge group contains three knowledge units, and their normalized weight values are a, b and c respectively. The average normalized weight value of the topic knowledge group is (a+b+c) divided by 3. After calculating the average normalized weight values of all the topic knowledge groups in the above manner, the topic knowledge groups are arranged in order from high to low according to the average normalized weight values, and a topic knowledge group sequence is formed. Through the above sorting, the entire weighted knowledge integration result can be presented in order of the importance of the topic, and the important topic knowledge groups are arranged in the front, which is convenient for subsequent generation of answer text.
[0165] Step S1446: Integrate each topic knowledge group in the topic knowledge group sequence, the knowledge units contained therein, the normalized weight values and the topic description information, and generate a weighted knowledge integration result.
[0166] After obtaining the sequence of topic knowledge groups, the sequence can be integrated to generate a weighted knowledge integration result. The integration process is to collect all the information contained in each topic knowledge group in the sequence, including the topic knowledge group itself, the knowledge units in the topic knowledge group after secondary sorting, the normalized weight value corresponding to each knowledge unit, and the topic description information of the topic knowledge group.
[0167] During integration, the relevant information of each topic knowledge group is sequentially incorporated into the weighted knowledge integration result according to the order of the topic knowledge groups in the sequence. For example, the topic description information, knowledge units and their normalized weight values of the topic knowledge group ranked first in the sequence are first integrated, followed by the relevant information of the topic knowledge group ranked second, and so on, until all the information of the topic knowledge groups is integrated. The weighted knowledge integration result generated in this way not only contains the core content of each topic, but also reflects the importance of different topics and different knowledge units within the same topic.
[0168] Step S150: generating a scientific research question answering text based on the weighted knowledge integration result, and feeding back the scientific research question answering text to the user interaction interface of the scientific research academic community for display.
[0169] After obtaining the weighted knowledge integration result, a scientific research question answering text can be generated based on the result for the user's scientific research question, and displayed on the user interaction interface, allowing the user to intuitively obtain the required information. This step requires parsing and processing the weighted knowledge integration result, organizing the knowledge units in the result into natural and fluent text according to a reasonable structure, and ensuring that the user can easily view and trace the relevant information.
[0170] Step S151: parsing the sequence of topic knowledge groups in the weighted knowledge integration result to determine the topic description information of each topic knowledge group and the basic scientific research knowledge units contained therein.
[0171] First, the sequence of topic knowledge groups in the weighted knowledge integration result is parsed. By traversing the sequence of topic knowledge groups, the topic description information of each topic knowledge group is extracted one by one, and the basic scientific research knowledge units contained in each topic knowledge group are also determined. For example, during the parsing process, it can be determined that the topic description information of the first topic knowledge group is "oxidative stress-related molecular mechanism analysis of the effect of type 2 diabetes mellitus combined with hypertension on kidney function", and all the basic scientific research knowledge units related to this mechanism contained in the topic knowledge group; the topic description information of the second topic knowledge group is "molecular marker screening for kidney function abnormalities related to type 2 diabetes mellitus combined with hypertension", and the corresponding basic scientific research knowledge units, etc. Through the above parsing, the core content and specific components of each topic knowledge group can be effectively obtained.
[0172] Step S152: generating a corresponding title paragraph for each topic knowledge group based on the sorting order of the topic knowledge group sequence, the title paragraph containing topic description information and an importance description of the topic knowledge group in the answer.
[0173] According to the sorting order of the topic knowledge group sequence, that is, in the order of the average normalized weight value of the topic knowledge group from high to low, a corresponding title paragraph is generated for each topic knowledge group. The content of the title paragraph mainly includes two parts, one part is the topic description information of the topic knowledge group, and the other part is the importance description of the topic knowledge group in the entire answer.
[0174] The importance description can be determined according to the proportion of the average normalized weight value of the topic knowledge group in all topic knowledge groups. For example, for the topic knowledge group ranked first in the sequence, the importance description can be "this part of the content is the core molecular mechanism to understand the problem, and has important significance for in-depth understanding of the related pathological process"; and for the topic knowledge group ranked later, the importance description can be "this part of the content provides reference information for related molecular marker screening, and has certain scientific research value", etc. By generating the above title paragraph, the user can have a general understanding of the core content and importance of each topic when reading the answer text.
[0175] Step S153: performing text integration processing on the basic scientific knowledge units in each topic knowledge group, and adding a reference mark to the integrated text content, the reference mark corresponding to the source literature mark of the knowledge unit, to facilitate the user to trace the original literature.
[0176] After generating the title paragraph, the basic scientific knowledge units in each topic knowledge group need to be subjected to text integration processing. Text integration processing is to reasonably organize and link the text content of the knowledge units in each topic knowledge group sorted according to the normalized weight value, forming a coherent text. In the integration process, attention should be paid to maintaining the logical relationship between the knowledge units to ensure that the integrated text content is smooth and easy to understand, and can accurately convey the information contained in the knowledge units.
[0177] Meanwhile, in order to facilitate users to trace the original literature, it is necessary to add citation marks to the integrated text content. The citation marks correspond to the source literature marks of the knowledge units. When each knowledge unit is integrated into the text content, the corresponding citation mark will be marked near the relevant content. For example, when a knowledge unit comes from a literature marked as "Lit-001", the above citation mark "[Lit-001]" can be added after the text content integrating the knowledge unit. In this way, during the reading process, if the user is interested in a part of the content or needs further verification, the corresponding original literature can be quickly found through the citation mark.
[0178] Step S154: Combine the title paragraphs of all topic knowledge groups and the integrated text content in order to generate a complete scientific research question answering text.
[0179] After completing the generation of the title paragraph of each topic knowledge group and the integration of the text content, combine all the title paragraphs and the integrated text content in the order of the sequence of the topic knowledge groups. That is, first place the title paragraph of the first topic knowledge group, followed by the integrated text content of the topic knowledge group; then the title paragraph of the second topic knowledge group and its integrated text content, and so on, until all the contents of the topic knowledge groups are combined.
[0180] During the combination process, attention should be paid to the connection between the parts to ensure that the structure of the entire answering text is clear and logically coherent. For example, transitional sentences can be appropriately added between the contents of two topic knowledge groups to make the text more natural and smooth. Through the above combination method, a complete answering text that can fully answer the user's scientific research question is finally generated.
[0181] Step S155: Send the scientific research question answering text to the user interaction interface of the scientific research academic community, display it below the user-submitted scientific research question text, and provide a link entry to view the original literature.
[0182] After generating the complete scientific research question answering text, send the text to the user interaction interface of the scientific research academic community. In the user interaction interface, the answering text will be displayed below the scientific research question text submitted by the user, facilitating the user to read the question and the answer in contrast.
[0183] At the same time, in order to further facilitate users to consult original literature, a link entry for viewing original literature will also be provided on the interactive interface. The link entry is associated with the source literature identifier corresponding to each knowledge unit. The user clicks the link entry corresponding to a reference identifier to jump to the acquisition page of the original literature, thereby obtaining more detailed literature content. Through the above display mode, not only does the user get a clear problem solution, but also provides the user with a way to further study, thereby improving the user experience.
[0184] In order to ensure the accuracy and efficiency of biomedical entity recognition, a biomedical entity recognition model needs to be pre-trained, which includes an embedding layer, a bidirectional recurrent neural network layer, and a conditional random field layer.
[0185] Step S211: Collecting text corpus in the biomedical field, which includes basic scientific research literature, scientific research teaching materials, and public scientific research literature resources, etc.
[0186] Firstly, a large amount of biomedical field text corpus is collected, which is widely sourced and covers basic scientific research literature, scientific research teaching materials, and public scientific research literature resources, etc. Basic scientific research literature can be obtained from major biomedical literature databases, including literature of different research directions and different publication times; scientific research teaching materials cover teaching materials content of basic life sciences, molecular biology, and other disciplines; public scientific research literature resources include desensitized experimental data and related scientific research reports. The collected text corpus needs to have a certain scale and diversity to ensure the effect of model training.
[0187] Step S212: Preprocessing the collected text corpus, including removing noise data, unifying text format, and performing word segmentation processing, etc.
[0188] After obtaining the text corpus, it needs to be preprocessed. Removing noise data means that information unrelated to biomedical content in the corpus, such as advertisements and irrelevant annotations, is removed to ensure the purity of the corpus. Unifying the text format is to convert text corpus of different sources and formats into a unified format for subsequent processing and training. Word segmentation processing is to split the text corpus into independent word units. The word segmentation tool used can be designed for the biomedical field to ensure the accuracy of word segmentation, especially for professional term segmentation.
[0189] Step S213: Labeling entity class labels for biomedical entities in the preprocessed text corpus, including cell type labels, related gene / protein name labels, molecular marker labels, and functional mechanism labels.
[0190] After the pre-processing, the biomedical entities in the text corpus need to be labeled with entity class labels. The labeling work can be done by professional life science researchers or with the help of existing labeling tools combined with manual review. During labeling, according to the entity type, the corresponding label is added, such as labeling "cardiomyocyte" as a cell type label, "TP53" as a related gene / protein name label, "inflammatory marker" as a molecular marker label, "apoptosis" as a functional mechanism label, etc. The labeled text corpus will be used as training data for model training.
[0191] Step S214: Construct a network structure of the biomedical entity recognition model, which includes an embedding layer, a bidirectional recurrent neural network layer, and a conditional random field layer in sequence.
[0192] According to the model design requirements, the corresponding network structure is constructed. The role of the embedding layer is to convert the word unit into a low-dimensional dense word vector; the bidirectional recurrent neural network layer is composed of forward and backward recurrent neural networks, which is used to extract the context features of the words; the conditional random field layer is used to label the sequence and output the entity class label. When constructing the network structure, the parameter settings of each layer need to be determined, such as the vector dimension of the embedding layer, the number of hidden layers and the number of neurons of the bidirectional recurrent neural network layer, etc.
[0193] Step S215: Divide the labeled text corpus into a training set, a validation set, and a test set, wherein the training set is used for model training, the validation set is used for adjusting model parameters, and the test set is used for evaluating model performance.
[0194] The labeled text corpus is divided into a training set, a validation set, and a test set according to a certain proportion. For example, according to the 7:2:1 proportion, 70% of the corpus is used as the training set for model training; 20% of the corpus is used as the validation set for evaluating model performance and adjusting parameters such as learning rate and iteration number during the training process; 10% of the corpus is used as the test set for final performance evaluation after training to determine whether the model achieves the expected recognition effect.
[0195] Step S216: Train the constructed biomedical entity recognition model using the training set, and monitor the model performance in real time through the validation set during the training process to adjust the model parameters.
[0196] The training set is input into the constructed biomedical entity recognition model to start the training process. The model continuously adjusts the network parameters to minimize the error between the predicted results and the actual labels according to the input text corpus and the corresponding entity class labels. Every certain number of training steps, the model performance is evaluated using the validation set to calculate accuracy, recall rate, etc. According to the indicators, the parameters are adjusted, such as reducing the learning rate or stopping training in advance when the performance of the validation set no longer improves, to avoid overfitting.
[0197] Step S217: After training is completed, the biomedical entity recognition model is evaluated for performance using a test set. If the evaluation results meet the preset requirements, the model is saved. If not, the model structure or training parameters are adjusted, and training is performed again.
[0198] After training is completed, the model performance is comprehensively evaluated using a test set. Indicators include accuracy, recall rate, F1 value, etc. If the results meet the preset requirements, the model achieves the expected performance and can be saved for subsequent biomedical entity recognition tasks. If not, the reasons are analyzed, which may be due to unreasonable model structure or improper training parameter settings. The model structure (such as the number of network layers and units) or training parameters (such as learning rate and training rounds) are adjusted according to the specific reasons, and retraining and evaluation are performed until the performance meets the standards.
[0199] In order to quickly select appropriate literature retrieval strategies according to different types of basic scientific research problems, a retrieval strategy library needs to be constructed in advance.
[0200] For example, first analyze common biomedical research problem types and divide them into molecular mechanism exploration, function verification, biomarker screening, and phenotype research.
[0201] First, collect and analyze common basic scientific research problems in the field of scientific research. According to the core content of the problem, it is divided into molecular mechanism exploration, function verification, biomarker screening, and phenotype research. Molecular mechanism exploration problems mainly focus on the mechanism of related genes, proteins, and signal pathways; function verification problems focus on experimental methods and functional research techniques; biomarker screening problems involve the identification and screening strategies of molecular markers; phenotype research problems focus on the phenotype changes of cell models or animal models, etc.
[0202] For each type of problem, design a corresponding literature retrieval strategy, which includes retrieval field combination method, logical operator configuration, and literature publication time range restriction.
[0203] For each type of problem, design a corresponding literature retrieval strategy. For example, for molecular mechanism exploration problems, the retrieval field combination can choose "abstract", "keyword", "introduction", "discussion", etc. These fields usually involve mechanism description; the logical operator is mainly "AND", connecting related genes, proteins, and mechanism words; the literature publication time range can be appropriately relaxed to obtain comprehensive research results.
[0204] For function verification type of problems, the search fields focus on "method", "result", "discussion" and other fields, and the experimental method and verification process are described in detail. In addition to "AND", "OR" can be used to connect different function verification terms. The literature time range can be set to a recent time period to obtain the latest experimental techniques and verification methods.
[0205] For biomarker screening type of problems, the search fields can be selected from "screening", "biomarker", "result" and other fields. The logic operator is mainly "AND", which connects the target molecules and screening terms. The literature time range should also be selected to be recent to reflect the current screening progress.
[0206] For phenotype research type of problems, the search fields include "phenotype", "function", "result" and other fields. The logic operator uses "AND" to connect the cell model or animal model and the phenotype-related terms. The literature time range can be set according to actual needs, which can include recent research or long-term observation results.
[0207] The designed search strategies are classified and stored according to the problem type to build a search strategy library.
[0208] The designed search strategies for each problem type are sorted and classified, and stored in the molecular mechanism exploration type, function verification type, biomarker screening type, and phenotype research type. During storage, each search strategy is added with the corresponding identifier to associate it with the corresponding problem type, making it easy to quickly call the corresponding search strategy according to the problem type in the subsequent search process. At the same time, a management mechanism for the search strategy library is established to update and optimize the search strategies according to actual application conditions.
[0209] In addition, the search strategies in the search strategy library are regularly evaluated and optimized, and the search field combination method, logic operator configuration, and literature publication time range limit are adjusted according to the search effect feedback.
[0210] After the search strategy library is put into use, the search strategies in it need to be evaluated regularly. The basis for evaluation is the search effect feedback, including the relevance of search results, recall rate, accuracy rate and other indicators. If the search strategy corresponding to a certain problem type has low relevance of search results or low recall rate in multiple uses, the search strategy needs to be optimized.
[0211] The optimization methods include adjusting the combination of search fields, increasing or decreasing search fields, modifying the configuration of logical operators to optimize the logical relationship between search terms, and adjusting the time range limit of document publication to make it more in line with the actual search needs. For example, if the search strategy for function verification type problems finds that there is insufficient relevant literature in recent years, the time range for document publication may be set too narrowly, in which case the time range can be appropriately expanded. If the search results contain a large number of irrelevant documents, it may be that the combination of search fields is unreasonable, in which case some fields that are prone to introducing irrelevant information can be reduced, or the use of logical operators can be adjusted to make the search conditions more precise.
[0212] The optimized search strategy is updated to the search strategy library, replacing the original search strategy. At the same time, the content and reasons for each evaluation and optimization are recorded to form a search strategy optimization log, providing a reference for further optimization. Through regular evaluation and optimization, the search strategies in the search strategy library can always maintain good applicability, improving the efficiency and accuracy of literature search.
[0213] In the above embodiments, when processing user input scientific research field text and related basic scientific research data, data privacy protection measures need to be taken to prevent sensitive information from being leaked. For example, sensitive information in scientific research data can be identified, including experimental sample numbers, laboratory internal codes, etc. During the collection and processing of scientific research data, sensitive information is first identified. Experimental sample numbers are unique codes used to identify cell lines or animal models; laboratory internal codes include experimental equipment numbers, research project numbers, etc. Once such information is leaked, it may affect research safety, so it needs to be identified and protected.
[0214] For identified sensitive information, appropriate desensitization measures are taken. Replacement is to replace sensitive information with meaningless symbols or codes, such as replacing sample numbers with "XXX" and experimental numbers with "************". Deletion means directly removing sensitive information from the data, such as deleting relevant identifiers when processing public data sets; encryption is to process sensitive information through encryption algorithms so that it cannot be interpreted without authorization, and only authorized personnel with decryption keys can view the original information.
[0215] Figure 2 A natural language processing combined outline content generation system 100 is shown in the embodiments of the present application, which includes a processor 1001, a memory 1003 and program code stored in the memory 1003, and the processor 1001 executes the above-mentioned program code to realize the steps of the natural language processing combined outline content generation method.
[0216] Figure 2The shown outline content generation system 100 combined with natural language processing comprises a processor 1001 and a memory 1003. The processor 1001 and the memory 1003 are connected, for example, through a bus 1002. Optionally, the outline content generation system 100 combined with natural language processing can further comprise a transceiver 1004, which can be used for data interaction, such as data sending and / or data receiving, between the outline content generation system combined with natural language processing and other outline content generation systems combined with natural language processing. It should be noted that the transceiver 1004 is not limited to one in actual scheduling, and the structure of the outline content generation system 100 combined with natural language processing does not constitute a limitation on the embodiments of the present application.
[0217] The memory 1003 is used to store program codes for implementing the embodiments of the present application, and is controlled by the processor 1001 to execute. The processor 1001 is used to execute the program codes stored in the memory 1003 to realize the steps shown in the foregoing method embodiments.
[0218] The embodiments of the present application provide a computer readable storage medium, which stores program codes. When the program codes are executed by a processor, the steps of the foregoing method embodiments and the corresponding contents can be realized.
[0219] It should be understood that, although the flowcharts of the embodiments of the present application indicate the respective operation steps by arrows, the implementation order of the steps is not limited to the order indicated by the arrows. Unless otherwise specified herein, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders based on requirements. In addition, part or all of the steps in each flowchart can include multiple sub-steps or multiple stages, part or all of the sub-steps or stages can be executed at the same time, and each of the sub-steps or stages can also be executed at different times. In the scenario where the execution times are different, the execution order of the sub-steps or stages can be flexibly configured based on requirements, and the embodiments of the present application do not limit this.
[0220] The above only describes optional implementation manners of some implementation scenarios of the present application. It should be noted that, for those skilled in the art, other similar implementation manners according to the technical concept of the present application can also be adopted without departing from the technical concept of the present application, and these also belong to the protection scope of the embodiments of the present application.
Claims
1. A method for generating outline content using natural language processing, characterized in that, The method includes: The system receives text from the scientific research field, performs semantic parsing on the text, and obtains a set of problem features corresponding to the text. The set of problem features includes relevant gene / protein names, scientific research problem type classifications, and text feature quantification parameters. The text feature quantification parameters are positively correlated with the number of qualifying words and entity names contained in the text. Based on the problem feature set, the biomedical literature database interface is called to perform basic scientific research literature knowledge retrieval processing to obtain a basic scientific research literature knowledge set that matches the problem feature set. The basic scientific research literature knowledge set includes the core viewpoints, descriptions of basic experimental methods, and descriptions of molecular mechanisms of multiple basic scientific research papers. The basic scientific research literature knowledge set is processed by knowledge unit extraction. Basic scientific research knowledge units with independent semantics are extracted from the core viewpoints, descriptions of basic experimental methods and descriptions of molecular mechanisms in the basic scientific research literature knowledge set, and a structured knowledge unit set is generated. The structured knowledge unit set is associated and matched with the problem feature set. Based on the scientific research problem type classification and text feature quantification parameters, the basic scientific research knowledge units in the structured knowledge unit set are weighted and weighted to generate a weighted knowledge integration result. Based on the weighted knowledge integration results, a text for answering scientific research questions is generated, and the text for answering scientific research questions is fed back to the user interface of the scientific research field for display. The process of associating and matching the structured knowledge unit set with the problem feature set, and assigning weights to the basic scientific research knowledge units in the structured knowledge unit set according to the scientific research problem type classification and text feature quantification parameters to generate a weighted knowledge integration result includes: Extract the text content of each basic scientific research knowledge unit in the structured knowledge unit set, match the text content of the basic scientific research knowledge unit with the entity name set in the problem feature set, and calculate the number of entity names and the matching degree contained in each basic scientific research knowledge unit. Based on the scientific research problem type classification in the problem feature set, determine the knowledge unit type weight corresponding to the scientific research problem type classification, and assign basic weight values to different types of basic scientific research knowledge units. Based on the text feature quantification parameters, the basic weight values are adjusted, and the weight values of each basic scientific research knowledge unit are normalized. The basic scientific research knowledge units in the structured knowledge unit set are sorted according to the normalized weight values. The sorted basic scientific research knowledge units are grouped and integrated according to the topic cluster identifier to generate a weighted knowledge integration result containing weight information and topic classification.
2. The outline content generation method combining natural language processing according to claim 1, characterized in that, The process involves receiving text from the scientific research field, performing semantic parsing on the text to obtain a set of problem features corresponding to the text, including: Receive text from the scientific research field, perform text format standardization processing on the text from the scientific research field, remove special symbols and irrelevant characters from the text from the scientific research field, and retain the valid text content containing basic scientific research terminology; The effective text content is segmented into words by calling a professional word segmentation tool to break it down into multiple word units, which include nouns, verbs and adjectives. Entity recognition processing is performed on the multiple word units. A pre-trained entity recognition model is used to identify entity words in the multiple word units that belong to the categories of related gene / protein names, cell types, molecular markers and experimental methods, forming a set of entity names. The effective text content is classified into scientific research question types. Based on the semantic relationship between the entity name set and word units, the text in the scientific research field is classified into mechanism research, function verification, biomarker screening or mechanism exploration, and a scientific research question type classification identifier is generated. The text feature quantification parameters in the effective text content are quantified, and the text feature quantification parameters are calculated based on the number of qualifying words, interrogative words and entity names contained in the text of the scientific research field. The entity name set, the scientific research question type classification identifier, and the text feature quantification parameters are integrated to generate a question feature set containing specific semantic information in the field of basic scientific research.
3. The outline content generation method combining natural language processing according to claim 2, characterized in that, The entity recognition process is performed on the multiple word units. A pre-trained entity recognition model identifies entity words belonging to the categories of related gene / protein names, cell types, molecular markers, and experimental methods within the multiple word units, forming an entity name set, including: The multiple word units are input into a pre-trained entity recognition model, which includes an embedding layer, a bidirectional recurrent neural network layer, and a conditional random field layer. The embedding layer converts each word unit into a low-dimensional dense word vector, preserving the semantic information of the word; The bidirectional recurrent neural network layer is used to extract contextual features from the word vectors, capturing the semantic association information of word units before and after the text sequence. The feature vectors output by the bidirectional recurrent neural network layer are sequence labeled through the conditional random field layer, and an entity category label is assigned to each word unit. The entity category label includes related gene / protein labels, cell type labels, molecular marker labels, and experimental method labels. The labeled word units are processed by entity boundary merging, which merges consecutive word units with the same entity category label into a complete entity name; Remove duplicate entities from the merged entity names, retain the unique entity names, and then categorize and organize all the unique entity names according to the entity category label to form an entity name set.
4. The outline content generation method combining natural language processing according to claim 1, characterized in that, The step of calling the biomedical literature database interface based on the problem feature set to perform basic scientific research literature knowledge retrieval processing and obtain a basic scientific research literature knowledge set matching the problem feature set includes: The entity name set in the problem feature set is parsed, and each entity name in the entity name set is converted into a standard search term supported by the biomedical literature database to generate a standardized search term list; Based on the scientific research problem type classification identifier in the problem feature set, the corresponding literature retrieval strategy is selected from the preset retrieval strategy library. The literature retrieval strategy includes the retrieval field combination method, logical operator configuration and literature publication time range restriction. The standardized search term list is combined with the document retrieval strategy to generate a structured search expression, which contains multiple search terms and their logical relationships. The structured search expression is sent to the biomedical literature database by calling the biomedical literature database interface, triggering the literature search operation and obtaining a preliminary search result set, which includes the titles, author information and abstract fragments of multiple basic scientific research papers. The preliminary search result set is sorted by relevance. Based on the text feature quantification parameters in the question feature set and the matching degree between the entity name set and the document content, the basic scientific research documents in the preliminary search result set are sorted, and a preset number of documents with the highest ranking are selected to form the target document set. Extract the core viewpoints, basic experimental method descriptions, and molecular mechanism descriptions from each basic research paper in the target literature set, and integrate them into a basic research literature knowledge set.
5. The outline content generation method combining natural language processing according to claim 4, characterized in that, The preliminary search result set is then subjected to relevance ranking processing. Based on the text feature quantification parameters in the question feature set and the matching degree between the entity name set and the document content, the basic scientific research documents in the preliminary search result set are ranked, and a predetermined number of documents with the highest ranking are selected to form a target document set, including: Extract the title and abstract fragments of each basic scientific research article from the preliminary search result set to form the article content text; The document content text is matched with the entity name set in the problem feature set. The number of entity names and the matching frequency contained in each basic scientific research document are calculated, and an entity matching score is generated. The semantic similarity calculation process is performed on the document content text, and the document content text and the scientific research field text are converted into semantic vectors. The cosine similarity between the semantic vectors is calculated as the semantic relevance score. Based on the text feature quantification parameters in the problem feature set, the entity matching score and semantic relevance score are weighted and summed to generate a comprehensive document relevance score. The higher the text feature quantification parameter, the greater the weight of the entity matching score. The basic research literature in the preliminary search results set is sorted from high to low according to the comprehensive relevance score of the literature; Based on the preset number of literatures to be selected, a target literature set is formed by selecting the top-ranked literatures from the sorted basic research literatures.
6. The outline content generation method combining natural language processing according to claim 1, characterized in that, The process of extracting knowledge units from the basic scientific research literature knowledge set involves extracting semantically independent basic scientific research knowledge units from the core viewpoint summaries, basic experimental method descriptions, and molecular mechanism descriptions in the basic scientific research literature knowledge set, generating a structured knowledge unit set, including: The core viewpoints, basic experimental methods, and molecular mechanisms in the aforementioned basic scientific research literature knowledge set are traversed, and the corresponding text content is divided into multiple text paragraph units. Semantic topic recognition is performed on each text paragraph unit. The core topic of each text paragraph unit is determined by a pre-trained topic model, and topic tags are generated. Based on the topic tags, the text paragraph units are clustered, and text paragraph units with the same or similar topic tags are classified into the same topic cluster. Each topic cluster corresponds to a basic scientific research knowledge topic. For each topic cluster, the text paragraph unit is processed for knowledge unit boundary identification. Based on the semantic rules and punctuation features of the basic scientific research field, the text paragraph unit is divided into multiple basic scientific research knowledge units with independent semantics. The basic scientific research knowledge units include declarative knowledge units, methodological knowledge units and conclusion-type knowledge units. Add source literature identifiers and topic cluster identifiers to each basic scientific research knowledge unit to establish the association between the basic scientific research knowledge unit and the original basic scientific research literature and topic cluster; All basic scientific research knowledge units are grouped and arranged according to topic cluster identifiers to generate a set of structured knowledge units containing topic classification information and source association information.
7. The outline content generation method combining natural language processing according to claim 6, characterized in that, The process involves performing knowledge unit boundary identification on text paragraph units within each topic cluster. Based on semantic rules and punctuation features in the field of basic scientific research, the text paragraph units are segmented into multiple basic scientific research knowledge units with independent semantics, including: Punctuation mark recognition processing is performed on the text paragraph units in each topic cluster to mark the positions of punctuation marks in the text paragraph units. The punctuation mark positions include the positions of period marks, semicolon marks, colon marks and dash marks. Based on semantic rules in basic scientific research, the system identifies conjunctions in text paragraph units that indicate causal, adversative, and progressive relationships, and determines the position of these conjunctions within the text paragraph units. By combining the positions of punctuation marks and conjunctions, the text paragraph units are initially divided into multiple semantic segments; Perform a basic scientific terminology integrity check on each semantic segment to ensure that the segmented semantic segments contain complete entity names and scientific concepts; A secondary segmentation process is performed on semantic segments containing multiple short sentences, and the segments are divided into independent basic scientific research knowledge units based on the logical relationships between the short sentences. Add type tags to each basic scientific research knowledge unit, and mark it as a declarative knowledge unit, a methodological knowledge unit, or a conclusion-type knowledge unit according to its content.
8. The outline content generation method combining natural language processing according to claim 1, characterized in that, The step of generating research question answer text based on the weighted knowledge integration result and displaying the research question answer text on the user interface in the research field includes: Analyze the sequence of topic knowledge groups in the weighted knowledge integration result to determine the topic description information and the basic scientific research knowledge units contained in each topic knowledge group; Based on the sorting order of the topic knowledge group sequence, a corresponding title paragraph is generated for each topic knowledge group. The title paragraph contains topic description information and an explanation of the importance of the topic knowledge group in the solution. The basic scientific research knowledge units in each topic knowledge group are processed by text integration, and citation marks are added to the integrated text content. The citation marks correspond to the source literature marks of the basic scientific research knowledge units, which makes it convenient for users to trace the original literature. Combine the title paragraphs of all the subject knowledge groups and the integrated text content in order to generate a complete text of the scientific research question answer; The text of the research question answer is sent to the user interaction interface in the research field, where it is displayed below the research text submitted by the user, and a link to view the original literature is provided.
9. A system for generating outline content using natural language processing, characterized in that, The method includes a processor and a computer-readable storage medium storing machine-executable instructions that, when executed by the processor, implement the outline content generation method incorporating natural language processing as described in any one of claims 1-8.
Citation Information
Patent Citations
Medical subject knowledge service method and device, electronic equipment and medium
CN119903129A
Method and apparatus for automatically generating inference questions and answers
WO2021184311A1