Method for generating multi-type question and answer data chain for science and technology literature question and answer system
By analyzing the content and structure of scientific and technological documents and generating multi-round question chains, the shortcomings of existing scientific and technological document question-and-answer systems in terms of multiple types of follow-up question modes are solved, enabling in-depth understanding and logical reasoning of scientific research content and improving the effectiveness and usability of the question-and-answer system.
Patent Information
- Application Number
- CN202511056708.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-07-30
AI Technical Summary
Existing scientific literature question-answering systems lack dynamic generation and chain structure control mechanisms for various types of follow-up questions. This results in a single type of question chain generated by the system and insufficient semantic coverage. It is difficult to systematically examine the AI model's deep understanding and logical reasoning ability of scientific research content, which affects the effectiveness and usability of scientific literature understanding systems in key scenarios such as academic assistance, scientific research evaluation, and intelligent recommendation.
By calling a pre-trained language model to perform content and structural hierarchy analysis on scientific and technological documents, initial questions are generated. Based on the document content features, initial question types, and answer content, the question-answer type library is matched, follow-up question types are dynamically selected, and multi-round question chains are constructed. Combined with knowledge relationship graphs, semantic hierarchy recognition and structural hierarchy distribution are performed to achieve automatic generation, evaluation, and optimization of question chains.
It improves the integrity of scientific literature question-answering systems in multi-type semantic understanding, multi-round logical modeling, and intelligent data generation, providing structured, high-quality, and highly adaptable question samples to support the deep understanding and logical reasoning capabilities of academic AI.
Smart Images

Figure CN120929568B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data generation technology, and in particular to a method for generating multi-type question-and-answer data chains for scientific and technological literature question-and-answer systems. Background Technology
[0002] Scientific literature question-answering systems are a key application area in the integration of artificial intelligence and natural language processing technologies in recent years, widely used in scenarios such as research assistance, academic retrieval, automatic summarization, and intelligent evaluation. Current mainstream scientific literature question-answering methods are mostly based on pre-trained language models or information retrieval models, understanding fragments of content in the literature and generating questions. Currently, most existing scientific literature question-answering systems only support shallow question-answering based on static templates or single-round information extraction, lacking comprehensive modeling capabilities for complex questioning logics such as multi-round, progressive, counterfactual, and evidence-based questions. They cannot realistically reproduce the actual process of researchers gradually reading literature, raising questions layer by layer, and gaining a deeper understanding.
[0003] In summary, existing technologies suffer from a lack of dynamic generation and chain-structure control mechanisms for multi-type questioning patterns. This results in question-answering systems generating single-type question chains, insufficient semantic coverage, and loose structures, making it difficult to systematically examine the AI model's deep understanding and logical reasoning capabilities regarding scientific research content. Consequently, this technology impacts the effectiveness and usability of scientific literature understanding systems in key scenarios such as academic assistance, research evaluation, and intelligent recommendation. Summary of the Invention
[0004] The purpose of this application is to provide a method for generating multi-type question-and-answer data chains for scientific literature question-and-answer systems. This method addresses the technical problem in existing technologies where the lack of dynamic generation and chain structure control mechanisms for multi-type follow-up question patterns leads to single-type question chains, insufficient semantic coverage, and loose structure. This makes it difficult to systematically examine the AI model's deep understanding and logical reasoning ability of scientific research content, and further affects the effectiveness and usability of scientific literature understanding systems in key scenarios such as academic assistance, scientific research evaluation, and intelligent recommendation.
[0005] In view of the above problems, this application provides a method for generating multi-type question-and-answer data chains for scientific and technological literature question-and-answer systems, including: calling a pre-trained language model to perform content and structural hierarchical parsing of the scientific and technological literature to be questioned and answered, obtaining literature content features, and generating initial questions based on the literature content features; using the literature content features, initial question types, and answer content as indexes, matching in a question-and-answer type library to obtain a set of follow-up question types; dynamically selecting at least one follow-up question type from the set of follow-up question types, constructing follow-up questions, and obtaining follow-up question answer content; performing re-matching in the set of follow-up question types based on the determination results of the follow-up question types and the follow-up question answer content to generate multiple rounds of questions; evaluating the question chain of initial question-follow-up question-multi-round questions based on the literature content features, performing question chain compensation optimization based on the evaluation results, and generating a question data chain using the question-and-answer and parsing data corresponding to the question chain that meets the evaluation objectives.
[0006] Preferably, the method for generating multi-type question-and-answer data chains for a scientific and technological literature question-and-answer system further includes: performing structured analysis on the scientific and technological literature to be questioned, extracting chapter elements, content types and distributions, and constructing a knowledge relationship graph; performing semantic hierarchical recognition on the content of each node based on the knowledge relationship graph, establishing structural hierarchical distribution and semantic relationships, and adding them to the knowledge relationship graph; and performing core content analysis based on the knowledge relationship graph to determine the literature type and content characteristics, thereby obtaining the literature content characteristics.
[0007] Preferably, the method for generating multi-type question-and-answer data chains for a scientific literature question-and-answer system further includes: the literature types include: review type, method type, experimental type, and theoretical type.
[0008] Preferably, the method for generating multi-type question-and-answer data chains for a scientific and technological literature question-and-answer system further includes: establishing a mapping relationship between the literature type, content features, and question-and-answer types, wherein the mapping relationship includes a first-level mapping corresponding to the literature type and a second-level mapping corresponding to the content features; performing first-level mapping and second-level mapping in the mapping relationship according to the literature content features to generate initial questions.
[0009] Preferably, the method for generating multi-type question-and-answer data chains for a scientific literature question-and-answer system further includes: establishing a primary mapping between each literature type and question-and-answer type based on historical question-and-answer samples of each literature type, wherein the question-and-answer types corresponding to review types include comparison chains, controversy chains, and developmental chains; the question-and-answer types corresponding to methodological types include causal chains, application chains, and algorithm optimization chains; the question-and-answer types corresponding to experimental types include evidence chains, variable control chains, and parameter sensitivity chains; and the question-and-answer types corresponding to theoretical types include definition chains, derivation chains, and abstract relationship chains; analyzing the response relationship between content features and question-and-answer types based on historical question-and-answer samples to establish a secondary mapping between content features and question-and-answer types; and obtaining the mapping relationship between the literature type, content features, and question-and-answer types based on the primary and secondary mappings.
[0010] Preferably, the method for generating multi-type question-and-answer data chains for a scientific literature question-and-answer system further includes: establishing follow-up question-and-answer types, including progressive, counterfactual, and evidentiary types; analyzing the matching conditions, generation mechanisms, and type reinforcement rules for each follow-up question-and-answer type; establishing the corresponding matching relationships between the literature content features, initial question types, and answer content, and the matching conditions, generation mechanisms, and type reinforcement rules, and constructing the question-and-answer type library.
[0011] Preferably, the method for generating multi-type question-and-answer data chains for a scientific literature question-and-answer system further includes: calling a predefined template for matching follow-up question types based on the generation mechanism and type enhancement rules; filling entity, relation, and numerical parameters from the structured literature data into the predefined template; and semantically enhancing the output of the predefined template through a large language model to generate follow-up questions that conform to scientific research logic.
[0012] Preferably, the method for generating multi-type question-and-answer data chains for a scientific literature question-and-answer system further includes: locating relevant literature content based on the type and type of follow-up question; extracting core keywords and core relationships based on the relevant literature content, and semantically expanding and adjusting the core keywords and core relationships using natural language to generate a reference answer; using the reference answer to determine the similarity of the follow-up answer content, obtaining a determination result, filtering follow-up questions based on the determination result, and performing re-matching using the filtered follow-up question types to generate multiple rounds of questions.
[0013] Preferably, the method for generating multi-type question-and-answer data chains for a scientific literature question-and-answer system further includes: using the initial question as the root node, follow-up questions and multi-round questions as child nodes, marking edge relationships through semantic similarity calculation and a logical relationship classifier, and annotating source text location information for cross-paragraph and document nodes to establish a question link graph; based on the question link graph, performing question content correlation and response judgment analysis on the answer content of each question to obtain a heat map of literature answer evaluation; based on the heat map of literature answer evaluation, performing coverage depth analysis through the literature content features to obtain the location distribution of missing documents; and based on the evaluation target of literature question-and-answer coverage depth, combining the heat map to perform question compensation on the location distribution of missing documents.
[0014] Preferably, the method for generating multi-type question-and-answer data chains for a scientific and technological literature question-and-answer system further includes: constructing a sample library through question data chains to analyze the topic types of various scientific and technological literatures, and constructing a question link reference library for each topic type. The question link reference library is used for the initial recommendation of question-and-answer data chains for different topic types in the later stage.
[0015] The technical solution provided in this application has at least the following technical effects or advantages: by achieving the technical goal of constructing a question-and-answer data chain based on the automatic generation of question chains driven by document structure and semantic features, adaptive matching of follow-up question types, and linkage control of question chain evaluation and compensation optimization, the technical effect of improving the ability of scientific literature question-and-answer systems in terms of multi-type semantic understanding, multi-round logical modeling and intelligent generation of data integrity is achieved, providing structured, high-quality and highly adaptable question samples to support academic AI.
[0016] The above description is merely an overview of the technical solution of this application. To better understand the technical means of this application and to facilitate its implementation according to the description, and to make the above and other objects, features, and advantages of this application more apparent, specific embodiments of this application are described below. It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent through the following description. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0018] Figure 1This is a flowchart illustrating the method for generating multi-type question-and-answer data chains for a scientific and technological literature question-and-answer system, as described in this application.
[0019] Figure 2 This is a flowchart illustrating the process of obtaining document content features in the multi-type question-and-answer data chain generation method for scientific and technological document question-and-answer systems in this application. Detailed Implementation
[0020] This application provides a method for generating multi-type question-and-answer data chains for scientific literature question-and-answer systems. It addresses the technical problem in existing technologies where the lack of dynamic generation and chain-structure control mechanisms for multi-type follow-up question patterns leads to single-type question chains, insufficient semantic coverage, and loose structures. This makes it difficult to systematically assess the AI model's deep understanding and logical reasoning capabilities regarding research content, further impacting the effectiveness and usability of scientific literature understanding systems in key scenarios such as academic assistance, research evaluation, and intelligent recommendation. The method achieves the technical goal of constructing question-and-answer data chains based on document structure and semantic features, adaptive matching of follow-up question types, and coordinated control of question chain evaluation and compensation optimization. This enhances the capabilities of scientific literature question-and-answer systems in multi-type semantic understanding, multi-round logical modeling, and the completeness of intelligently generated data, providing structured, high-quality, and highly adaptable question samples to support academic AI.
[0021] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. It should be understood that this application is not limited to the exemplary embodiments described herein. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application. It should also be noted that, for ease of description, only the parts related to this application are shown in the accompanying drawings, not all of them.
[0022] Please see the appendix Figure 1 This application provides a method for generating multi-type question-and-answer data chains for scientific and technological literature question-and-answer systems, specifically including the following steps:
[0023] The pre-trained language model is invoked to perform content and structural hierarchical analysis on the scientific and technological documents to be answered, obtain the document content features, and generate initial questions based on the document content features.
[0024] Specifically, pre-trained language models can identify semantic relationships and linguistic features in scientific and technological literature. The scientific and technological literature to be questioned refers to the scientific papers or technical documents to be processed. The pre-trained language model, trained on a large-scale text corpus, is invoked to understand and process the scientific and technological literature to be questioned. This involves content parsing, extracting concepts, definitions, research objectives, methods, experimental procedures, data results, etc., and establishing a hierarchical content structure to facilitate the generation and location of subsequent question-and-answer questions. Then, structural hierarchy parsing is performed, further analyzing the organization of chapters, paragraphs, and sections based on content parsing. For example, typical structural modules such as introduction, methods, results, and discussion are identified, and hierarchical relationships are established to obtain the content characteristics of the literature.
[0025] By analyzing the structured information extracted from the scientific and technological literature to be questioned, such as chapter titles, paragraph themes, semantic levels, and keyword distribution, the key content in the literature is identified. Then, questions that can summarize or guide the content are constructed, thus generating the initial questions.
[0026] Using the document content features, initial question type, and answer content as indexes, a set of follow-up question types is obtained by matching in the question-and-answer type library.
[0027] Specifically, the system uses document content features, initial question type, and answer content as indexes, combining these three key pieces of information as search criteria for comprehensive matching. Each serves a different function: document content features provide contextual background, such as whether the document is a review or experimental study; the initial question type reflects the logical intent of the current question, such as whether it is a definitional, methodological, or hypothetical question; and the answer content analyzes whether the current answer involves elements such as process, assumptions, and data. By combining these three elements, a ternary index structure is formed for accurately finding the most suitable follow-up question type. Matching is performed in a question-and-answer type library, which calls a pre-built question-and-answer type knowledge base to match the input ternary index information. The question-and-answer type library stores the applicable conditions, template styles, and triggering mechanisms for different follow-up question types. Common follow-up question types include progressive, counterfactual, and evidential. The matching process mainly relies on structured rules, semantic tags, and language models to achieve automatic identification and return potentially suitable follow-up question types. For example, if the index information shows that the initial question is methodological and the answer contains a statement comparing it to existing methods, progressive follow-up questions will be matched first. After matching is completed, a set of one or more candidate follow-up question types is obtained, which represents the follow-up question type set that may be used in the current context.
[0028] Dynamically select at least one follow-up question type from the set of follow-up question types, construct follow-up questions, and obtain follow-up question answers.
[0029] Specifically, at least one type of follow-up question is dynamically selected from a set of question types, such as progressive, counterfactual, or evidential. Based on the semantic features of the initial question and the current answer, as well as the context of the document structure, the most suitable follow-up questioning method is automatically chosen to further explore the question. Dynamic selection means that the optimal follow-up path is determined in real-time based on each specific question and answer. For example, when the initial question asks about the effect of a method, a progressive follow-up question might be chosen: "Under what conditions is this method most effective?"; while when a hypothetical statement appears in the answer, a counterfactual follow-up question might be chosen: "If the experimental temperature is lowered by 20 degrees Celsius, will it still maintain stability?" Predefined templates are called, and key fields are filled in with keywords, parameters, causal relationships, etc., from the structured data of the document to construct follow-up questions, ensuring that the follow-up questions are professional and coherent. By locating the paragraphs in the document most relevant to the follow-up questions, key information is extracted and rewritten to obtain the follow-up question answers, generating response content with semantic consistency and logical coherence. For example, when asked "Did the authors provide experimental evidence to support this conclusion?", the system might retrieve paragraphs containing specific sample sizes, experimental results, and significance levels. These paragraphs could then be organized into concise answer paragraphs, forming extension units of the question-and-answer chain. This allows for a layered examination of model understanding and reasoning abilities. Each additional round of follow-up questions not only increases the depth of the question-and-answer process but also broadens the coverage of detailed understanding of the literature. The resulting question chain can better simulate the interactive exploration process in real scientific research.
[0030] Based on the determination results of the follow-up question type and the content of the follow-up question answer, a rematch is performed on the set of follow-up question types to generate multiple rounds of questions.
[0031] Specifically, after generating follow-up questions, the system will jointly analyze the type of the follow-up question (such as progressive, counterfactual, or evidentiary) and the corresponding answer content to determine whether the current follow-up question has achieved the expected knowledge expansion effect or whether there is still room for further exploration.
[0032] The judgment result is obtained by evaluating indicators such as semantic similarity calculation, logical consistency detection or keyword coverage. For example, if the follow-up answer only involves the surface phenomenon and does not touch the causal relationship, the current follow-up question is considered to be insufficient in depth.
[0033] Subsequently, rematching is performed on the set of follow-up question types. Based on the current content, it is determined whether the follow-up question type needs to be changed or expanded, leading to the next round of question generation and the formation of a more complex question-and-answer chain. For example, if the first round of follow-up questions is evidence-based and the answer lacks specific data, the rematch may shift to more exploratory counterfactual or progressive follow-up questions, forming a new question such as "If this method is not used, can the same performance still be achieved?" This iterative process, building upon each round of follow-up questions, forms a multi-round question-and-answer chain consisting of the initial question, the first follow-up question, the second follow-up question, and so on. Through this multi-round design, the literature content can be examined at different dimensions and depths. For example, first ask "What is the core principle of this method?", then "Is this principle applicable to different data scales?", and then "If the input data dimension increases tenfold, how should this method be adjusted?" Each round increases complexity and challenge, allowing the model's understanding ability to be more comprehensively tested and trained in a real scientific research context.
[0034] The question chain of initial question - follow-up question - multiple round questions is evaluated based on the characteristics of the document content. The question chain is compensated and optimized according to the evaluation results. Question data chain is generated by using the question and answer and parsing data corresponding to the question chain that meets the evaluation objectives.
[0035] Specifically, based on the structural characteristics, semantic elements, and knowledge density of scientific literature, the question-and-answer chain is comprehensively scored from dimensions such as coverage, logical consistency, and semantic depth. Literature content characteristics may include chapter structure, paragraph themes, frequency of key terms, and density of causal chains. For example, if a methodological document contains frequently occurring algorithmic terms and step descriptions, the evaluation will focus more on whether the question chain covers the methodological logic and its improvement basis. Next, if a section of the literature is found to be uncovered by questions, or if a certain type of follow-up question is missing, such as a lack of counterfactual exploration or parameter sensitivity analysis, supplementary questions are automatically generated to complete the chain. For example, if a follow-up question related to variable control is missing, a supplementary question might be added: "If other conditions remain unchanged, will changing only parameter X affect the result?" Finally, comprehensive, logically rigorous, and semantically profound question-and-answer chains are selected, and their questions, answers, and their location relationships within the literature are packaged as meta-information and used as standardized training or evaluation data samples to form a question data chain. Taking a 5,000-word scientific paper as an example, 30 initial questions and their multiple rounds of follow-up questions can be generated from it. Finally, a high-quality chain covering 80% of the paragraphs and with an average follow-up question level of more than three rounds is retained, forming a set of training data that is both challenging and interpretable, serving the intelligent modeling and evaluation tasks of scientific research question answering systems.
[0036] Please see the appendix Figure 2Furthermore, this application also includes: performing structured analysis on the scientific and technological documents to be questioned, extracting chapter elements, content types and distributions, and constructing a knowledge relationship graph; performing semantic hierarchical recognition on the content of each node based on the knowledge relationship graph, establishing structural hierarchical distribution and semantic relationships, and adding them to the knowledge relationship graph; performing core content analysis based on the knowledge relationship graph, determining the document type and content characteristics, and obtaining the document content characteristics.
[0037] Specifically, the scientific literature to be questioned undergoes structured parsing, processing it according to its formatted features such as chapters, paragraphs, and titles, transforming unstructured text information into a structured representation with clear hierarchical relationships. Then, the chapter elements of the scientific literature to be questioned are identified, such as introduction, related work, methods, experiments, and conclusions, and further, the content types within each chapter are extracted, such as method descriptions, performance analyses, and hypothesis statements, while recording the specific location distribution of content types within the scientific literature to be questioned. Based on the extracted structured information, the core entities, attributes, and logical relationships in the scientific literature to be questioned are organized in the form of a graph, constructing a knowledge relationship graph and forming a semantic network composed of nodes and edges.
[0038] Based on a knowledge relationship graph, semantic hierarchy identification is performed on the content of each node. The content represented by each node is semantically categorized, and its position within the semantic understanding hierarchy is determined. For example, definitional content is at the basic level, principle-based content is at the intermediate level, and reasoning or innovation points are at the higher level, thus establishing a hierarchical distribution and semantic relationships. By establishing this hierarchical distribution and semantic relationships, the role of each piece of content in the overall semantic structure of the scientific document to be answered becomes clear. The hierarchical distribution and semantic relationships are then re-added to the knowledge relationship graph, continuously enriching its content. This includes not only entities and connections but also the semantic hierarchy labels and structural positions of each node, giving the knowledge relationship graph stronger representational and semantic organization capabilities, which helps to accurately locate specific question-and-answer target content within the document.
[0039] Core content analysis based on knowledge relationship graphs, utilizing structural hierarchy and semantic relationships, can identify the main points of the scientific literature to be questioned, such as main methods, key parameters, and core conclusions. This allows for further determination of the overall type of the literature. For example, review articles involve comparisons of multiple methods, methodological articles emphasize new algorithmic structures, experimental articles emphasize experimental design and data validation, while theoretical articles often involve logical reasoning and mathematical derivation. Based on this, the type of literature and its specific content characteristics, such as content emphasis, expression style, and language style, are determined, thus obtaining the literature's content characteristics—a comprehensive description of the literature's content. A structured model of the semantic distribution and logical construction of the literature provides a solid knowledge foundation for subsequent question generation and follow-up matching.
[0040] Furthermore, this application also includes: the document types include: review type, method type, experimental type, and theoretical type.
[0041] Specifically, document type refers to the classification of scientific and technological documents according to their writing purpose, content structure, and research focus. It serves as the basis for subsequent question-and-answer generation and structure matching. In the process of constructing the question-and-answer data chain, accurately identifying document type helps to rationally plan the semantic paths of initial questions and follow-up questions, thereby improving the depth and professionalism of the question-and-answer chain.
[0042] Literature types include review articles, methodological articles, experimental articles, and theoretical articles. Review articles are scientific literature whose main goal is to systematically review and summarize existing achievements in a research field. They contain numerous references and compare, classify, and evaluate existing methods, technologies, and viewpoints. The structure of review articles typically revolves around the development history, current research status, advantages, and disadvantages, making them suitable for generating comparative, controversial, and trend-related questions. Methodological articles primarily focus on the proposal and implementation process of new technologies or algorithms, emphasizing a detailed explanation of the innovative points, structural design, and logical principles of the method. Methodological articles are suitable for generating causal, optimization, and application mechanism-related questions and answers. Experimental articles focus on demonstrating experimental design, data collection, result analysis, and comparative verification processes to evaluate the practical performance of a method. Experimental articles are suitable for constructing evidence chains, variable control chains, or parameter sensitivity chains. Theoretical articles focus on the proposal of basic concepts, the construction of mathematical models, or the process of logical reasoning; their core is rigorous logical relationships and abstract expression. Theoretical articles are suitable for generating definition chains, derivation chains, or abstract relational chains.
[0043] Furthermore, this application also includes: establishing a mapping relationship between the document type, content features, and question-and-answer type, wherein the mapping relationship includes a first-level mapping corresponding to the document type and a second-level mapping corresponding to the content features; performing first-level mapping and second-level mapping in the mapping relationship according to the document content features to generate an initial question.
[0044] Specifically, this involves establishing a mapping relationship between document type, content features, and question-and-answer type. Based on a large collection of scientific literature and question-and-answer samples, statistical analysis and pattern induction are used to construct a structured association mechanism that can be readily applied. This mechanism matches different types of literature and their internal content features with the most suitable question-and-answer type for generation. The mapping relationship is divided into two levels. First, the primary mapping corresponding to document type is based on the classification attributes of the entire document, such as review, methodological, experimental, or theoretical, establishing a correspondence with broad question-and-answer types. For example, review documents tend to generate comparison chains and controversy chains. Second, the secondary mapping corresponding to content features is a matching relationship established based on the fine-grained content within the document. For example, the content features describing the trend of technological evolution in a certain section of the text are suitable for mapping to generate developmental context-based questions and answers. This two-level mapping mechanism can more accurately capture the role and semantic guidance of the overall structure and local information of the document in the question-and-answer construction process.
[0045] Based on the characteristics of the document content, first-level mapping and second-level mapping are performed in the mapping relationship to generate an initial question. That is, after obtaining the document to be processed, the possible question types are first determined in the first-level mapping table based on the type information of the entire document. Then, based on the specific content characteristics in the document, more refined question types that fit the contextual semantics are selected in the second-level mapping table. Finally, an initial question with logical clarity and professionalism is synthesized.
[0046] Furthermore, this application also includes: establishing a primary mapping between each document type and question-and-answer type based on historical question-and-answer samples of each document type, wherein the question-and-answer types corresponding to review document types include comparison chains, controversy chains, and developmental chains; those corresponding to methodological document types include causal chains, application chains, and algorithm optimization chains; those corresponding to experimental document types include evidence chains, variable control chains, and parameter sensitivity chains; and those corresponding to theoretical document types include definition chains, derivation chains, and abstract relationship chains; establishing a secondary mapping between content features and question-and-answer types based on the response relationship between content features and question-and-answer types analyzed from historical question-and-answer samples; and obtaining the mapping relationship between the document type, content features, and question-and-answer types based on the primary and secondary mappings.
[0047] Specifically, based on historical question-and-answer samples for each document type, that is, by analyzing a large amount of existing question-and-answer data, the system categorizes the question-and-answer types associated with each document type, thus forming a one-to-one or one-to-many mapping relationship, and establishing a primary mapping between each document type and question-and-answer type. This primary mapping reflects the question-and-answer logical paths that different types of documents tend to adopt during the question-and-answer process. For example, for review documents, the system observes the frequent occurrence of questions such as "What are the performance differences between Model A and Model B?" or "In what context has this method been questioned?" in their historical question-and-answer pairs, thus summarizing that comparison chains, controversy chains, and development context chains are typical question-and-answer types for this type of document. For methodological documents, questions such as "What are the key improvement steps of this method?" or "In which scenarios can this technology be applied?" abstract causal chains, application chains, and algorithm optimization chains as their main question-and-answer types. Similarly, experimental documents focus on experimental design and parameter changes, thus establishing connections with evidence chains, variable control chains, and parameter sensitivity chains. Theoretical documents revolve around concept definitions and logical reasoning, with typical question-and-answer types including definition chains, derivation chains, and abstract relation chains.
[0048] By analyzing historical question-and-answer samples, a two-level mapping between content features and question-and-answer types can be established. For example, in the same review article, if keywords and semantic features such as "advantages and disadvantages analysis," "performance comparison," or "time evolution trend" frequently appear in a certain section, the appropriate question-and-answer type can be determined based on the content features, suggesting a comparison chain or development chain. By analyzing the response patterns of different content elements in historical question-and-answer data, a more granular mapping between content and question-and-answer types can be established, i.e., a two-level mapping.
[0049] Based on primary and secondary mappings, the mapping relationship between document type, content features, and question-and-answer type is obtained. This involves combining the overall feature mapping at the document level with the local feature mapping at the content level to form a complete triplet mapping. This allows for a preliminary judgment of the possible appropriate question-and-answer type based on the document type, and further refinement by combining the content features of specific paragraphs within the document, thereby supporting more accurate question type selection.
[0050] Furthermore, this application also includes: establishing follow-up question and answer types, including progressive, counterfactual, and evidentiary types; analyzing the matching conditions, generation mechanisms, and type reinforcement rules for each follow-up question and answer type; establishing the corresponding matching relationships between the document content features, initial question types, and answer content, and the matching conditions, generation mechanisms, and type reinforcement rules, and constructing the question and answer type library.
[0051] Specifically, this paper establishes three types of follow-up questions and answers, classifying them into three representative approaches based on different research literature question-and-answer needs: progressive, counterfactual, and evidential. Progressive follow-up questions emphasize further exploration beyond the original answer, such as expanding from methodological principles to optimization strategies. Counterfactual follow-up questions raise "what if...?" questions based on changes in hypothetical conditions, simulating reverse reasoning about variable relationships. Evidential follow-up questions require answers to provide empirical evidence such as data or charts to verify the credibility and support of the arguments. These three types of follow-up questions help cover common paths to in-depth understanding and logical extensions found in research literature.
[0052] This paper analyzes the matching conditions, generation mechanisms, and type reinforcement rules for each type of follow-up question and answer, further defining the mechanisms of the three types of follow-up questions. Matching conditions determine whether a follow-up question and answer type is applicable to the current question or answer context, such as the existence of a causal structure or supporting data. Generation mechanisms define the operational methods for question construction; for example, progressive follow-up questions generate questions by identifying process verbs and temporal relationships, counterfactual follow-up questions rely on the extraction of conditional variables and reverse semantic construction, and evidential follow-up questions drive question generation by identifying charts, numerical values, or experimental indicators. Type reinforcement rules strengthen the priority or enforceability of a certain type of follow-up question generation; for example, when hypothetical statements appear in the literature, counterfactual questions are prioritized. The clarification of mechanisms and rules such as matching conditions, generation mechanisms, and type reinforcement rules makes the question-answer chain construction more accurate and logically sound.
[0053] This process establishes a matching relationship between the content characteristics of the scientific literature to be questioned, the initial question type, and the answer content, and the corresponding matching conditions, generation mechanisms, and type reinforcement rules. This constructs a question-and-answer type library, mapping the semantic structure information of the scientific literature to be questioned, the starting point characteristics of the questions, and the semantic elements of the answers to the control conditions required for the three types of follow-up questions, and systematically organizing this into a structured question-and-answer type library. For example, when the initial question contains process-oriented descriptive words such as "steps," "process," or "next," it is preferentially routed to progressive follow-up questions; when the answer contains hypothetical expressions such as "if...then...", counterfactual follow-up questions are automatically triggered; if the literature contains structural information such as charts, graphs, or numerical analysis, an evidence-based follow-up question generation path is enforced. By constructing this question-and-answer type library that combines semantic driving, conditional response, and control rules, the most suitable follow-up question structure can be flexibly invoked when facing different literature scenarios.
[0054] Furthermore, this application also includes: calling a predefined template for matching follow-up question types based on the generation mechanism and type enhancement rules; filling the predefined template with entity, relation, and numerical parameters from the structured literature data; and semantically enhancing the output of the predefined template through a large language model to generate follow-up questions that conform to scientific research logic.
[0055] Specifically, based on the generation mechanism and type-enhancing rules, predefined templates for matching follow-up question types are invoked. According to the identified follow-up question type and its corresponding generation logic, a suitable template is selected as the basic framework for generating the question. The generation mechanism refers to the semantic path and logical order followed by different follow-up question types during the generation process. For example, progressive questions typically delve deeper into the details of existing answers, while counterfactual questions rely on changes in the underlying assumptions. Type-enhancing rules are supplementary constraints set to highlight the semantic features of various follow-up questions, such as forcibly introducing the "what if..." assumption structure in counterfactual follow-up questions. Predefined templates are pre-designed question construction frameworks. When invoking a predefined template, matching is performed based on the follow-up question type, ensuring logical consistency and linguistic standardization in the generated question.
[0056] The process involves filling entities, relationships, and numerical parameters from structured literature data into predefined templates, embedding template variables, and generating specific questions. Structured literature data typically originates from literature analysis results and includes chapter tags, variable definitions, parameter values, entity names, and their logical relationships. For example, if the template is "What are the differences between this method and [comparison method] in [parameter name]?", it will be filled with "convolutional neural network" and "learning rate" as entities and parameters, thus generating a question with specific context. This ensures a close correspondence between the follow-up questions and the literature content, improving the professionalism and accuracy of the questions.
[0057] After completing the basic template filling, a pre-trained large language model is invoked to semantically enhance the output of the predefined template, generating follow-up questions that conform to scientific research logic. The large language model further processes and optimizes the question sentences to ensure that the final generated questions not only conform to grammatical rules but also have rigorous logic and conform to the style of scientific research language. The semantic enhancement process typically includes operations such as sentence structure adjustment, expression polishing, logical completion, and context integration. For example, the question generated by the original template might be "Is method X better than method Y on Z?", which, after optimization by the large language model, can be transformed into "Compared with method Y, what performance advantages does method X exhibit under the adjustment of parameter Z?", thus better conforming to scientific research expression habits.
[0058] Furthermore, this application also includes: locating relevant content in the literature based on the type and type of follow-up question; extracting core keywords and core relationships based on the relevant content in the literature, and semantically expanding and adjusting the core keywords and core relationships using natural language to generate a reference answer; using the reference answer to determine the similarity of the follow-up answer content, obtaining a determination result, and filtering follow-up questions based on the determination result; and performing re-matching on the filtered follow-up question types to generate multiple rounds of questions.
[0059] Specifically, based on the type and nature of the follow-up question, relevant content in the literature is located. That is, after obtaining the specific follow-up question type and its corresponding question, the system retrospectively searches the original scientific literature for paragraphs or information locations semantically closely related to that follow-up question. The type of follow-up question determines the retrieval strategy. For example, if it is an evidence-based question, paragraphs containing charts, data, or citations are prioritized for searching; if it is a counterfactual question, text containing conditional sentences or variable descriptions is preferred, thus ensuring that the most relevant information is extracted from the original text for subsequent analysis and avoiding answers that deviate from the original context.
[0060] Based on the relevant content of the literature, core keywords and core relationships are extracted. Then, natural language processing is used to semantically expand and adjust these core keywords and relationships to generate a reference answer. This involves extracting key terms, variables, experimental subjects, and logical relationships from the relevant literature content. For example, when describing the performance of a method, keywords might be "accuracy," "dataset name," and "experimental results," while the core relationship might be "how much performance was improved under what conditions." Subsequently, a natural language processing model is used to convert this information into a complete answer sentence that conforms to scientific language norms. Semantic expansion refers to supplementing and integrating the original literature information, for example, expanding "X improved Y" to "On dataset Z, method X improved the accuracy metric by Y percentage points compared to the baseline method."
[0061] The automatically generated reference answer is semantically matched with the existing follow-up answers, and a similarity calculation model is used to determine whether the two are consistent or close. If the similarity is low, it indicates that the current follow-up question may not have a clear answer or is irrelevant; in this case, the follow-up question is marked as low-quality or redundant and discarded or adjusted. Subsequently, the next suitable follow-up question type is re-matched from the remaining set of follow-up question types to generate a new round of questions, achieving multi-round extension of the follow-up question chain, thereby effectively improving the coherence and relevance of the question chain.
[0062] Furthermore, this application also includes: establishing a question link graph by using the initial question as the root node, follow-up questions and multi-round questions as child nodes, marking edge relationships through semantic similarity calculation and a logical relationship classifier, and annotating source text location information for cross-paragraph and document nodes; based on the question link graph, performing question content correlation and response judgment analysis on the answer content of each question to obtain a heat map of document answer evaluation; based on the heat map of document answer evaluation, performing coverage depth analysis through the document content features to obtain the distribution of missing document locations; and based on the evaluation target of document question-answer coverage depth, combining the heat map to perform question compensation on the distribution of missing document locations.
[0063] Specifically, a question-and-answer chain is represented as a graph structure, with the initial question at the top, called the root node. The first and subsequent follow-up questions centered around the initial question are its subordinate nodes, called child nodes. A semantic similarity algorithm is used to determine the content relevance between each pair of questions, and a logical relationship classifier is used to establish directed connections between questions, such as labeling edge attributes like "cause and effect," "contrast," and "supplement." Furthermore, if a question and answer originate from different paragraphs, their specific sources are recorded in the graph, i.e., the source text location information, thus forming a visualized and structured question chain graph, establishing a question link graph.
[0064] After establishing a complete question-and-answer chain graph, the matching degree and contextual relationship between each question and its answer are further analyzed. By calculating the semantic consistency between the question content and the answer content, and judging whether the answer is complete and whether it answers the core intent of the question, each question node can be scored. The scores are then mapped to the original document paragraphs to form a heat map, which can intuitively display the document sections with high-quality and comprehensive answers, as well as the document sections with gaps or weak response areas.
[0065] Based on the low-coverage areas shown in the heatmap, and combined with the structural information and semantic features of the documents themselves, it is determined whether the current question-and-answer chain adequately covers the core content of the documents. Coverage depth analysis not only examines the number of questions and answers, but also includes indicators such as whether the questions and answers hit key paragraphs and whether they involve core methods or important variables. Finally, a set of "under-covered" document fragment locations is output, i.e., the distribution of missing document locations, to guide subsequent question supplementation.
[0066] After identifying coverage blind spots, new questions are automatically generated to fill the gaps in the existing question-and-answer chain. The compensation process comprehensively considers factors such as the position of paragraphs with lower scores in the heat map, their hierarchical position in the document structure, and related content themes, thereby proposing targeted questions. For example, it may add follow-up questions on the boundary conditions of a certain algorithm or supplement the analysis of the reasons for the failure of a certain experiment. This can further improve the question-and-answer chain and enhance the dataset's comprehensive coverage of scientific literature and its ability to express diverse ideas.
[0067] Furthermore, this application also includes: constructing a sample library through question data chains to analyze the topic types of various scientific and technological documents, constructing a question link reference library for each topic type, and the question link reference library is used for preliminary recommendation of question-answer data chains for different topic types in the later stage.
[0068] Specifically, a question data chain refers to a complete chain consisting of an initial question and its corresponding multiple rounds of follow-up questions and answers, reflecting the understanding and discussion of different levels of literature. By integrating question chain data automatically generated from a large number of scientific and technological documents into a sample library, the various scientific and technological documents are categorized by topic. Through accumulating question data chains, based on the terminology, structural features, and semantic clues involved in the question-and-answer content, the specific topic type of the document can be identified and labeled, such as biomedicine, engineering technology, and computer technology.
[0069] Based on the question chains already categorized by topic in the sample database, representative, comprehensive, and logically rigorous question-and-answer chain structures are extracted and compiled into a reference template library as a question chain reference library. This library is then categorized by topic, ensuring that each type of literature has a corresponding high-quality question chain as a reference model, facilitating the rapid processing of similar subsequent literature. The question data chains in the question chain reference library not only include questions and their logical relationships but may also involve typical follow-up questioning patterns, key indicator focus points, and common causal paths.
[0070] When new scientific and technological literature is received, the system can first use the corresponding reference question chain template based on its subject type (such as medicine, chemistry, materials science, etc.) as a starting point or guiding principle for constructing new question-and-answer data chains. This improves the relevance and professionalism of the generated question-and-answer chains. For example, when processing medical papers, priority can be given to generating question types such as evidence chains around clinical conclusions, comparative chains of treatment plans, and sensitive chains of experimental sample parameters. When processing materials science literature, the system tends to generate chains of development trends and comparative chains of chemical raw material selection.
[0071] In summary, the multi-type question-and-answer data chain generation method for scientific literature question-and-answer systems provided in this application has the following technical effects: by achieving the technical goals of automatic generation of question chains driven by document structure and semantic features, adaptive matching of follow-up question types, and linkage control of question chain evaluation and compensation optimization in question-and-answer data chain construction, it enhances the ability of scientific literature question-and-answer systems in terms of multi-type semantic understanding, multi-round logical modeling, and intelligent generation of data integrity, and provides structured, high-quality, and highly adaptable question sample support for academic AI.
[0072] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0073] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application also intends to include such modifications and variations.
Claims
1. A method for generating multi-type question-and-answer data chains for scientific and technological literature question-and-answer systems, characterized in that, include: The pre-trained language model is invoked to perform content and structural hierarchical analysis on the scientific and technological documents to be answered, obtain the document content features, and generate initial questions based on the document content features; Using the document content features, initial question type, and answer content as indexes, a set of follow-up question types is obtained by matching in the question-and-answer type library; Dynamically select at least one type of follow-up question from the set of follow-up question types, construct follow-up questions, and obtain the content of follow-up question answers; Based on the determination results of the follow-up question type and the content of the follow-up question answer, a rematch is performed on the set of follow-up question types to generate multiple rounds of questions; The question chain of initial question - follow-up question - multiple round questions is evaluated based on the characteristics of the document content. The question chain is compensated and optimized according to the evaluation results. Question data chain is generated by using the question and answer and parsing data corresponding to the question chain that meets the evaluation objectives.
2. The method for generating multi-type question-and-answer data chains for a scientific and technological literature question-and-answer system according to claim 1, characterized in that, The obtained document content features include: The scientific and technological documents to be questioned are subjected to structured analysis to extract chapter elements, content types and distribution, and a knowledge relationship graph is constructed. Based on the knowledge relationship graph, semantic hierarchy recognition is performed on the content of each node, structural hierarchy distribution and semantic relationships are established, and added to the knowledge relationship graph; Based on the knowledge relationship graph, the core content is analyzed to determine the document type and content characteristics, and the content characteristics of the document are obtained.
3. The method for generating multi-type question-and-answer data chains for a scientific and technological literature question-and-answer system according to claim 2, characterized in that, The document types include: review, methodological, experimental, and theoretical.
4. The method for generating multi-type question-and-answer data chains for a scientific and technological literature question-and-answer system according to claim 3, characterized in that, Initial questions are generated based on the characteristics of the document content, including: Establish a mapping relationship between the document type, content features and question-and-answer type, wherein the mapping relationship includes a first-level mapping corresponding to the document type and a second-level mapping corresponding to the content features; Based on the characteristics of the document content, first-level mapping and second-level mapping are performed in the mapping relationship to generate the initial question.
5. The method for generating multi-type question-and-answer data chains for a scientific and technological literature question-and-answer system according to claim 4, characterized in that, Establish the mapping relationship between the document type, content features, and question-and-answer type, including: Based on historical question-and-answer samples of various document types, a first-level mapping between each document type and question-and-answer type is established. Among them, the question-and-answer types corresponding to review document types include comparison chain, controversy chain, and development path chain; the question-and-answer types corresponding to methodological document types include causal chain, application chain, and algorithm optimization chain; the question-and-answer types corresponding to experimental document types include evidence chain, variable control chain, and parameter sensitivity chain; and the question-and-answer types corresponding to theoretical document types include definition chain, derivation chain, and abstract relationship chain. Based on the analysis of historical question-and-answer samples, the response relationship between content features and question-and-answer types is analyzed, and a two-level mapping between content features and question-and-answer types is established. Based on the first-level and second-level mappings, the mapping relationship between the document type, content features, and question-and-answer type is obtained.
6. The method for generating multi-type question-and-answer data chains for a scientific and technological literature question-and-answer system according to claim 5, characterized in that, Using the document content characteristics, initial question type, and answer content as indexes, a match is performed in the question-and-answer type database, prior to which includes: Establish question-and-answer types, including progressive, counterfactual, and evidentiary types; The matching conditions, generation mechanism, and type enhancement rules for each type of follow-up question and answer are analyzed separately. Establish the corresponding matching relationships between the document content features, initial question types and answer content, and the matching conditions, generation mechanisms and type reinforcement rules, and construct the question-answer type library.
7. The method for generating multi-type question-and-answer data chains for a scientific and technological literature question-and-answer system according to claim 6, characterized in that, Dynamically select at least one follow-up question type from the set of follow-up question types to construct follow-up questions, including: Based on the generation mechanism, type-enhancing rules invoke predefined templates for type matching; Populate entities, relations, and numerical parameters from the structured data of the literature into the predefined template; By semantically enhancing the output of predefined templates using a large language model, follow-up questions that conform to scientific research logic are generated.
8. The method for generating multi-type question-and-answer data chains for a scientific and technological literature question-and-answer system according to claim 1, characterized in that, Based on the determination results of the follow-up question type and the content of the follow-up question answer, a re-matching is performed on the set of follow-up question types to generate multiple rounds of questions, including: Based on the type and nature of the follow-up questions, locate relevant content in the literature; Based on the relevant content of the literature, core keywords and core relationships are extracted, and semantic expansion and adjustment are performed using natural language based on the core keywords and core relationships to generate reference answers; The similarity of the follow-up question answers is determined using the reference answer, and the result is obtained. Follow-up questions are filtered based on the result, and the filtered follow-up question types are then re-matched to generate multiple rounds of questions.
9. The method for generating multi-type question-and-answer data chains for a scientific and technological literature question-and-answer system according to claim 8, characterized in that, The question chain of initial question - follow-up question - multiple rounds of questions is evaluated based on the characteristics of the document content. Based on the evaluation results, the question chain is compensated and optimized, including: Using the initial question as the root node, follow-up questions and multi-round questions as child nodes, the edge relationships are marked by semantic similarity calculation and logical relationship classifier, and the source text location information is marked for cross-paragraph and document nodes to establish a question link graph; Based on the aforementioned question link diagram, the content correlation and response judgment analysis of the answers to each question are performed to obtain a heat map of the literature response evaluation. Based on the heat map of the literature response evaluation, a coverage depth analysis is performed using the literature content characteristics to obtain the location distribution of missing literature. Based on the evaluation objective of the depth of literature question-and-answer coverage, and combined with the heat map, the distribution of the missing literature is compensated for.
10. The method for generating multi-type question-and-answer data chains for a scientific and technological literature question-and-answer system according to claim 1, characterized in that, Also includes: By constructing a sample library through question data chains, the topic types of various scientific and technological documents are analyzed, and a question link reference library for each topic type is constructed. The question link reference library is used for the initial recommendation of question and answer data chains for different topic types in the later stage.
Citation Information
Patent Citations
Question and answer pair data mining method and device and electronic equipment
CN109657038A
Professional field-oriented question and answer knowledge database construction method and device, equipment and medium
CN119378672A