Scientific discovery intelligent system embedded with scientific and technological text reverse reasoning mechanism

Through the scientific discovery intelligent system embedded in the scientific text reverse reasoning mechanism, the problem that the analysis results in the existing technology are limited by experts' knowledge reserves and biases, and the diversity and accuracy of the direction of scientific discovery are improved, and subjective dependence is reduced.

CN120146190AActive Publication Date: 2025-06-13DOCUMENT & INFORMATION CENT OF CHINESE ACAD OF SCI

Patent Information

Application Number
CN202510215561.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

The analysis results in the prior art are limited by experts' knowledge reserves and biases, and the large model mainly relies on basic knowledge for positive derivation, lacking the fusion of fine-grained knowledge, resulting in strong subjective dependence and a single reasoning path.

Method used

Provides a scientific discovery intelligent system that embeds the reverse reasoning mechanism of scientific text, including planning modules, storage modules, tool modules and execution modules, and uses the reverse reasoning mechanism to reverse future research directions from existing research limitations to achieve full-process support for the scientific discovery process.

Benefits of technology

Through the reverse reasoning mechanism, the system can reverse the possible future research directions from existing research limitations, reduce subjective dependence, improve the diversity and accuracy of inference paths, and ensure the interpretability and reliability of the direction of scientific discovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146190A_ABST
    Figure CN120146190A_ABST
Patent Text Reader

Abstract

The invention provides a scientific discovery intelligent system embedded with a scientific and technological text reverse reasoning mechanism, and relates to the technical field of data processing, and the system comprises a planning module which is used for extracting and clustering research problems, research methods and research limitations, and generating structured information; the storage module is used for forming a management knowledge base; the tool module is used for providing operation tools for the planning module, the storage module and the execution module; and the execution module is used for reasoning and optimizing a scientific discovery direction by combining a reverse reasoning mechanism based on the management knowledge base. The technical problems that in the prior art, an analysis result is limited by knowledge reserve and prejudice of experts, a current large model mainly depends on basic knowledge to conduct forward derivation, fusion of fine-grained knowledge is lacked, subjective dependence is high, and a reasoning path is single exist, the possible future research direction can be reversely deduced from existing research limitations, and the research efficiency is improved. And full-flow support of the scientific discovery process is realized, so that the accuracy of the reasoning direction is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to an intelligent scientific discovery system incorporating a reverse reasoning mechanism for scientific texts. Background Art

[0002] In the traditional approach, researchers mainly deduce future research directions through methods such as manual literature analysis, statistical analysis, and expert discussions. Specifically, researchers first need to read a large number of literatures, sort and classify them, and sort out the evolution law of research hotspots in chronological order, record and summarize the research methods, contents, and conclusions of each literature, and find research gaps through comparison. At the same time, statistical analysis methods are also used to conduct multi-dimensional statistics on the number of literature publications, perform word frequency analysis of keywords, construct citation networks, and use bibliometric methods for trend analysis. On this basis, by organizing expert seminars, Delphi methods, etc., the empirical judgments of domain experts are gathered to form a consensus prediction of future research directions. Finally, researchers will combine the literature analysis results with expert opinions, consider the laws of disciplinary development and social needs, evaluate the technical feasibility and research conditions, and finally form a recommendation report on future research directions.

[0003] When large models are applied to deduce future research directions, it is mainly achieved through a training method based on large-scale corpora. By pre-training on a vast amount of scientific and technological literature, large models have mastered domain knowledge and language understanding capabilities. When inputting research questions in a specific domain, they can provide relevant research direction suggestions based on the knowledge in their training data. Specifically, large models will perform semantic understanding on the input research topic, then retrieve relevant research progress in their knowledge bases, identify research hotspots and development trends by analyzing the relevance and temporal relationships between literatures. At the same time, large models can also speculate on possible research gaps and innovation points through understanding research methods, experimental results, and theoretical frameworks. However, this method mainly relies on the historical data used during model training, and the reasoning process is relatively closed, lacking real-time grasp of the latest research dynamics, and there are also certain limitations in the interpretability and reliability of the deduced results.

[0004] In the existing technology, there are technical problems that the analysis results are limited by the knowledge reserve and biases of experts, and current large models mainly rely on basic knowledge for forward deduction, lacking the integration of fine-grained knowledge, resulting in strong subjective dependence and a single reasoning path. Summary of the Invention

[0005] The present application provides an intelligent scientific discovery system embedded with a reverse reasoning mechanism for scientific and technological texts, which is used to solve the technical problems existing in the prior art. Due to the fact that the analysis results are limited by the knowledge reserve and bias of experts, and the current large models mainly rely on basic knowledge for forward derivation, lacking the integration of fine-grained knowledge, resulting in strong subjective dependence and a single reasoning path.

[0006] In view of the above problems, the present application provides an intelligent scientific discovery system embedded with a reverse reasoning mechanism for scientific and technological texts, including: a planning module, which is used to call the knowledge extraction and clustering tools of the tool module to plan the input scientific and technological literature set, extract and cluster research questions, research methods, and research limitations, and generate structured information; a storage module, which is used to hierarchically manage the structured information extracted by the planning module to form a management knowledge base; a tool module, which is used to provide operation tools for the planning module, the storage module, and the execution module; an execution module, which is used to infer and optimize the scientific discovery direction based on the management knowledge base in combination with the reverse reasoning mechanism.

[0007] One or more technical solutions provided in the present application have at least the following technical effects or advantages:

[0008] The intelligent scientific discovery system embedded with a reverse reasoning mechanism for scientific and technological texts provided by the present application solves the technical problems existing in the prior art. Due to the fact that the analysis results are limited by the knowledge reserve and bias of experts, and the current large models mainly rely on basic knowledge for forward derivation, lacking the integration of fine-grained knowledge, resulting in strong subjective dependence and a single reasoning path. By planning the input scientific and technological literature set through the planning module, the extraction and clustering processes of research questions, research methods, and research limitations are respectively obtained. The scientific and technological literature information and processing results are stored and managed by the storage module. With the help of the knowledge extraction, clustering analysis, and large model fine-tuning tools provided by the tool module, the execution module executes specific operation tasks. At the same time, by introducing the reverse reasoning mechanism, it is possible to reverse infer possible future research directions from the existing research limitations, realizing full-process support for the scientific discovery process, thereby ensuring the accuracy of the reasoning direction. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 It is a schematic structural diagram of the intelligent scientific discovery system embedded with a reverse reasoning mechanism for scientific and technological texts provided by the present application;

[0010] Figure 2 It is a schematic flow diagram of constructing structured information in the intelligent scientific discovery system embedded with a reverse reasoning mechanism for scientific and technological texts provided by the present application.

[0011] Description of the reference numerals: planning module 11, storage module 12, tool module 13, execution module 14. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0012] This application provides an intelligent scientific discovery system embedded with a reverse reasoning mechanism for scientific and technological texts, aiming to solve the technical problems in the prior art. Due to the analysis results being limited by the knowledge reserve and bias of experts, and the current large models mainly relying on basic knowledge for forward derivation, lacking the integration of fine-grained knowledge, the subjective dependence is strong and the reasoning path is single.

[0013] As Figure 1 shown, the embodiment of this application provides an intelligent scientific discovery system embedded with a reverse reasoning mechanism for scientific and technological texts, including:

[0014] A planning module 11, configured to call the knowledge extraction and clustering tool of the tool module 13, plan the input scientific and technological literature set, extract and cluster research questions, research methods, and research limitations, and generate structured information.

[0015] Specifically, the core function of the planning module 11 is to systematically plan the input scientific and technological literature set by calling the knowledge extraction and clustering tool in the tool module 13. Specifically, this process first includes identifying and extracting research questions, research methods, and research limitations in the literature. Research questions refer to the main issues or challenges discussed in the literature, usually topics that have not been solved or need further exploration in a certain field; research methods refer to the specific methods or technical means used in the literature to study and solve problems, which may involve experimental design, theoretical analysis, or data processing methods; research limitations refer to the deficiencies or limitations of the research pointed out in the literature, usually reflecting the inevitable limitations in the research process, such as insufficient sample size, method limitations, etc.

[0016] To effectively organize this information, the clustering tool classifies and clusters the extracted research questions, methods, and limitations. The clustering process automatically groups similar or related research questions, methods, and limitations into the same group through an algorithm, thereby constructing an ordered and hierarchical structure. This not only helps to identify the core issues in the literature but also effectively extracts the relationships between different research methods, as well as the commonalities and differences of each research limitation. For example, when planning the scientific and technological literature in a certain field, if the literature set discusses issues related to "data security" and "privacy protection", the planning module 11 will identify and cluster these issues into one group; if there are different research methods, such as "encryption technology" or "anonymization technology", these will be clustered into another group and establish a connection with the research question group. Finally, the generated structured information can clearly show the research status in this field, including key issues, methods, and limitations, facilitating subsequent in-depth analysis and optimization. It can not only help to construct a comprehensive knowledge framework but also provide accurate basic data for the subsequent reverse reasoning mechanism.

[0017] The storage module 12 is used to hierarchically manage the structured information extracted by the planning module to form a management knowledge base.

[0018] Specifically, the main task of the storage module 12 is to hierarchically manage the structured information extracted by the planning module 11 and finally form an efficient and sustainable update management knowledge base. In this process, the storage module 12 organizes information such as research questions, research methods, and research limitations in a hierarchical manner to ensure that the access and update of knowledge become more orderly and efficient. Specifically, multiple sub-libraries are constructed according to different levels and categories of structured information. For example, for the "research question" part, research questions in different fields or disciplines are classified and divided into appropriate hierarchical structures based on the descriptions in the literature. Similarly, the "research method" part is classified and managed according to characteristics such as the type of method, application scenario, or research goal. The "research limitation" part is hierarchically managed according to the type of limitation or influencing factors. Taking actual operation as an example, assume that the literature in a certain field involves multiple different research questions, such as "data privacy protection", "fairness of intelligent systems", and "interpretability of machine learning". According to the relevance and research background of these questions, they are stored in different sub-libraries respectively to obtain a management knowledge base, which is managed in the form of a hierarchical structure, ensuring that when a certain type of question needs to be retrieved, relevant data can be quickly located without the need for a long search from the entire literature library.

[0019] During the storage process, the structured information will not only be a static document or data set, but a dynamic and updatable knowledge base at any time. New research literature can be continuously input, and the storage module 12 will automatically adjust and optimize the existing hierarchical structure according to the new data, thereby ensuring the accuracy and timeliness of the knowledge base. Through this hierarchical management, the storage module 12 not only enhances the maintainability of knowledge, but also provides convenient data support for the subsequent execution module 14, enabling the system to reason and optimize based on the latest research results.

[0020] The tool module 13 is used to provide operation tools for the planning module 11, the storage module 12, and the execution module 14.

[0021] Specifically, the tool module 13 is a crucial part of the entire system. Its main function is to provide necessary operation tools for the planning module 11, the storage module 12, and the execution module 14, ensuring efficient collaborative work among the modules. The tool module 13 includes a variety of functional tools, and the specific tool forms and functions will vary according to the needs of each module. For example, there are knowledge extraction tools and clustering analysis tools. When processing the input scientific and technological literature, the knowledge extraction tools are used to extract key research questions, research methods, research limitations, and other information from a large number of literatures. After preliminary screening, the clustering analysis tools will cluster the extracted data, classify this information based on its similarity and relevance, and form preliminary structured data. This process ensures that the planning module can efficiently sort out and plan a large number of scientific and technological literatures and generate valuable knowledge networks.

[0022] Secondly, the tool module 13 provides support for the storage module 12 to help it efficiently manage and organize structured information. The hierarchical management and knowledge base construction in the storage module 12 rely on the data mapping tools in the tool module 13. The data mapping tools convert data such as research questions, methods, and limitations in the structured information into a format that is convenient for storage and query, and organize it according to a predetermined hierarchical structure. For example, the classification of research questions may require the mapping tool to automatically assign them to the corresponding knowledge base, while research methods and limitations are classified and stored according to different feature vectors. This process effectively reduces the need for manual intervention and improves storage and retrieval efficiency.

[0023] Finally, the tool module 13 also provides necessary operation tools for the execution module 14. Especially in the process of large model fine-tuning and reverse reasoning, the fine-tuning tools and reasoning tools of the tool module 13 play a crucial role. The large model fine-tuning tools enable the execution module 14 to adjust the reasoning model through the reverse reasoning mechanism according to the structured information extracted from the storage module 12 and optimize the scientific discovery direction.

[0024] The execution module 14 is used to reason and optimize the scientific discovery direction based on the management knowledge base in combination with the reverse reasoning mechanism.

[0025] Specifically, the task of the execution module 14 is to deduce and optimize the direction of scientific discovery based on the management knowledge base through a reverse reasoning mechanism. This module not only relies on the previously stored and planned knowledge base but also can self-adjust and optimize the reasoning process, thereby providing precise directional guidance for scientific research. During the execution process, first, information such as structured research questions, research methods, and research limitations is obtained from the management knowledge base. Based on this stored information, an analysis is carried out, combined with the reverse reasoning mechanism. By analyzing the relationship between research questions and methods, the most promising direction of scientific discovery is speculated. The core of the reverse reasoning mechanism is to reverse-deduce possible causes or research paths from known results. Specifically, the reverse reasoning mechanism first analyzes the current known scientific research limitations based on the research question and method network in the knowledge base, identifies which research questions have not been fully solved, and which methods can be further optimized. It can not only discover existing scientific discovery directions but also provide new ideas for unsolved research questions and propose innovative research methods. Taking a specific scientific research field as an example, if the current research direction is the high-efficiency charging technology of new energy batteries, through the research question network in the management knowledge base, the limitations in the current charging technology are identified, such as poor low-temperature performance and slow charging speed. Then, combined with the reverse reasoning mechanism, a charging optimization method based on new electrolyte materials may be deduced, thus proposing a new and potential scientific discovery direction.

[0026] Furthermore, as shown in the appendix Figure 2 The planning module in the intelligent scientific discovery system embedded with the reverse reasoning mechanism for scientific and technological texts includes: a research question extraction unit, which is used to extract research questions from the scientific and technological literature set, classify and identify problem-related sentences, and perform dynamic clustering to generate a research question knowledge network; a research method extraction unit, which is used to extract research method-related sentences from the scientific and technological literature set through multi-stage training and perform incremental clustering to generate a problem-oriented research method knowledge network; a research limitation extraction unit, which is used to extract limitation information related to research questions and methods from the scientific and technological literature set, perform incremental clustering on the limitation information for each question-method combination to generate a scientific research knowledge discovery network; and an information construction unit, which is used to construct the structured information with the research question knowledge network, the research method knowledge network, and the scientific research knowledge discovery network.

[0027] Specifically, research questions are extracted from the input collection of scientific and technological literature. To this end, the planning module uses classification and recognition techniques to identify sentences related to research questions. For example, from the full texts of 500 scientific and technological literature input by the user, key sentences related to research questions are selected through manual annotation, including problem background sentences, problem description sentences, problem significance sentences, etc. The initial corpus contains 3,000 annotated data, and positive and negative samples are configured in a 1:1 ratio. In the preprocessing stage, the text is cleaned and tokenized, Chinese word segmentation is performed, and stop words and special characters are removed. This batch of high-quality manually annotated corpus is used to preliminarily train and fine-tune the DeBERTa model to construct a classification model capable of identifying sentences related to research questions. The initial model achieves an accuracy of 85% on the validation set. Subsequently, the preliminarily trained model is used to semi-automatically annotate the input scientific and technological literature, and the sentences related to the identified research questions are manually reviewed and verified, and iterative optimization is continuously carried out. Finally, 15,000 high-quality research question fine-tuning corpus is obtained. These corpus are used to deeply fine-tune the DeBERT model. With the same training parameter configuration, the model performance is further improved to 92% accuracy. The optimized model is used to automatically extract sentences related to research questions from the scientific and technological literature collection and perform classification and dynamic clustering. Specifically, the incremental clustering method is adopted. First, the sentences related to the extracted research questions are processed, and natural classification is performed according to topic similarity. Accurate class labels and key feature descriptions are generated for each category to achieve automatic clustering. Each time a new research question is received, the semantic similarity between the new question and the existing categories is calculated (the similarity threshold is set to 0.85). For questions with similarity exceeding the threshold, they are directly classified into the corresponding category; for questions with similarity lower than the threshold, they are separately extracted and re-clustered to form new categories. In this way, the dynamic incremental clustering of research questions is realized, ensuring the natural evolution of categories and the continuous accumulation of knowledge. During the processing, detailed reasoning bases are generated for each classification decision, and the feature descriptions of each category are dynamically updated to maintain the structured organization of the clustering system, and finally an ever-expanding knowledge network of research questions is formed.

[0028] Next, through multi-stage training, relevant sentences about research methods are further extracted from the collection of scientific and technological literature. Exemplarily, first, from the full texts of 500 scientific and technological literature input by users, key sentences related to research methods are selected through manual annotation, including method description sentences, experimental design sentences, technical route sentences, etc. The initial corpus contains 4,000 annotated data, and positive and negative samples are configured in a 1:1 ratio. In the preprocessing stage, the text is cleaned and tokenized. Chinese word segmentation is performed using existing tokenization tools, stop words and special characters are removed, and the RoBERTa model is preliminarily trained and fine-tuned using this batch of high-quality manually annotated corpus to build a classification model capable of identifying sentences related to research methods. The preliminary model achieves an accuracy of 87% on the validation set; subsequently, the preliminarily trained model is used to semi-automatically annotate the full texts of the input scientific and technological literature, and the identified sentences related to research methods are manually reviewed and verified, and iterative optimization is continuously carried out. Finally, 18,000 high-quality research method fine-tuning corpus are obtained, and the RoBERTa model is deeply fine-tuned using these corpus. With the same training parameter configuration, the model performance is further improved to 94% accuracy, and the optimized model is used to automatically extract research method knowledge from the collection of scientific and technological literature, realizing the accurate extraction from literature to method. Further, an incremental clustering method is adopted for each problem category. First, the initial 300 research method corpus input under each problem category is processed, and natural classification is carried out based on method type and application scenario to generate accurate category labels and method feature descriptions for each category. Through in-depth semantic understanding, the 300 corpus under each problem category are initially clustered into 10-15 main method categories, and each category generates detailed method feature descriptions and application boundaries by the model. In the incremental clustering stage, each time 300 newly added research methods under the problem category are received, the semantic similarity between the new methods and the existing categories is calculated (the similarity threshold is set to 0.82). For methods with similarity exceeding the threshold, they are directly classified into the corresponding category; for methods with similarity lower than the threshold, they are separately extracted and re-clustered to form new categories. In this way, the dynamic incremental clustering of research methods under each problem category is realized, ensuring the natural evolution of method categories and the continuous accumulation of method knowledge. During the processing, detailed reasoning bases are generated for each classification decision, and at the same time, the method feature descriptions of each category are dynamically updated to maintain the structured organization of the problem-method system, and finally a problem-oriented research method knowledge network is formed, which can not only provide a solid knowledge foundation for subsequent scientific discovery reasoning, but also help scientific researchers quickly master the research progress and difficulties in specific fields in a systematic way.

[0029] Furthermore, incremental clustering of limited information is performed to generate a scientific research knowledge discovery network. Specifically, for the scientific and technological literature under each problem-method combination, research limitations are extracted using a multi-stage training approach. First, from the literature corresponding to the problem-method (about 15 - 20 pieces of literature for each combination), key sentences related to research limitations are selected through manual annotation, including limitation description sentences, deficiency explanation sentences, improvement suggestion sentences, etc. The initial corpus contains 2,000 pieces of annotated data, and positive and negative samples are configured in a 1:1 ratio. In the preprocessing stage, the text is cleaned and tokenized, and Chinese word segmentation is performed using existing tokenization tools to remove stop words and special characters. This batch of high-quality manually annotated corpus is used to preliminarily train and fine-tune the XLNet model to build a classification model capable of identifying sentences related to research limitations. The preliminary model achieves an accuracy of 86% on the validation set; subsequently, the preliminarily trained model is used to semi-automatically annotate the literature under each problem-method combination, and the identified sentences related to research limitations are manually reviewed and verified, and iterative optimization is continuously carried out. Finally, 12,000 pieces of high-quality research limitation fine-tuning corpus are obtained. Using these corpora to deeply fine-tune the XLNet model with the same training parameter configuration, the model performance is further improved to 93% accuracy, and the optimized model is used to automatically extract research limitation knowledge from the literature corresponding to each problem-method combination, realizing the accurate extraction from the problem-method combination to the limitation. For each problem-method combination, incremental clustering is used for clustering to obtain a scientific research knowledge discovery network, and the structured information is constructed with the research problem knowledge network, the research method knowledge network, and the scientific research knowledge discovery network. The finally generated structured information framework can clearly mark which research problems are the most urgent, which methods are the most effective, and which limitations need to be broken through the most, thus greatly improving the efficiency and accuracy of scientific discovery.

[0030] Furthermore, the storage module 12 in the scientific discovery intelligent system embedded with the reverse reasoning mechanism of scientific and technological texts includes: a research problem knowledge base construction unit, which is used to extract and map problem descriptions, problem categories, and problem feature vectors based on the research problem knowledge network in the structured information, and establish a research problem knowledge base; a research method knowledge base construction unit, which is used to extract and map method descriptions, method categories, and method feature vectors based on the research method knowledge network in the structured information, and establish a research method knowledge base; a research limitation knowledge base construction unit, which is used to extract and map limitation descriptions, limitation categories, and limitation feature vectors based on the scientific research knowledge discovery network in the structured information, and establish a research limitation knowledge base; a knowledge base integration unit, which hierarchically manages the research problem knowledge base, the research method knowledge base, and the research limitation knowledge base to form the management knowledge base.

[0031] Specifically, based on the research question knowledge network in the structured information, each research question is described in detail, and the category of the question and the feature vector of the question are extracted, including information such as the question description, question category, and question feature vector. The question description is a concise elaboration of each research question, covering the core content and background of the question; the question category classifies the research questions by field, topic, or technical method, etc., facilitating subsequent search and comparison. The question feature vector is to transform each research question into a mathematical vector through technical means. Next, based on the research method knowledge network in the structured information, each research method is described, classified, and its feature vector is extracted, including information such as method description, method category, and method feature vector. The method description includes a detailed explanation of each research method, including its scope of application, operation process, and scientific principles, etc. The method category classifies these methods according to the research field or technical direction, such as "experimental method", "simulation method", or "theoretical analysis method". The method feature vector quantifies these research methods into a set of numerical values. By comparing the similarity and effectiveness between research methods, the best method suitable for the current research question can be efficiently found.

[0032] Finally, based on the scientific research knowledge discovery network, the limitations related to research questions and research methods are described. For each question-method combination, hierarchical storage management of research limitations is carried out, and a dynamically updated research limitation knowledge base corresponding to the question-method is established, including information such as limitation description, limitation category, and limitation feature vector.

[0033] Through the above steps, three hierarchical knowledge bases are finally formed: the research question knowledge base, the research method knowledge base, and the research limitation knowledge base. Each knowledge base contains the description, classification, and feature vector of questions, methods, and limitations, facilitating the subsequent efficient management and retrieval of knowledge.

[0034] Furthermore, the tools in the tool module 13 include a knowledge extraction tool, a clustering analysis tool, and a large model fine-tuning tool.

[0035] Specifically, the knowledge extraction tool is a tool used to automatically extract key information from the input scientific and technological literature collection, and can identify and extract key information in sentences, paragraphs, or paragraphs related to research questions, methods, and limitations. Through natural language processing (NLP) technology, the knowledge extraction tool can identify important elements such as terms, professional vocabulary, and research conclusions in scientific and technological literature. The clustering analysis tool plays a role based on knowledge extraction. Its main task is to cluster the information extracted from scientific and technological literature, group it according to characteristics such as theme, method, or limitation. The clustering analysis tool uses various machine learning and data mining algorithms (such as K-means, hierarchical clustering, etc.) to identify the similarities of different research directions, research questions, and their associated methods. Finally, the large model fine-tuning tool is a key tool for improving the system's reasoning ability. Through large model fine-tuning, it can perform fine-tuning in a specific field based on the existing large pre-trained model to better adapt to the task requirements.

[0036] Furthermore, the execution module 14 in the scientific discovery intelligent system embedded with the reverse reasoning mechanism of scientific and technological texts includes: a prompt word generation unit, which is used to generate scientific discovery prompt words based on the management knowledge base in the form of question-method-limit combinations, and construct a training data set for scientific discovery prompt words; a fine-tuning training unit, which is used to adopt the training data set and perform prompt word fine-tuning training based on the reverse reasoning mechanism through the large model fine-tuning tool to obtain an optimized scientific discovery direction.

[0037] Specifically, the generation of question-method-limit combinations is to organically combine the research questions, corresponding research methods, and relevant research limitation information extracted from the management knowledge base. These combinations not only cover the existing academic research framework but also provide specific reasoning guidance for scientific discovery. For example, if a question is "how to improve the charging efficiency of the battery?", the corresponding research method may be "optimize the charging algorithm", and the limitation information may include "the heat management problem during the charging process". By integrating these elements, a specific prompt word is generated to help the subsequent reasoning system focus on the relevant scientific direction. Once the prompt words are generated, a training data set for scientific discovery prompt words will be constructed based on these prompt words. The construction of the training data set not only includes the above question-method-limit combinations but also takes into account the context information of each prompt word to ensure that the data set reflects the complexity and diversity of scientific research questions in multiple dimensions. The accuracy and comprehensiveness of the training data set directly affect the effect of large model fine-tuning training, so the construction of the data set is a very crucial step.

[0038] Next, using this training dataset, prompt fine-tuning training based on the reverse reasoning mechanism is carried out through the large model fine-tuning tool. The role of the large model fine-tuning tool is to use the existing large-scale pre-trained model and, based on the existing model, through targeted dataset fine-tuning, make the model more adaptable to the needs of specific fields. During the fine-tuning process, the model performs reverse reasoning according to the generated scientific discovery prompts, combined with the input scientific and technological literature and the existing structured knowledge, and gradually adjusts the reasoning ability and weights of the model. Specifically, the reverse reasoning mechanism allows not only to derive the scientific discovery direction from the existing knowledge during the reasoning process, but also to dynamically adjust the reasoning path to adapt to new discoveries or proposed research questions. Finally, after repeated training and fine-tuning, an optimized scientific discovery direction is obtained, which can not only combine existing research, refine effective scientific research directions, but also reveal new research potential to a certain extent and promote scientific and technological innovation. Therefore, the scientific discovery prompts generated based on the management knowledge base and the training process of large model fine-tuning form a closed loop, and through repeated optimization, the accuracy and efficiency of the system in scientific discovery direction reasoning are continuously improved.

[0039] Furthermore, the scientific discovery intelligent system embedded with the reverse reasoning mechanism of scientific and technological texts further includes a cross-domain learning module. The cross-domain learning module is used for: a domain sample collection unit, which is used to collect literature annotation data in multiple different fields, identify the key research questions, methods and their limitations in each field, and obtain a structured sample database for each field; a domain focus project configuration unit, which is used to configure each adaptive loss function and each reasoning feature priority information for the scientific reasoning focus projects in each field based on the structured sample database for each field; a weight adjustment unit, which is used to automatically adjust the weights of the model when performing prompt fine-tuning training based on the reverse reasoning mechanism through the large model fine-tuning tool according to the domain type of the input scientific and technological literature set based on the respective adaptive loss functions and the respective reasoning feature priority information.

[0040] Specifically, in the cross - domain learning module, it is first necessary to collect labeled data of scientific and technological literature from multiple different scientific research fields, and use natural language processing (NLP) technology to identify the key research questions, research methods and their limitations in each field. The acquisition of labeled data usually relies on the combination of manual annotation and automated extraction to ensure the accuracy and coverage of the data. Through this process, a structured sample database covering multiple disciplines can be constructed. This database not only contains the core research directions of each field, but also covers the common methodologies and their limitations within the field, thus forming an interdisciplinary knowledge framework. After obtaining the structured sample database, for the scientific reasoning focus items in different fields, configure each adaptive loss function and the priority information of reasoning features. The role of the adaptive loss function is to dynamically adjust the penalty degree of the model for wrong predictions during the reasoning process according to the characteristics of different research fields. For example, in the field of medical image analysis, the accuracy of research methods is usually the core, so the loss function may give a higher penalty for wrong method identification. While in materials science research, the rationality of research limitations may be more important, and the system may assign a higher loss weight to wrong limitation predictions. In addition, the setting of the priority information of reasoning features can guide the system to weigh the weights of research questions, methods and limitations in different fields. For example, in artificial intelligence research, the innovation of new algorithms may be given priority, while in environmental science, the applicability and generalizability of methods may be more concerned.

[0041] Based on these adaptive loss functions and the priority information of reasoning features, when performing the prompt - tuning training of the large - model fine - tuning tool based on the reverse reasoning mechanism, the weights of the model can be automatically adjusted. When processing scientific and technological literature in different fields, the parameters of the neural network will be dynamically adjusted according to the characteristics of the field to make it more in line with the reasoning needs of the field. For example, in the field of energy science, the model may be more inclined to reason about new material optimization schemes, while in the field of biomedicine, the model may be more concerned with the improvement direction of experimental design. Such a weight - adjustment mechanism ensures that a high reasoning accuracy can be maintained in multi - disciplinary scenarios and can be flexibly adapted according to scientific research trends, improving the accuracy and applicability of scientific discoveries.

[0042] Furthermore, the execution module 14 in the intelligent scientific discovery system embedded with the reverse reasoning mechanism of scientific and technological texts includes: a similarity analysis unit, which is used to compare and analyze the first scientific discovery direction generated by the large model fine-tuning tool with the standard answer through a semantic similarity algorithm to obtain the first reasoning similarity; a similarity comparison unit, which is used to output the first scientific discovery direction as the optimized scientific discovery direction when the first reasoning similarity is greater than the first preset threshold; a loop analysis unit, which is used to correct the first scientific discovery direction and retrain the large model fine-tuning tool when the first reasoning similarity is less than or equal to the first preset threshold, and iterate until the reasoning similarity obtained by iteration is greater than the first preset threshold.

[0043] Specifically, a semantic similarity algorithm is used to calculate the similarity between the first scientific discovery direction and the standard answer. The semantic similarity algorithm usually uses deep learning models (such as BERT, RoBERTa, SBERT, etc.) to vectorize the text and calculates the semantic proximity between the two through methods such as cosine similarity, Euclidean distance, or dot product. When the first reasoning similarity is greater than the first preset threshold (for example, 0.9), it is considered that the reasoning result of this scientific discovery direction meets the expectations, so this result is directly output as the optimized scientific discovery direction. If the first reasoning similarity is lower than or equal to the first preset threshold (such as less than 0.85), it indicates that the generated scientific discovery direction may be biased or not precise enough, and it will enter the correction and retraining process. During the correction process, the first scientific discovery direction will be adjusted to correct the gap between it and the standard answer. For example, the reason for the deviation may be analyzed and it is found that a certain key factor is missing (such as the innovation point of the research method), or the reasoning direction is not specific enough (such as the proposed research direction is too general). Combining with the standard answer, the reasoning result is corrected manually or automatically, and the corrected answer is used as the new standard answer. Then, these corrected data are used as additional training samples to retrain the large model fine-tuning tool. This process will be carried out in an iterative loop. After each training, a new scientific discovery direction will be generated, and its similarity to the standard answer will be calculated again. When the reasoning similarity obtained by iteration exceeds the first preset threshold (such as reaching more than 0.95), it indicates that the model has been fully learned and optimized, and the iteration will stop, and the final reasoning result will be output as the optimized scientific discovery direction.

[0044] This iterative optimization mechanism based on semantic similarity can not only ensure that the inferred scientific discovery direction meets the actual scientific research needs, but also continuously optimize the reasoning ability of the model, making it more accurate and efficient in dealing with new scientific research problems in the future.

[0045] Furthermore, the data in the training dataset of the scientific discovery prompt words is in the form of triples, where the triple form includes inference task definition information, structured knowledge extracted from the storage module 12, and the scientific discovery direction inferred based on the combination relationship of problem-method-limit.

[0046] Specifically, when constructing the training dataset of the scientific discovery prompt words, the triple form is used for data representation and organization. Each triple consists of three main components: inference task definition information, structured knowledge extracted from the storage module 12, and the scientific discovery direction inferred based on the combination relationship of problem-method-limit. This way of organizing triples can provide a clearer and more structured input for the large model fine-tuning tool, thus promoting the learning and optimization of the model under the reverse inference mechanism.

[0047] First of all, the inference task definition information is the first component in the triple, mainly used to describe the background and objectives of the current inference task. The second part is the structured knowledge extracted from the storage module 12. The storage module 12 manages information such as research questions, methods, and limitations, and organizes them into multiple knowledge bases (such as research question knowledge base, research method knowledge base, research limitation knowledge base). In the training dataset, the structured knowledge is extracted from these knowledge bases and mapped into specific information, such as the detailed description of a specific research method or the explanation of a specific research limitation.

[0048] Finally, the third part of the triple is the scientific discovery direction inferred based on the combination relationship of problem-method-limit. In this part, by combining the known research question, research method, and research limitation information, a potential scientific discovery direction is obtained using the reverse inference mechanism. By combining these three parts into a triple, not only can each inference task be clearly defined, but also it can ensure that the large model can fully understand and utilize the existing scientific research knowledge during the fine-tuning process, and at the same time guide it to infer a scientific discovery direction that meets the actual needs.

[0049] Furthermore, when the large model fine-tuning tool performs prompt word fine-tuning training based on the reverse inference mechanism, a hierarchical learning rate strategy is adopted, and the optimal hyperparameter combination is selected through cross-validation.

[0050] Specifically, the hierarchical learning rate strategy is a training strategy that adjusts the learning rate layer by layer, aiming to enable different layers of the model to be updated at different learning rates during the fine-tuning process. For deep learning models, the parameters of different layers usually exhibit different update requirements during the learning process. The parameters of the lower layers usually have learned more general features through the pre-trained model, while the parameters of the upper layers need to be more finely tuned according to specific tasks. Therefore, during fine-tuning, a lower learning rate is used to update the parameters of the lower layers, while a higher learning rate is used for the parameters of the top layer. This strategy can ensure that the model does not destroy the general features already learned during fine-tuning, and at the same time can accelerate the adaptive adjustment of high-level features, so as to better adapt to specific tasks in the scientific research field.

[0051] For example, when dealing with research questions, methods, and limitations in scientific literature, the lower layers may include the word embedding layer, syntactic analysis layer, etc. of the language model, while the upper layers include layers specifically for reasoning and task definition. Under the hierarchical learning rate strategy, the update speed of the lower-layer lexical information is slow, while the higher-layer task-specific layers quickly adapt to specific scientific reasoning tasks.

[0052] Cross-validation is another technique for optimizing the training process. It divides the dataset into multiple subsets and takes each subset as the validation set in turn to test the model, thereby effectively evaluating the performance of the model on different data. Through cross-validation, the system can ensure that the selected combination of hyperparameters is not only applicable to a certain part of the data, but can achieve stable and efficient performance on the overall data. When selecting hyperparameters used in the fine-tuning process (such as learning rate, regularization parameter, batch size, etc.), cross-validation can help identify the optimal combination of parameters. By validating and adjusting the hyperparameters multiple times, the parameter configuration that performs best on all data subsets can be selected. Finally, based on the results of this cross-validation, the efficiency and accuracy of the model in the inference task can be achieved.

[0053] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A scientific discovery agent system embedded with a reverse reasoning mechanism for scientific and technological texts, characterized by: include: The planning module is used to call the knowledge extraction and clustering tools of the tool module to plan the input scientific and technological literature set, extract and cluster research problems, research methods and research limitations, and generate structured information; A storage module, used for hierarchically managing the structured information extracted by the planning module to form a management knowledge base; A tool module is used to provide operation tools for the planning module, storage module and execution module; The execution module is used to infer and adjust the direction of scientific discovery based on the management knowledge base and in combination with the reverse reasoning mechanism.

2. The scientific discovery agent system embedded with the reverse reasoning mechanism of scientific and technological text as claimed in claim 1 is characterized in that: The planning module includes: A research question extraction unit, used to extract research questions from the scientific and technological literature set, and to classify and identify question-related sentences and dynamically cluster them to generate a research question knowledge network; A research method extraction unit, used for extracting sentences related to research methods from the scientific and technological literature set through multi-stage training, and performing incremental clustering to generate a problem-oriented research method knowledge network; A research limitation extraction unit, used to extract limitation information related to research problems and methods from the scientific literature set, perform incremental clustering on the limitation information of each problem-method combination, and generate a scientific research knowledge discovery network; An information construction unit is used to construct the structured information using the research problem knowledge network, the research method knowledge network and the scientific research knowledge discovery network.

3. The scientific discovery agent system embedded with the reverse reasoning mechanism of scientific and technological text as claimed in claim 2 is characterized in that: The storage module comprises: A research problem knowledge base construction unit, used to extract and map problem descriptions, problem categories, and problem feature vectors based on the research problem knowledge network in the structured information to establish a research problem knowledge base; A research method knowledge base construction unit, used to extract and map method descriptions, method categories, and method feature vectors based on the research method knowledge network in the structured information to establish a research method knowledge base; A research limitation knowledge base construction unit, used to extract and map limitation descriptions, limitation categories and limitation feature vectors based on the scientific research knowledge discovery network in the structured information, and establish a research limitation knowledge base; The knowledge base integration unit is used to form the management knowledge base by hierarchical management of the research problem knowledge base, the research method knowledge base and the research limitation knowledge base.

4. The scientific discovery agent system embedded with the reverse reasoning mechanism of scientific and technological text as claimed in claim 1 is characterized in that: The tools in the tool module include knowledge extraction tools, cluster analysis tools and large model fine-tuning tools.

5. The scientific discovery agent system embedded with the reverse reasoning mechanism of scientific and technological text as claimed in claim 1 is characterized in that: The execution module includes: A prompt word generating unit, used to generate scientific discovery prompt words based on the management knowledge base by using a combination of problem-method-limitation to construct a training data set of scientific discovery prompt words; The fine-tuning training unit is used to adopt the training data set and perform prompt word fine-tuning training based on the reverse reasoning mechanism through the large model fine-tuning tool to obtain the optimized scientific discovery direction.

6. The scientific discovery agent system embedded with the reverse reasoning mechanism of scientific and technological text as claimed in claim 5 is characterized in that: It also includes a cross-domain learning module, which includes: The field sample collection unit is used to collect literature annotation data in multiple different fields, identify key research issues, methods and their limitations in each field, and obtain a structured sample database in each field; A domain focus item configuration unit, configured to configure each adaptive loss function and each reasoning feature priority information for the scientific reasoning focus items in each domain based on the structured sample database of each domain; The weight adjustment unit is used to automatically adjust the weight of the model based on the various adaptive loss functions and the various reasoning feature priority information, according to the field type of the input scientific and technological literature set, when the prompt word fine-tuning training is performed based on the reverse reasoning mechanism through the large model fine-tuning tool.

7. The scientific discovery agent system embedded with the reverse reasoning mechanism of scientific and technological text as claimed in claim 5 is characterized in that: The execution module includes: A similarity analysis unit is used to compare and analyze the first scientific discovery direction generated by the large model fine-tuning tool with the standard answer through a semantic similarity algorithm to obtain a first reasoning similarity; a similarity comparison unit, configured to output the first scientific discovery direction as an optimized scientific discovery direction when the first reasoning similarity is greater than a first preset threshold; A loop analysis unit is used to correct the first scientific discovery direction and retrain the large model fine-tuning tool when the first reasoning similarity is less than or equal to a first preset threshold, and iterate the loop until the iteratively obtained reasoning similarity is greater than the first preset threshold.

8. The scientific discovery agent system embedded with the reverse reasoning mechanism of scientific and technological text as claimed in claim 7 is characterized in that: The data in the training data set of the scientific discovery prompt words are in the form of triples, wherein the triples include reasoning task definition information, structured knowledge extracted from the storage module, and scientific discovery directions derived based on reasoning about the problem-method-limitation combination relationship.

9. The scientific discovery agent system embedded with the reverse reasoning mechanism of scientific and technological text as claimed in claim 5, It is characterized in that When fine-tuning the prompt words based on the reverse reasoning mechanism through the large model fine-tuning tool, A stratified learning rate strategy is adopted, and the optimal hyperparameter combination is selected through cross-validation.

Citation Information

Patent Citations

  • Knowledge reasoning method based on deep migration reinforcement learning

    CN118734968A

  • Complex problem reasoning method based on dynamic collaboration of large language model and domain knowledge base

    CN118798368A

  • Platform for providing high value-added intelligent research information based on prescriptive analysis and a method thereof

    KR102156287B1

Cited By

  • Material synthesis data extraction method and system based on knowledge enhancement large model

    CN121051250A