An intelligent scientific discovery system embedded with reverse reasoning mechanism of scientific and technological texts
Through the scientific discovery intelligent system embedded in the scientific text reverse reasoning mechanism, the problem that the analysis results in the existing technology are limited by expert knowledge reserves and biases, and the full process support of the scientific discovery process and the accuracy of the reasoning direction is achieved.
Patent Information
- Application Number
- CN202510215561.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-02-26
AI Technical Summary
The analysis results in the prior art are limited by experts' knowledge reserves and biases, and the large model mainly relies on basic knowledge for positive derivation, lacking the fusion of fine-grained knowledge, resulting in strong subjective dependence and a single reasoning path.
A scientific discovery intelligent system that embeds the reverse reasoning mechanism of scientific texts, including planning modules, storage modules, tool modules and execution modules, uses planning modules, extraction and clustering of scientific literature collections to study problems, methods and limitations, and combines the reverse reasoning mechanism to generate structured information and optimize the direction of scientific discovery.
It realizes full-process support for the scientific discovery process, ensures the accuracy and diversity of the reasoning direction, reduces the dependence on expert knowledge reserves and prejudice, and improves the accuracy and efficiency of the scientific research direction.
Smart Images

Figure CN120146190B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a scientific discovery intelligent system embedded with a reverse reasoning mechanism for scientific and technological texts. Background Art
[0002] In traditional methods, researchers mainly use methods such as manual literature analysis, statistical analysis, and expert discussions to deduce future research directions. Specifically, researchers first need to read a large amount of literature, organize and classify the literature, and sort out the evolution of research hotspots in chronological order. They also need to record and summarize the research methods, content, and conclusions of each literature, and find research gaps through comparison. At the same time, statistical analysis methods will be used to conduct multi-dimensional statistics on the number of published documents, conduct keyword frequency analysis, construct citation networks, and use bibliometric methods to conduct trend analysis. On this basis, by organizing expert seminars, the Delphi method, and other methods, the experience and judgment of experts in the field are gathered to form a consensus prediction on future research directions. Finally, researchers will combine the results of the literature analysis with expert opinions, consider the development laws of the discipline and social needs, evaluate technical feasibility and research conditions, and finally form a recommendation report on future research directions.
[0003] When large models are applied to infer future research directions, they primarily employ a large-scale corpus training approach. By pre-training on vast troves of scientific and technological literature, large models acquire domain knowledge and language comprehension capabilities. When presented with a research question in a specific field, they can provide relevant research direction recommendations based on the knowledge in their training data. Specifically, the large model performs semantic understanding of the input research topic, then searches its knowledge base for relevant research advances. By analyzing the relevance and temporal relationships between the literature, it identifies research hotspots and development trends. Furthermore, large models can infer potential research gaps and innovations through an understanding of research methods, experimental results, and theoretical frameworks. However, this approach primarily relies on the historical data used during model training, and the inference process is relatively closed, lacking real-time access to the latest research trends. This also limits the interpretability and reliability of the derived results.
[0004] However, existing technologies have technical problems such as the analysis results being limited by the experts' knowledge reserves and biases, and the current large models mainly relying on basic knowledge for forward deduction and lacking the integration of fine-grained knowledge, resulting in strong subjective dependence and a single reasoning path. Summary of the Invention
[0005] This application provides an intelligent scientific discovery system embedded with a reverse reasoning mechanism for scientific and technological texts, which is used to solve the technical problems existing in the existing technology, such as the analysis results are limited by the knowledge reserves and biases of experts, and the current large models mainly rely on basic knowledge for forward deduction and lack the integration of fine-grained knowledge, resulting in strong subjective dependence and a single reasoning path.
[0006] In view of the above problems, the present application provides an intelligent scientific discovery system embedded with a reverse reasoning mechanism for scientific and technological texts, including: a planning module, which is used to call the knowledge extraction and clustering tools of the tool module, plan the input scientific and technological literature set, extract and cluster research problems, research methods and research limitations, and generate structured information; a storage module, which is used to hierarchically manage the structured information extracted by the planning module to form a management knowledge base; a tool module, which is used to provide operating tools for the planning module, storage module and execution module; and an execution module, which is used to infer and optimize the direction of scientific discovery based on the management knowledge base and combined with the reverse reasoning mechanism.
[0007] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0008] The scientific discovery intelligent system provided by this application, which is embedded with a reverse reasoning mechanism for scientific and technological texts, solves the technical problems of the existing technology that the analysis results are limited by the knowledge reserves and biases of experts, and the current large model mainly relies on basic knowledge for forward deduction, lacks the integration of fine-grained knowledge, resulting in strong subjective dependence and a single reasoning path. The input scientific and technological literature set is planned through the planning module to obtain the extraction and clustering processes of research problems, research methods and research limitations respectively. The storage module is used to store and manage scientific and technological literature information and processing results. With the help of the knowledge extraction, cluster analysis and large model fine-tuning tools provided by the tool module, the execution module performs specific operation tasks. At the same time, the introduction of the reverse reasoning mechanism can reversely infer possible future research directions from the limitations of existing research, realize full-process support for the scientific discovery process, and thus ensure the accuracy of the reasoning direction. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 A schematic diagram of the structure of the scientific discovery intelligent system embedded with the reverse reasoning mechanism of scientific and technological texts provided in this application;
[0010] Figure 2 A schematic diagram of the process of constructing structured information in a scientific discovery intelligent system with an embedded reverse reasoning mechanism for scientific texts provided in this application.
[0011] Description of the accompanying drawings: planning module 11, storage module 12, tool module 13, execution module 14. DETAILED DESCRIPTION
[0012] This application provides a scientific discovery intelligent system embedded with a reverse reasoning mechanism for scientific and technological texts to solve technical problems in the existing technology, such as the analysis results are limited by the experts' knowledge reserves and biases, and the current large models mainly rely on basic knowledge for forward deduction and lack the integration of fine-grained knowledge, resulting in strong subjective dependence and a single reasoning path.
[0013] like Figure 1 As shown, the embodiment of the present application provides a scientific discovery intelligent system embedded with a reverse reasoning mechanism for scientific and technological texts, including:
[0014] The planning module 11 is used to call the knowledge extraction and clustering tool of the tool module 13 to plan the input scientific and technological literature set, extract and cluster research problems, research methods and research limitations, and generate structured information.
[0015] Specifically, the core function of the planning module 11 is to systematically plan the input set of scientific and technological literature by calling the knowledge extraction and clustering tools in the tool module 13. Specifically, this process first includes identifying and extracting the research questions, research methods and research limitations in the literature. Research questions refer to the main problems or challenges discussed in the literature, which are usually topics that have not been solved or need further exploration in a certain field; research methods refer to the specific methods or technical means used in the literature to study and solve problems, which may involve experimental design, theoretical analysis or data processing methods; research limitations refer to the shortcomings or limitations of the research pointed out in the literature, which usually reflect the inevitable limitations in the research process, such as insufficient sample size, method limitations, etc.
[0016] To effectively organize this information, the clustering tool categorizes and clusters the extracted research questions, methods, and limitations. The clustering process automatically groups similar or related research questions, methods, and limitations into the same group through an algorithm, thereby constructing an orderly, hierarchical structure. This not only helps identify the core issues in the literature, but also effectively extracts the relationships between different research methods, as well as the similarities and differences between various research limitations. For example, when planning scientific literature in a certain field, if the literature focuses on issues related to "data security" and "privacy protection," the planning module 11 will identify and cluster these issues into one group; if there are different research methods, such as "encryption technology" or "anonymization technology," these will be clustered into another group and connected to the research question group. Ultimately, the generated structured information can clearly demonstrate the current research status of the field, including key issues, methods, and limitations, facilitating subsequent in-depth analysis and optimization. It not only helps build a comprehensive knowledge framework, but also provides accurate basic data for subsequent reverse reasoning mechanisms.
[0017] The storage module 12 is used to manage the structured information extracted by the planning module in a hierarchical manner to form a management knowledge base.
[0018] Specifically, the main task of the storage module 12 is to manage the structured information extracted by the planning module 11 in a hierarchical manner, and ultimately form an efficient and continuously updated management knowledge base. In this process, the storage module 12 ensures that the access and update of knowledge become more orderly and efficient by organizing information such as research problems, research methods and research limitations in a hierarchical manner. Specifically, multiple sub-libraries are constructed according to the different levels and categories of structured information. For example, for the "research problems" section, research problems in different fields or disciplines are classified and divided into appropriate hierarchical structures based on the descriptions in the literature. Similarly, the "research methods" section is classified and managed according to characteristics such as the type of method, application scenario or research objective. The "research limitations" section will be managed in layers according to the type of limitation or influencing factors. Taking actual operation as an example, suppose the literature in a certain field involves multiple different research issues, such as "data privacy protection", "fairness of intelligent systems" and "interpretability of machine learning". According to the relevance and research background of these issues, they are stored in different sub-libraries respectively to obtain a management knowledge base, which is managed in the form of a hierarchical structure. This ensures that when a certain type of problem needs to be retrieved, the relevant data can be quickly located without the need for lengthy searches from the entire literature library.
[0019] During storage, structured information becomes more than just a static document or dataset; it becomes a dynamic, constantly updated knowledge base. New research literature can be continuously input, and the storage module 12 automatically adjusts and optimizes the existing hierarchical structure based on this new data, ensuring the accuracy and timeliness of the knowledge base. This hierarchical management not only enhances the maintainability of knowledge but also provides convenient data support for subsequent execution modules 14, enabling the system to reason and optimize based on the latest research findings.
[0020] The tool module 13 is used to provide operating tools for the planning module 11 , the storage module 12 and the execution module 14 .
[0021] Specifically, the tool module 13 is a vital part of the entire system. Its main function is to provide the necessary operating tools for the planning module 11, the storage module 12 and the execution module 14 to ensure that the modules can work together efficiently. The tool module 13 includes a variety of functional tools. The specific tool forms and functions will vary according to the needs of each module, such as knowledge extraction tools and clustering analysis tools. When processing the input scientific and technological literature, the knowledge extraction tool is used to extract key research questions, research methods, research limitations and other information from a large amount of literature. After the information is initially screened, the clustering analysis tool will cluster the extracted data and classify the information according to its similarity and correlation to form preliminary structured data. This process ensures that the planning module can efficiently sort out and plan a large amount of scientific and technological literature and generate valuable knowledge networks.
[0022] Secondly, the tool module 13 provides support for the storage module 12, helping it to efficiently manage and organize structured information. The hierarchical management and knowledge base construction in the storage module 12 rely on the data mapping tool in the tool module 13. The data mapping tool converts data such as research questions, methods, and limitations in the structured information into a format that can be easily stored and queried, and organizes them according to a predetermined hierarchical structure. For example, the classification of research questions may require the mapping tool to automatically assign them to the corresponding knowledge base, while research methods and limitations are classified and stored according to different feature vectors. This process effectively reduces the need for manual intervention and improves storage and retrieval efficiency.
[0023] Finally, the tool module 13 provides the necessary operational tools for the execution module 14. These tools are particularly crucial during large-scale model fine-tuning and reverse reasoning. These large-scale model fine-tuning tools enable the execution module 14 to adjust the inference model through reverse reasoning based on the structured information extracted from the storage module 12, thereby optimizing the direction of scientific discovery.
[0024] The execution module 14 is used to infer and optimize the direction of scientific discovery based on the management knowledge base in combination with the reverse reasoning mechanism.
[0025] Specifically, the task of Execution Module 14 is to deduce and optimize the direction of scientific discovery through a reverse reasoning mechanism based on the management knowledge base. This module not only relies on the knowledge base stored and planned previously, but also can self-adjust and optimize the reasoning process, thereby providing precise directional guidance for scientific research. During the execution process, it first obtains information such as structured research questions, research methods, and research limitations from the management knowledge base. Based on this stored information, it analyzes the relationship between research questions and methods, combined with the reverse reasoning mechanism, and infers the most promising direction of scientific discovery by analyzing the relationship between research questions and methods. The core of the reverse reasoning mechanism is to reversely deduce possible causes or research paths from known results. Specifically, the reverse reasoning mechanism first analyzes the current known limitations of scientific research based on the network of research questions and methods in the knowledge base, identifying which research questions have not been fully addressed and which methods can be further optimized. This mechanism can not only discover existing directions of scientific discovery, but also provide new ideas for unresolved research problems and propose innovative research methods. Taking a specific scientific research field as an example, the current research direction is the efficient charging technology of new energy batteries. By managing the research problem network in the knowledge base, the limitations of current charging technology can be identified, such as poor low-temperature performance and slow charging speed. Then, combined with the reverse reasoning mechanism, it is possible to derive a charging optimization method based on new electrolyte materials, thereby proposing a new and potential direction for scientific discovery.
[0026] Furthermore, as attached Figure 2 As shown, the planning module in the scientific discovery intelligent system embedded with the reverse reasoning mechanism of scientific and technological texts includes: a research problem extraction unit, which is used to extract research problems from the scientific and technological literature set, and perform classification and identification and dynamic clustering of problem-related sentences to generate a research problem knowledge network; a research method extraction unit, which is used to extract research method-related sentences from the scientific and technological literature set through multi-stage training, and perform incremental clustering to generate a problem-oriented research method knowledge network; a research limitation extraction unit, which is used to extract limitation information related to research problems and methods from the scientific and technological literature set, and perform incremental clustering on the limitation information of each problem-method combination to generate a scientific research knowledge discovery network; an information construction unit, which is used to construct the structured information with the research problem knowledge network, the research method knowledge network and the scientific research knowledge discovery network.
[0027] Specifically, research questions are extracted from the input scientific literature. To this end, the planning module uses classification and recognition technology to identify sentences related to the research questions. For example, from the full text of 500 scientific literature input by the user, key sentences related to the research questions are selected through manual annotation, including question background sentences, question description sentences, and question significance sentences. The initial corpus contains 3,000 annotated data, and positive and negative samples are configured in a 1:1 ratio. In the preprocessing stage, the text is cleaned and segmented, and Chinese word segmentation is performed to remove stop words and special characters. This batch of high-quality manually annotated corpora was used to perform preliminary training and fine-tuning on the DeBERTa model, and a classification model capable of identifying sentences related to research questions was constructed. The preliminary model achieved an accuracy of 85% on the validation set. The preliminary trained model was then used to semi-automatically annotate the input scientific literature, and the identified sentences related to research questions were manually reviewed and verified, and continuous iterative optimization was performed. Ultimately, 15,000 high-quality research question fine-tuning corpora were obtained. These corpora were used to deeply fine-tune the DeBERT model. Using the same training parameter configuration, the model performance was further improved to an accuracy of 92%. The optimized model was used to automatically extract sentences related to research questions from the scientific literature collection, and perform classification and dynamic clustering. Specifically, an incremental clustering method is employed. First, extracted sentences related to research questions are processed and naturally categorized according to topic similarity. Accurate category labels and key feature descriptions are generated for each category, achieving automatic clustering. Each time a new research question is received, the semantic similarity between the new question and the existing categories is calculated (with a similarity threshold of 0.85). Questions with similarity exceeding the threshold are directly assigned to the corresponding category; questions with similarity below the threshold are individually extracted and re-clustered to form new categories. This approach achieves dynamic, incremental clustering of research questions, ensuring the natural evolution of categories and the continuous accumulation of knowledge. During the processing, detailed reasoning is generated for each classification decision, while the feature descriptions of each category are dynamically updated to maintain the structured organization of the clustering system, ultimately forming an ever-expanding research question knowledge network.
[0028] Next, through multi-stage training, sentences related to research methods are further extracted from the scientific literature. For example, key sentences related to research methods are first selected from the full texts of 500 scientific literature input by users through manual annotation, including method description sentences, experimental design sentences, technical route sentences, etc. The initial corpus contains 4,000 annotated data, and positive and negative samples are configured in a 1:1 ratio. In the preprocessing stage, the text is cleaned and segmented, and the existing segmentation tools are used for Chinese segmentation to remove stop words and special characters. This batch of high-quality manually annotated corpora is used to perform preliminary training and fine-tuning on the RoBERTa model to construct a recognition model. A classification model for sentences related to research methods was developed, and the preliminary model achieved an accuracy of 87% on the validation set. The preliminary trained model was then used to semi-automatically annotate the full text of the input scientific literature, and the identified sentences related to research methods were manually reviewed and verified. After continuous iterative optimization, 18,000 high-quality research method fine-tuning corpora were finally obtained. These corpora were used to deeply fine-tune the RoBERTa model. With the same training parameter configuration, the model performance was further improved to an accuracy of 94%. The optimized model was used to automatically extract research method knowledge from the scientific literature set, realizing the precise extraction from literature to methods. Furthermore, an incremental clustering method is employed for each problem category. First, the initial input of 300 research method corpora for each problem category is processed, naturally categorized based on method type and application scenario, and accurate category labels and method feature descriptions are generated for each category. Deep semantic understanding is then used to initially cluster the 300 corpora for each problem category into 10-15 major method categories. The model generates detailed method feature descriptions and application boundaries for each category. In the incremental clustering phase, each time 300 new research methods are received for a problem category, the semantic similarity between the new methods and existing categories is calculated (with a similarity threshold of 0.82). Methods with similarities exceeding the threshold are directly assigned to the corresponding category; methods with similarities below the threshold are individually extracted and re-clustered to form new categories. This approach achieves dynamic incremental clustering of research methods within each problem category, ensuring the natural evolution of method categories and the continuous accumulation of method knowledge. During the processing, detailed reasoning basis is generated for each classification decision, and the method feature descriptions of each category are dynamically updated to maintain the structured organization of the problem-method system, and finally form a problem-oriented research method knowledge network, which can not only provide a solid knowledge foundation for subsequent scientific discovery reasoning, but also help scientific researchers quickly grasp the research progress and difficulties in specific fields in a systematic way.
[0029] Furthermore, we incrementally clustered the limitation information to generate a scientific research knowledge discovery network. Specifically, we extracted research limitations from the scientific literature for each problem-method combination using a multi-stage training approach. First, we manually annotated key sentences related to research limitations from the literature corresponding to the problem-method (approximately 15-20 documents per combination), including limitation descriptions, deficiency explanations, and improvement suggestions. The initial corpus contained 2,000 annotated data items, with positive and negative samples configured in a 1:1 ratio. In the preprocessing stage, the text was cleaned and segmented, using existing segmentation tools for Chinese word segmentation, removing stop words and special characters. This high-quality, manually annotated corpus was used to initially train and fine-tune the XLNet model, constructing a classification model capable of identifying sentences related to research limitations. The initial model achieved 86% accuracy on the validation set. This initially trained model was then used to semi-automatically annotate the literature for each problem-method combination. The identified sentences related to research limitations were manually reviewed and verified, and continuous iterative optimization was performed. Ultimately, 12,000 high-quality, fine-tuned corpora for research limitations were obtained. Using this corpus, the XLNet model was deeply fine-tuned. Using the same training parameter configuration, model performance was further improved to 93% accuracy. The optimized model was then used to automatically extract research limitation knowledge from the literature corresponding to each problem-method combination, achieving precise extraction of limitations from problem-method combinations. Each problem-method combination was clustered using an incremental clustering method to obtain a research knowledge discovery network. The structured information was constructed using the research problem knowledge network, the research method knowledge network, and the research knowledge discovery network. The resulting structured information framework can clearly identify which research issues are most pressing, which methods are most effective, and which limitations need to be overcome, thereby greatly improving the efficiency and accuracy of scientific discovery.
[0030] Furthermore, the storage module 12 in the scientific discovery intelligent system embedded with the reverse reasoning mechanism of scientific and technological text includes: a research problem knowledge base construction unit, which is used to extract and map problem descriptions, problem categories and problem feature vectors based on the research problem knowledge network in the structured information, and establish a research problem knowledge base; a research method knowledge base construction unit, which is used to extract and map method descriptions, method categories and method feature vectors based on the research method knowledge network in the structured information, and establish a research method knowledge base; a research limitation knowledge base construction unit, which is used to extract and map limitation descriptions, limitation categories and limitation feature vectors based on the scientific research knowledge discovery network in the structured information, and establish a research limitation knowledge base; a knowledge base integration unit, which is used to hierarchically manage the research problem knowledge base, the research method knowledge base and the research limitation knowledge base to form the management knowledge base.
[0031] Specifically, based on the research question knowledge network contained in the structured information, a detailed description of each research question is prepared, and the problem category and feature vector are extracted. This information includes the problem description, problem category, and problem feature vector. The problem description is a concise explanation of each research question, covering the core content and background of the question. The problem category categorizes the research questions by field, theme, or technical approach, facilitating subsequent search and comparison. The problem feature vector converts each research question into a mathematical vector through technical means. Next, based on the research method knowledge network contained in the structured information, each research method is described, categorized, and a feature vector is extracted. This feature vector includes information such as the method description, method category, and method feature vector. The method description includes a detailed explanation of each research method, including its scope of application, operational procedures, and scientific principles. The method category categorizes these methods by research field or technical direction, such as "experimental method," "simulation method," or "theoretical analysis method." The method feature vector quantifies these research methods into a set of numerical values. By comparing the similarity and effectiveness of research methods, the optimal method for the current research question can be efficiently identified.
[0032] Finally, based on the scientific research knowledge discovery network, the limitations related to research questions and research methods were described, and hierarchical storage management of research limitations was performed for each question-method combination. A dynamically updated research limitation knowledge base corresponding to the question-method was established, which contained information such as limitation description, limitation category, and limitation feature vector.
[0033] Through the above steps, we ultimately formed three hierarchical knowledge bases: a research problem knowledge base, a research method knowledge base, and a research limitation knowledge base. Each knowledge base contains descriptions, classifications, and feature vectors of the problem, method, and limitation, facilitating efficient subsequent knowledge management and retrieval.
[0034] Furthermore, the tools in the tool module 13 include knowledge extraction tools, cluster analysis tools and large model fine-tuning tools.
[0035] Specifically, knowledge extraction tools are used to automatically extract key information from an input set of scientific literature. They can identify and extract key information within sentences, paragraphs, or paragraphs related to research questions, methods, and limitations. Using natural language processing (NLP) technology, knowledge extraction tools can identify important elements in scientific literature, such as terminology, professional vocabulary, and research conclusions. Cluster analysis tools function on the basis of knowledge extraction. Their main task is to cluster information extracted from scientific literature, grouping it by characteristics such as topic, method, or limitation. Cluster analysis tools utilize various machine learning and data mining algorithms (such as K-means and hierarchical clustering) to identify similarities between different research directions, research questions, and their associated methods. Finally, large model fine-tuning tools are key tools for improving the system's reasoning capabilities. Through large model fine-tuning, existing large pre-trained models can be fine-tuned for specific domains to better suit task requirements.
[0036] Furthermore, the execution module 14 in the scientific discovery intelligent system embedded with the reverse reasoning mechanism of scientific and technological text includes: a prompt word generation unit, which is used to generate scientific discovery prompt words based on the management knowledge base and the problem-method-limitation combination, and construct a training data set of scientific discovery prompt words; a fine-tuning training unit, which is used to use the training data set to perform prompt word fine-tuning training based on the reverse reasoning mechanism through a large model fine-tuning tool to obtain an optimized scientific discovery direction.
[0037] Specifically, the generation of problem-method-limitation combinations involves organically combining research questions, corresponding research methods, and relevant research limitation information extracted from the management knowledge base. These combinations not only encompass existing academic research frameworks but also provide specific reasoning guidance for scientific discovery. For example, if a question is "How to improve battery charging efficiency?", the corresponding research method might be "optimizing charging algorithms," while limitation information might include "thermal management issues during charging." By integrating these elements, a specific prompt word is generated, helping the subsequent reasoning system focus on relevant scientific directions. Once the prompt words are generated, a training dataset of scientific discovery prompt words is constructed based on these prompt words. The construction of the training dataset not only incorporates the aforementioned problem-method-limitation combination but also takes into account the contextual information of each prompt word, ensuring that the dataset reflects the complexity and diversity of scientific research questions in multiple dimensions. The accuracy and comprehensiveness of the training dataset directly impact the effectiveness of large-scale model fine-tuning training, making dataset construction a critical step.
[0038] Next, using this training dataset, the large-scale model fine-tuning tool fine-tunes the cue words based on a reverse reasoning mechanism. This tool leverages an existing large-scale pre-trained model and fine-tunes it on a targeted dataset, making it more adaptable to specific domain needs. During the fine-tuning process, the model performs reverse reasoning based on the generated scientific discovery cue words, combined with input scientific literature and existing structured knowledge, to gradually adjust the model's reasoning capabilities and weights. Specifically, the reverse reasoning mechanism allows the system to not only derive scientific discovery directions from existing knowledge but also dynamically adjust the reasoning path to accommodate new discoveries or proposed research questions. Ultimately, through repeated training and fine-tuning, an optimized scientific discovery direction is obtained. This not only integrates existing research to extract effective research directions, but also reveals new research potential to a certain extent, promoting scientific and technological innovation. Therefore, the training process of generating scientific discovery cue words based on the management knowledge base and fine-tuning the large model forms a closed loop. Through repeated optimization, the accuracy and efficiency of the system's inference of scientific discovery directions are continuously improved.
[0039] Furthermore, the scientific discovery intelligent system embedded with the reverse reasoning mechanism of scientific and technological texts also includes a cross-domain learning module, which is used for: a domain sample collection unit, which is used to collect literature annotation data from multiple different fields, identify key research problems, methods and limitations of each field, and obtain a structured sample database for each field; a domain focus project configuration unit, which is used to configure each adaptive loss function and each reasoning feature priority information for the scientific reasoning focus projects in each field based on the structured sample database of each field; a weight adjustment unit, which is used to automatically adjust the weight of the model based on the each adaptive loss function and the each reasoning feature priority information, according to the domain type of the input scientific and technological literature set, when the prompt word fine-tuning training is performed based on the reverse reasoning mechanism through the large model fine-tuning tool.
[0040] Specifically, in the cross-disciplinary learning module, we first collect annotated scientific literature data from multiple scientific research fields and use natural language processing (NLP) techniques to identify key research questions, research methods, and limitations in each field. The acquisition of annotated data typically relies on a combination of manual annotation and automated extraction to ensure data accuracy and coverage. This process enables the construction of a structured sample database covering multiple disciplines. This database not only encompasses the core research directions in each field but also covers common methodologies and their limitations within that field, thereby forming an interdisciplinary knowledge framework. After obtaining the structured sample database, we configure adaptive loss functions and inference feature priority information for scientific reasoning projects in different fields. The adaptive loss function dynamically adjusts the model's penalty for mispredictions during inference based on the characteristics of different research fields. For example, in medical image analysis, the accuracy of research methods is often paramount, so the loss function may impose a higher penalty on misidentification of methods. However, in materials science research, the rationality of research limitations may be more important, so the system may assign a higher penalty to mispredictions of limitations. Furthermore, the setting of inference feature priority information guides the system in weighing the weights of research questions, methods, and limitations across different fields. For example, in artificial intelligence research, priority may be given to the innovativeness of new algorithms, while in environmental science, more attention may be paid to the applicability and generalizability of methods.
[0041] Based on these adaptive loss functions and inference feature priority information, the model weights can be automatically adjusted when executing the large model fine-tuning tool to perform prompt word fine-tuning training based on the reverse reasoning mechanism. When processing scientific literature from different fields, the parameters of the neural network will be dynamically adjusted according to the characteristics of the field to make it more suitable for the reasoning needs of the field. For example, in the field of energy science, the model may be more inclined to infer new material optimization solutions, while in the field of biomedicine, the model may be more focused on improving the direction of experimental design. This weight adjustment mechanism ensures that high inference accuracy can be maintained in multidisciplinary scenarios, and can flexibly adapt to scientific research trends, thereby improving the accuracy and applicability of scientific discoveries.
[0042] Furthermore, the execution module 14 in the scientific discovery intelligent system embedded with the reverse reasoning mechanism of scientific and technological text includes: a similarity analysis unit, which is used to compare and analyze the first scientific discovery direction generated by the large model fine-tuning tool with the standard answer through a semantic similarity algorithm to obtain a first reasoning similarity; a similarity comparison unit, which is used to output the first scientific discovery direction as an optimized scientific discovery direction when the first reasoning similarity is greater than a first preset threshold; a loop analysis unit, which is used to correct the first scientific discovery direction and retrain the large model fine-tuning tool when the first reasoning similarity is less than or equal to the first preset threshold, and iterate the loop until the reasoning similarity obtained by the iteration is greater than the first preset threshold.
[0043] Specifically, a semantic similarity algorithm is used to calculate the similarity between the first scientific discovery direction and the standard answer. Semantic similarity algorithms typically employ deep learning models (such as BERT, RoBERTa, and SBERT) to vectorize text and calculate the semantic proximity between the two using methods such as cosine similarity, Euclidean distance, or dot product. When the first inference similarity is greater than a first preset threshold (e.g., 0.9), the inference result for the scientific discovery direction is considered to meet expectations and is therefore directly output as the optimized scientific discovery direction. If the first inference similarity is less than or equal to the first preset threshold (e.g., less than 0.85), the generated scientific discovery direction may be biased or inaccurate, and a correction and retraining process will be initiated. During the correction process, the first scientific discovery direction is adjusted to correct the discrepancy with the standard answer. For example, analysis may reveal the cause of the deviation, revealing a missing key factor (e.g., an innovative research method) or a lack of specificity in the inference direction (e.g., an overly generalized proposed research direction). Based on the standard answer, the inference result is manually or automatically corrected, and the corrected answer is used as the new standard answer. The large model fine-tuning tool is then retrained using this corrected data as additional training samples. This process is repeated in an iterative loop, where new scientific discovery directions are generated after each training session and their similarity to the standard answer is recalculated. When the inference similarity obtained from the iteration exceeds a first preset threshold (e.g., above 0.95), indicating that the model has been fully learned and optimized, iteration is stopped, and the final inference result is output as the optimized scientific discovery direction.
[0044] This iterative optimization mechanism based on semantic similarity can not only ensure that the inferred direction of scientific discovery meets actual scientific research needs, but also continuously optimize the model's reasoning ability, making it more accurate and efficient when dealing with new scientific research problems in the future.
[0045] Furthermore, the data in the training data set of the scientific discovery prompt words are in the form of triples, wherein the triple form includes reasoning task definition information, structured knowledge extracted from the storage module 12, and scientific discovery directions derived based on reasoning of the problem-method-limitation combination relationship.
[0046] Specifically, when constructing a training dataset for scientific discovery prompts, data is represented and organized in the form of triples. Each triple consists of three main components: reasoning task definition information, structured knowledge extracted from storage module 12, and the direction of scientific discovery derived through reasoning based on the problem-method-limitation relationship. This triple organization provides clearer and more structured input for large model fine-tuning tools, thereby promoting model learning and optimization under the reverse reasoning mechanism.
[0047] First, the reasoning task definition information is the first component of the triple, primarily used to describe the context and objectives of the current reasoning task. The second component is structured knowledge extracted from storage module 12. Storage module 12 manages information such as research questions, methods, and limitations, and organizes this information into multiple knowledge bases (e.g., research question knowledge base, research method knowledge base, and research limitation knowledge base). In the training dataset, structured knowledge is extracted from these knowledge bases and mapped into specific information, such as a detailed description of a specific research method or an explanation of a specific research limitation.
[0048] Finally, the third component of the triplet is the direction of scientific discovery, derived from reasoning about the problem-method-limitation relationship. This component uses a reverse reasoning mechanism to derive a potential direction for scientific discovery by combining known research questions, research methods, and research limitations. Combining these three components into a triplet not only clearly defines each reasoning task but also ensures that the large model fully understands and leverages existing scientific knowledge during fine-tuning, guiding it to infer directions of scientific discovery that meet practical needs.
[0049] Furthermore, when fine-tuning the prompt words through the large model fine-tuning tool based on the reverse reasoning mechanism, a hierarchical learning rate strategy is adopted, and the optimal hyperparameter combination is selected through cross-validation.
[0050] Specifically, the hierarchical learning rate strategy is a training strategy that adjusts the learning rate layer by layer. It aims to enable different layers of the model to be updated at different learning rates during the fine-tuning process. For deep learning models, the parameters of different layers usually have different update requirements during the learning process. The parameters of the bottom layer have usually learned more general features through pre-trained models, while the parameters of the upper layer require more fine-tuning based on specific tasks. Therefore, during fine-tuning, a lower learning rate is used to update the bottom-level parameters, while a higher learning rate is used for the top-level parameters. This strategy can ensure that the model does not destroy the general features that have been learned during fine-tuning, while it can also accelerate the adaptive adjustment of high-level features, thereby adapting to specific tasks in the field of scientific research more quickly.
[0051] For example, when addressing research questions, methods, and limitations in scientific literature, the bottom layer might include word embeddings and syntactic analysis from a language model, while the upper layers might include specialized layers for reasoning and task definition. Under a hierarchical learning rate strategy, the underlying lexical information is updated more slowly, while higher-level task-specific layers quickly adapt to specific scientific reasoning tasks.
[0052] Cross-validation is another technique for optimizing the training process. It effectively evaluates the performance of the model on different data by dividing the dataset into multiple subsets and using each subset as a validation set to test the model in turn. Through cross-validation, the system can ensure that the selected hyperparameter combination is not only applicable to a certain part of the data, but can achieve stable and efficient performance on the overall data. When selecting hyperparameters used in the fine-tuning process (such as learning rate, regularization parameter, batch size, etc.), cross-validation can help identify the optimal parameter combination. By verifying and adjusting hyperparameters multiple times, the parameter configuration that performs best on all data subsets can be selected. Ultimately, based on the results of this cross-validation, the model can achieve high efficiency and accuracy in inference tasks.
[0053] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A scientific discovery agent system embedded with a reverse reasoning mechanism for scientific texts, characterized by: include: The planning module is used to call the knowledge extraction and clustering tools of the tool module to plan the input scientific and technological literature set, extract and cluster research questions, research methods and research limitations, and generate structured information; A storage module, configured to manage the structured information extracted by the planning module in a hierarchical manner to form a management knowledge base; The tool module is used to provide operation tools for the planning module, storage module and execution module; An execution module, configured to infer and optimize the direction of scientific discovery based on the management knowledge base and in combination with a reverse reasoning mechanism; Wherein, the execution module includes: A prompt word generation unit is used to generate scientific discovery prompt words based on the management knowledge base using a problem-method-limitation combination to construct a training data set of scientific discovery prompt words; A fine-tuning training unit is used to use the training data set to perform prompt word fine-tuning training based on a reverse reasoning mechanism through a large model fine-tuning tool to obtain an optimized scientific discovery direction; Wherein, the execution module includes: A similarity analysis unit is configured to compare and analyze the first scientific discovery direction generated by the large model fine-tuning tool with the standard answer using a semantic similarity algorithm to obtain a first reasoning similarity; a similarity comparison unit, configured to output the first scientific discovery direction as an optimized scientific discovery direction when the first reasoning similarity is greater than a first preset threshold; A loop analysis unit is used to correct the first scientific discovery direction and retrain the large model fine-tuning tool when the first reasoning similarity is less than or equal to a first preset threshold, and iterate the loop until the reasoning similarity obtained by the iteration is greater than the first preset threshold.
2. The scientific discovery agent system embedded with the reverse reasoning mechanism of scientific and technological texts as claimed in claim 1 is characterized in that: The planning module includes: A research question extraction unit is used to extract research questions from the scientific literature set, and to classify and identify sentences related to the questions and dynamically cluster them to generate a research question knowledge network; A research method extraction unit is used to extract sentences related to research methods from the scientific literature set through multi-stage training, and perform incremental clustering to generate a problem-oriented research method knowledge network; A research limitation extraction unit is used to extract limitation information related to research questions and methods from the scientific literature set, perform incremental clustering on the limitation information of each question-method combination, and generate a scientific research knowledge discovery network; An information construction unit is used to construct the structured information using the research problem knowledge network, the research method knowledge network and the scientific research knowledge discovery network.
3. The scientific discovery agent system embedded with the reverse reasoning mechanism of scientific and technological texts as claimed in claim 2 is characterized in that: The storage module includes: A research question knowledge base construction unit, configured to extract and map question descriptions, question categories, and question feature vectors based on the research question knowledge network in the structured information, and to establish a research question knowledge base; A research method knowledge base construction unit is used to extract and map method descriptions, method categories, and method feature vectors based on the research method knowledge network in the structured information to establish a research method knowledge base; A research limitation knowledge base construction unit is used to extract and map limitation descriptions, limitation categories, and limitation feature vectors based on the scientific research knowledge discovery network in the structured information to establish a research limitation knowledge base; The knowledge base integration unit is used to hierarchically manage the research problem knowledge base, the research method knowledge base and the research limitation knowledge base to form the management knowledge base.
4. The scientific discovery agent system embedded with the reverse reasoning mechanism of scientific and technological texts as claimed in claim 1 is characterized in that: The tools in the tool module include knowledge extraction tools, cluster analysis tools and large model fine-tuning tools.
5. The scientific discovery agent system embedded with the reverse reasoning mechanism of scientific and technological texts as claimed in claim 1 is characterized in that: It also includes a cross-disciplinary learning module, which includes: The domain sample collection unit is used to collect literature annotation data from multiple different fields, identify key research issues, methods and limitations in each field, and obtain a structured sample database in each field; A domain focus item configuration unit, configured to configure each adaptive loss function and each reasoning feature priority information for the scientific reasoning focus items in each domain based on the structured sample database of each domain; The weight adjustment unit is used to automatically adjust the weight of the model based on the various adaptive loss functions and the various reasoning feature priority information, according to the field type of the input scientific and technological literature set, when the prompt word fine-tuning training is performed based on the reverse reasoning mechanism through the large model fine-tuning tool.
6. The scientific discovery agent system embedded with the reverse reasoning mechanism of scientific and technological texts as claimed in claim 1 is characterized in that: The data in the training data set of the scientific discovery prompt words is in the form of triples, wherein the triple form includes reasoning task definition information, structured knowledge extracted from the storage module, and scientific discovery direction derived based on the problem-method-limitation combination relationship reasoning.
7. The scientific discovery agent system embedded with the reverse reasoning mechanism of scientific and technological texts as claimed in claim 1 is characterized in that: When fine-tuning the prompt words based on the reverse reasoning mechanism through the large model fine-tuning tool, a hierarchical learning rate strategy is adopted, and the optimal hyperparameter combination is selected through cross-validation.
Citation Information
Patent Citations
Knowledge reasoning method based on deep migration reinforcement learning
CN118734968A
Complex problem reasoning method based on dynamic collaboration of large language model and domain knowledge base
CN118798368A