Scientific and technological research achievement tracking method and system supported by big data
Through big data technology and knowledge graph analysis, the problem that the value of failed research is not valued, and the in-depth exploration and reuse of failed research is achieved, the scientificity and accuracy of scientific research evaluation is improved, and the allocation of scientific research resources is optimized.
Patent Information
- Application Number
- CN202510176559.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing scientific and technological research results tracking methods fail to effectively analyze the value of failed research, resulting in the role of failed research in the knowledge system that has not been paid attention to and reflected as due. At the same time, there is a lack of comprehensive capture of the changing trends of the dynamic impact of research results, and the evaluation of the value of research results is not scientific and accurate enough.
Failed research and successful research data are obtained through big data technology, the knowledge graph of failed research and successful research is constructed, dynamic correlation relationships are analyzed, causal inference models are established, failed research is quantitatively evaluated, contribution scores are generated, and the potential of failed research is evaluated through time series analysis and potential prediction models, low-citation but high-potent failed research is screened out, and failed research with low citation but high potential is performed for reuse matching analysis.
It has achieved in-depth exploration and reuse of the value of failed research, improved the scientific nature and accuracy of scientific research evaluation, dynamically tracked the changing trends of research results, and optimized the rational allocation and utilization of scientific research resources.
Smart Images

Figure CN120067343A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of research result tracking, and particularly relates to a method and system for tracking scientific and technological research results supported by big data. Background Art
[0002] The technical field of scientific and technological research result tracking refers to a technical and methodological system for collecting, analyzing, evaluating, and managing various results (including papers, patents, experimental data, failed research, etc.) generated during the scientific research process. Its purpose is to reveal the academic influence, application value, and development trend through the dynamic tracking and in-depth analysis of research results, providing scientific decision-making support for scientific research institutions, academia, and policymakers.
[0003] This field covers a variety of technical means and methods, including result retrieval and aggregation based on big data, natural language processing (NLP) technology for semantic analysis, time series analysis for dynamic citation trend research, knowledge graphs for constructing result association relationships, and machine learning models for predicting and evaluating potential value. In addition, this field also focuses on the dissemination and cross-influence of research results among disciplines, especially in data-driven scientific research, emphasizing the reuse of failed research and lowly cited results.
[0004] The technology of scientific and technological research result tracking is widely applied in aspects such as the rational allocation of scientific research resources, forward-looking prediction of research directions, evaluation of innovation outputs, and scientific research management and policy formulation. With the development of big data, artificial intelligence, and information technology, this field is developing rapidly, aiming to build a more open, dynamic, and efficient scientific research ecosystem.
[0005] Current tracking methods tend to focus on the result data of successful research, while the analysis of failed research is generally insufficient, resulting in the role of failed research in the knowledge system not being given due attention and reflection. Secondly, these methods have limitations in the dynamic tracking of research results, lacking time series analysis means for citation data and being unable to comprehensively capture the changing trend of the influence of research results at different stages, especially the potential value that failed research may generate as the field progresses has not been discovered in a timely manner. In addition, for the evaluation of the value of research results, existing methods often rely on a single indicator and fail to conduct a comprehensive potential prediction by combining multi-dimensional factors (such as citation trends, semantic associations, contribution relationships), thus reducing the scientificity and accuracy of the evaluation. Finally, in terms of the reuse of research resources, current methods lack effective technical means to conduct in-depth matching analysis between failed research and current research needs, resulting in these potentially valuable data not being fully integrated and utilized. Summary of the Invention
[0006] The purpose of the present invention is to provide a method for tracking scientific and technological research results supported by big data, aiming to solve the technical problems existing in the prior art identified in the background art.
[0007] The present invention is implemented as follows. A method for tracking scientific and technological research results supported by big data, the method comprising:
[0008] Obtain the failed research data and successful research data in the relevant field of the current research, fuse and construct a research result dataset, analyze each failed research, and package and store the analyzed failure reasons, data details, and research background.
[0009] Based on the research result dataset, construct a knowledge graph of failed research and successful research, analyze the dynamic association relationship between failed research and successful research, establish a causal inference model, quantitatively evaluate the contribution of failed research, and generate a contribution score for failed research.
[0010] Conduct time series analysis on the citation data of failed research and successful research, establish a citation curve for each research, identify the change trend of the citation volume of each failed research, and based on the citation curve, contribution score, and change trend, establish a potential prediction model, output the potential score of each failed research, and measure the future change trend of the citation volume.
[0011] Set a citation volume threshold and a potential score threshold, screen the failed research with a current citation volume lower than the citation volume threshold but a potential score higher than the potential score threshold, construct a reuse model for failed research, match and analyze the data content in the failed research with the current research requirements, and generate a list of available data.
[0012] As a further solution of the present invention, the obtaining of the failed research data and successful research data in the relevant field and the fusion and construction of the research result dataset specifically include:
[0013] Identify the field characteristics of the current research, retrieve and identify the existing failed research data and successful research data in the associated field.
[0014] Classify the retrieved data according to the success or failure status, field, and theme of the research.
[0015] Fuse the failed and successful research data into a structured research result dataset, and store the data classified according to the classification results.
[0016] According to the descriptive fields in the data, extract the reasons for failure, identify the key factors for failure, simultaneously extract the experimental design content, construct a causal chain of the experimental process, and identify the research background field, and synchronously package and store the key factors for failure, the causal chain of the experimental process, and the research background field in the research result dataset.
[0017] As a further aspect of the present invention, the quantitative evaluation of the contribution of failed research is carried out to generate a contribution score for failed research, which specifically includes:
[0018] Adopt natural language processing to perform text content analysis on the failed research results and successful research results, and calculate the semantic similarity between the failed research results and the successful research results;
[0019] Based on the semantic similarity, construct a semantic association matrix representing the exercise intensity between failed research and successful research;
[0020] Compare the causal chain of the experimental process of failed research with that of successful research, identify the overlapping parts in experimental design and variable selection between the two, mark the experimental steps in failed research that are used in subsequent successful research, and generate a heuristic association model based on this;
[0021] Match the background fields of failed research and successful research, identify the potential inspiration of failed research for subsequent research in the same background field, and dynamically update the heuristic association model;
[0022] Using the failed research results and successful research results as nodes and the outputs of the semantic association matrix and the heuristic association model as edges, construct a knowledge graph of failed research and successful research;
[0023] Based on the experimental causal chain, extract the interrelated nodes from the technical routes of failed research and successful research, establish an inference model, analyze whether the failed research constitutes a causal relationship with the successful research, and quantify the contribution degree of the failed research nodes, calculate the influence value of the failed research on the successful research, and calculate the contribution score:
[0024] Contribution score = A 1 · Semantic similarity weight + A 2 · Experimental causal chain coverage + A 3 · Citation impact factor;
[0025] Wherein, A 1 、A 2 、A 3 Are weight parameters.
[0026] As a further aspect of the present invention, the potential score of each failed research is output to measure the future citation volume change trend, which specifically includes:
[0027] Read the citation data sources of each failed research and successful research in the research result dataset, sort the citation data in chronological order, and create a time series dataset;
[0028] Construct the citation curves of each failed research and successful research to dynamically capture the change trend of the citation volume of research results;
[0029] Based on the citation volume change trend, the contribution score of failed research, and semantic similarity, a potential prediction model is constructed to generate a potential score for failed research:
[0030] Potential score = B 1 · Citation change growth rate + B 2 · Contribution score + B 3 · Field relevance;
[0031] Among them, B 1 、B 2 、B 3 are weight parameters.
[0032] As a further solution of the present invention, the failed research with the current citation volume lower than the citation volume threshold but the potential score higher than the potential score threshold is screened to construct a failed research reuse model, which specifically includes:
[0033] Define the citation volume threshold according to the average citation volume in the field, and define the potential score threshold according to the distribution state of the potential score;
[0034] Take the citation volume threshold and the potential score threshold as parallel conditions and combine them into a screening formula:
[0035] Citation volume < citation volume threshold AND potential score > potential score threshold;
[0036] Screen the failed research according to the screening formula, obtain all the failed research that meets the parallel conditions, and organize them into a screening data set;
[0037] Based on the semantic matching algorithm, establish a failed research reuse model, compare the content of the failed research with the research requirements, including: research direction similarity, technical requirement similarity, and adaptation requirement similarity, and add them up to obtain the final research matching degree to judge the feasibility of reusing the failed research data;
[0038] Integrate the feasibility data according to the research direction similarity, technical requirement similarity, and adaptation requirement similarity respectively, and output a list of available data.
[0039] Another object of the present invention is to provide a scientific and technological research result tracking system supported by big data, and the system includes:
[0040] A research data extraction module, which is used to obtain the failed research data and successful research data in the relevant field of the current research, fuse and construct a research result data set, analyze each failed research, and package and store the analyzed failure reasons, data details, and research backgrounds;
[0041] A knowledge graph construction module, which is used to construct a knowledge graph of failed and successful studies based on a research result dataset, analyze the dynamic correlation relationship between failed and successful studies, establish a causal inference model, quantitatively evaluate the contributions of failed studies, and generate contribution scores for failed studies;
[0042] A potential prediction module, which is used to perform time series analysis on the citation data of failed and successful studies, establish a citation curve for each study, identify the change trend of the citation volume of each failed study, and establish a potential prediction model based on the citation curve, contribution score and change trend, and output the potential score of each failed study to measure the future change trend of the citation volume;
[0043] A reuse analysis module, which is used to set a citation volume threshold and a potential score threshold, screen failed studies whose current citation volume is lower than the citation volume threshold but whose potential score is higher than the potential score threshold, construct a failed study reuse model, match and analyze the data content in the failed studies with the current research requirements, and generate a list of available data.
[0044] As a further solution of the present invention, the research data extraction module includes:
[0045] A research field feature recognition unit, which is used to recognize the field features of the current research, retrieve and identify the failed research data and successful research data existing in the associated field;
[0046] A research data classification unit, which is used to classify the retrieved data according to the success or failure status, field and theme of the research;
[0047] A research result dataset fusion unit, which is used to fuse the failed and successful research data into a structured research result dataset, and store the data classified according to the classification results;
[0048] A causal chain construction unit, which is used to extract the reasons for failure, identify the key factors for failure according to the descriptive fields in the data, extract the experimental design content at the same time, construct a causal chain of the experimental process, and identify the research background field, and synchronously package and store the key factors for failure, the causal chain of the experimental process and the research background field in the research result dataset.
[0049] As a further solution of the present invention, the knowledge graph construction module includes:
[0050] A semantic similarity calculation unit, which is used to perform text content analysis on the failed research results and successful research results by using natural language processing, and calculate the semantic similarity between the failed research results and the successful research results;
[0051] A semantic association matrix construction unit, configured to construct a semantic association matrix representing the exercise intensity between failed studies and successful studies based on semantic similarity;
[0052] An experimental process overlap analysis unit, configured to compare the experimental process causal chains of failed studies with those of successful studies, identify the overlapping parts in experimental design and variable selection between the two, mark the experimental steps in failed studies that are utilized by subsequent successful studies, and generate a heuristic association model based on this;
[0053] A background field matching unit, configured to match the background fields of failed studies and successful studies, identify the potential inspiration of failed studies for subsequent studies within the same background field, and dynamically update the heuristic association model;
[0054] A knowledge graph construction unit, configured to construct a knowledge graph of failed studies and successful studies with the results of failed studies and successful studies as nodes and the outputs of the semantic association matrix and the heuristic association model as edges;
[0055] A contribution score quantification unit, configured to extract the mutually related nodes from the technical routes of failed studies and successful studies based on the experimental causal chain, establish an inference model, analyze whether the failed studies constitute a causal relationship with the successful studies, and quantify the contribution degree of the failed study nodes, and calculate the influence value of the failed studies on the successful studies and calculate the contribution score.
[0056] As a further solution of the present invention, the potential prediction module includes:
[0057] A time series dataset creation unit, configured to read the citation data sources of each failed study and successful study in the research result dataset, sort the citation data in chronological order, and create a time series dataset;
[0058] A dynamic trend capture unit, configured to construct the citation curves of each failed study and successful study and dynamically capture the changing trend of the citation volume of the research results;
[0059] A potential score generation unit, configured to construct a potential prediction model based on the citation volume change trend, the contribution score of the failed study, and the semantic similarity, and generate a potential score for the failed study.
[0060] As a further solution of the present invention, the reuse analysis module includes:
[0061] A potential score threshold definition unit, configured to define a citation volume threshold according to the average citation volume in the field and define a potential score threshold according to the distribution state of the potential scores;
[0062] A failed study screening unit, configured to use the citation volume threshold and the potential score threshold as parallel conditions to form a screening formula;
[0063] A screening dataset sorting unit, which is used to screen failed studies according to a screening formula, obtain all failed studies that meet the parallel conditions, and sort them into a screening dataset;
[0064] A matching degree calculation unit, which is used to establish a reuse model for failed studies based on a semantic matching algorithm, compare the content of failed studies with research requirements, including: similarity of research directions, similarity of technical requirements, and similarity of adaptation requirements, and add them up to obtain the final research matching degree, and judge the feasibility of reusing failed research data;
[0065] An available data list integration unit, which is used to integrate feasible data according to the similarity of research directions, the similarity of technical requirements, and the similarity of adaptation requirements respectively, and output an available data list.
[0066] The beneficial effects of the present invention are as follows:
[0067] This method accurately tracks and analyzes scientific and technological research results through big data technology, especially the value mining and reuse of failed studies, realizing the in-depth optimization of scientific research evaluation and resource utilization. The traditional scientific research system often ignores failed studies, while this method breakthroughly incorporates failed studies into the contribution evaluation system, and reveals their potential promoting effect on scientific progress through knowledge graphs and causal models. At the same time, the combination of citation time series analysis and potential prediction models makes the research tracking change from static to dynamic, and can accurately predict the future academic value of failed studies, providing data support for the reasonable allocation of scientific research resources. In addition, this method constructs a reuse model through a semantic matching algorithm, which promotes low-citation but high-potential failed studies to be effectively integrated into the current research requirements and transformed into available resources, significantly reducing the waste of scientific research resources and duplicate costs. Description of the Drawings
[0068] Figure 1 It is a flowchart of the method for tracking scientific and technological research results supported by big data provided by an embodiment of the present invention;
[0069] Figure 2 It is a flowchart of obtaining failed research data and successful research data in related fields, and fusing and constructing a research result dataset provided by an embodiment of the present invention;
[0070] Figure 3 It is a flowchart of quantitatively evaluating the contribution of failed studies and generating a contribution score for failed studies provided by an embodiment of the present invention;
[0071] Figure 4 It is a flowchart of outputting the potential score of each failed study and measuring the future citation volume change trend provided by an embodiment of the present invention;
[0072] Figure 5A flowchart for screening failed studies with a current citation count lower than a citation count threshold but a potential score higher than a potential score threshold provided by an embodiment of the present invention, and constructing a reuse model for failed studies;
[0073] Figure 6 A structural block diagram of a scientific research result tracking system supported by big data provided by an embodiment of the present invention;
[0074] Figure 7 A structural block diagram of a research data extraction module provided by an embodiment of the present invention;
[0075] Figure 8 A structural block diagram of a knowledge graph construction module provided by an embodiment of the present invention;
[0076] Figure 9 A structural block diagram of a potential prediction module provided by an embodiment of the present invention;
[0077] Figure 10 A structural block diagram of a reuse analysis module provided by an embodiment of the present invention. Detailed implementation manners
[0078] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0079] It can be understood that the terms "first", "second", etc. used in the present application can be used in this document to describe various elements, but unless otherwise specified, these elements are not limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the present application, the first xx script can be called the second xx script, and similarly, the second xx script can be called the first xx script.
[0080] Figure 1 A flowchart of a scientific research result tracking method supported by big data provided by an embodiment of the present invention, as Figure 1 shown, the method includes:
[0081] S100, obtaining failed research data and successful research data in the relevant field of the current research, fusing and constructing a research result data set, and analyzing each failed research, and packaging and storing the analyzed failure reasons, data details, and research background;
[0082] In this step, by paying attention to the hotspots, key technical issues in the current research field, as well as existing academic literature, experimental reports, etc., relevant failed research data and successful research data are obtained using big data retrieval tools (such as academic databases, patent databases, etc.). For the acquisition of failed research data, not only obvious failures (such as research that clearly states that the experiment did not meet the expected goals) should be concerned, but also through text mining techniques and natural language processing tools, those implicit failures (such as statements in the research conclusions about experimental limitations or the non - establishment of certain assumptions) should be identified. The acquisition of successful research data requires emphasizing high quality and authority to ensure its reliability for subsequent analysis.
[0083] Secondly, according to the content of the acquired data, it needs to be classified in multiple dimensions, including the classification of research topics, experimental methods, application fields, and research time. The classification process needs to combine artificial intelligence techniques, using classification algorithms and domain knowledge graphs to ensure the accuracy and fine - granularity of the classification. In the process of data fusion, unified data standards and formats must be adopted to facilitate subsequent analysis and processing, ensuring data compatibility and operability.
[0084] Then, the analysis of failed research data is crucial. Through text mining techniques, the core information of the failure reasons is extracted, including key defects in experimental design, insufficient theoretical basis, limitations of technical means, etc. Subsequently, through causal analysis, a causal chain of the experimental process is constructed, which can not only trace the root cause of the failure but also provide directions for improvement in future research. When analyzing the background of failed research, by combining domain knowledge and external information, the technical environment, theoretical basis, and application background on which the failed research depends are accurately identified, thus providing comprehensive information support for the packaging and storage of data. The reasons for failed research, the experimental causal chain, and background information are finally stored together with the data of successful research in a unified research result dataset, laying a foundation for subsequent knowledge graph construction and causal inference.
[0085] This step can extract the hidden value from failed research and convert it into a reference basis for guiding future research. By deeply analyzing the reasons for failure and their backgrounds, S100 can reveal potential innovation points or technical short - boards in some failed research, helping researchers avoid existing problems and optimize the research path. In addition, combining the storage of failed research and successful research data can not only form global data support but also improve the efficiency and accuracy of subsequent analysis through the structured and standardized processing of data. More importantly, this step makes full use of the rich and often overlooked information resources in failed research, providing a solid foundation for establishing a more comprehensive and dynamic knowledge system in the scientific research field and helping scientific research move towards a more efficient and innovative direction.
[0086] Such as Figure 2As shown, obtaining the failed research data and successful research data in the relevant field and integrating them to construct a research result dataset specifically includes:
[0087] S110, identifying the field characteristics of the current research, retrieving and identifying the failed research data and successful research data existing in the associated field;
[0088] S120, classifying the retrieved data according to the success or failure status, field, and theme of the research;
[0089] S130, fusing the failed and successful research data into a structured research result dataset, and storing the data classified according to the classification results;
[0090] S140, extracting the reasons for failure, identifying the key factors of failure according to the descriptive fields in the data, simultaneously extracting the experimental design content, constructing the causal chain of the experimental process, and identifying the research background field, and synchronously packing and storing the key factors of failure, the causal chain of the experimental process, and the research background field in the research result dataset.
[0091] S200, constructing a knowledge graph of failed research and successful research based on the research result dataset, analyzing the dynamic association relationship between failed research and successful research, establishing a causal inference model, quantitatively evaluating the contribution of failed research, and generating a contribution score for failed research;
[0092] In this step, content analysis is performed on the data texts of failed research and successful research through natural language processing (NLP) technology. This analysis is not limited to extracting keywords or themes, but generates context-related semantic embedding representations based on a deep learning model to capture the potential semantic relationships between failed and successful research. Based on these embedding representations, semantic similarity is calculated to generate a quantifiable semantic association matrix, and the values in the matrix reflect the content similarity and correlation degree between research results. This analysis method can effectively identify the key parts in failed research that may inspire successful research.
[0093] Secondly, for the experimental processes of failed research and successful research, S200 further identifies the overlapping parts in experimental design, variable selection, and experimental steps by comparing the experimental causal chains. Through these overlapping points, it can be intuitively marked which experimental steps in failed research are borrowed and optimized by subsequent successful research to generate an inspiration association model. This model dynamically captures the contribution of failed research in the scientific research chain, making failed research no longer simply classified as "useless", but becoming an important "inspiration source" in the process of scientific exploration.
[0094] Meanwhile, S200 also combines the matching analysis in the background field to identify the potential promoting effect of failed studies on subsequent studies in the same field. For example, when the similarity of the background fields of failed studies and successful studies is relatively high, some theoretical assumptions or data results in the failed studies may directly influence subsequent successful studies. By dynamically updating the heuristic association model, the correlation representation between the two is further improved.
[0095] In the construction of the knowledge graph, S200 uses failed studies and successful studies as nodes, and the outputs of the semantic association matrix and the heuristic association model as edges to construct a knowledge graph oriented to the association between failed and successful studies. This graph is not only a relational network of research data, but also a visualization tool for the dynamic development of research, showing the interaction and influence path between failed studies and successful studies.
[0096] Finally, a causal inference model is established through the experimental causal chain. The associated nodes of failed-success studies are extracted from the technical route, and further analysis is carried out on whether the failed studies constitute a causal relationship with the successful studies. For example, it is judged whether a certain failed study provides important proof-of-concept or experimental basis, directly triggering the progress of subsequent successful studies. On this basis, S200 quantifies the contribution of failed studies to generate a contribution score. The contribution score is jointly composed of the semantic similarity weight, the coverage of the experimental causal chain, and the citation impact factor. The weight parameters of A 1 、A 2 、A 3 can be adjusted according to the characteristics of the field research to ensure the rationality and scientificity of the score.
[0097] This step comprehensively reveals the multiple association relationships between failed studies and successful studies through multi-dimensional semantic, experimental, and background matching analyses, avoiding the simplistic understanding of the value of failed studies in traditional research. By introducing the knowledge graph and the causal inference model, the failed studies are transformed from "static data" into "dynamic knowledge", demonstrating their key role in the scientific research chain. Secondly, the quantitative evaluation method for the contribution of failed studies is highly scientific and objective. It not only considers semantic similarity, but also combines the coverage of the experimental causal chain and the citation impact factor, ensuring the diversification and comprehensiveness of the evaluation results. In addition, the introduction of the causal inference model provides a unique perspective for scientific research, enabling the identification of hidden innovation points or theoretical contributions in failed studies, and helping researchers to more profoundly understand the significance of failure in the scientific process. Through the visualization of the knowledge graph, researchers can intuitively discover the profound impact of failed studies on successful studies, thus effectively promoting the scientific practice of reusing failed studies.
[0098] As Figure 3 shown, the quantitative evaluation of the contribution of failed studies to generate a contribution score for failed studies specifically includes:
[0099] Perform text content analysis on failed research results and successful research results using natural language processing, and calculate the semantic similarity between failed research results and successful research results;
[0100] Construct a semantic association matrix representing the exercise intensity between failed research and successful research based on semantic similarity;
[0101] Compare the causal chain of the experimental process of failed research with that of successful research, identify the overlapping parts in experimental design and variable selection between the two, label the experimental steps in failed research that are used in subsequent successful research, and generate a heuristic association model based on this;
[0102] Match the background fields of failed research and successful research, identify the potential inspiration of failed research for subsequent research in the same background field, and dynamically update the heuristic association model;
[0103] Construct a knowledge graph of failed research and successful research with failed research results and successful research results as nodes and the outputs of the semantic association matrix and heuristic association model as edges;
[0104] Based on the experimental causal chain, extract the interrelated nodes from the technical routes of failed research and successful research, establish an inference model, analyze whether failed research constitutes a causal relationship with successful research, and quantify the contribution degree of failed research nodes, calculate the influence value of failed research on successful research, and calculate the contribution score:
[0105] Contribution score = A 1 · Semantic similarity weight + A 2 · Coverage of experimental causal chain + A 3 · Citation impact factor;
[0106] Where A 1 、A 2 、A 3 Are weight parameters.
[0107] S300, conduct time series analysis on the citation data of failed research and successful research, establish a citation curve for each research, identify the change trend of the citation volume of each failed research, and based on the citation curve, contribution score and change trend, establish a potential prediction model, output the potential score of each failed research, and measure the future change trend of the citation volume;
[0108] This step is based on the citation data of each failed study and successful study in the research result dataset, comprehensively reads the citation sources and sorts them in chronological order to construct a time series dataset. In this process, it is necessary to make full use of big data processing technologies, integrate citation data from various sources such as academic journals, conference papers, patent citations, white papers, etc., and standardize the data in combination with timestamps. At the same time, filter out invalid citations or non-academic citations to ensure the accuracy and authority of the data. Through the refined collation of time series data, it can intuitively reflect the academic attention and citation dynamics of each research result at different time stages.
[0109] Subsequently, based on the time series data, construct the citation curves for each failed study and successful study. The citation curves visually show the changing trend of the citation quantity of the research results over time. For example, some failed studies may have a low citation volume in the short term, but with the progress of the research field or technological breakthroughs, their citation volume will show a significant increase.
[0110] Next, construct a scientific potential prediction model through the citation volume change trend, the contribution score of failed studies, and semantic similarity. The citation volume change trend reflects the dynamic attention degree of failed studies in the academic community, while the contribution score is derived from the quantitative evaluation of the value of failed studies in S200, reflecting the inspiring contributions of failed studies in aspects such as experimental methods, theoretical basis, or data support. Semantic similarity combines the relevance between failed studies and current research topics in related fields to measure their content adaptability and background relevance. The potential prediction model generates a potential score for each failed study by integrating these factors. In the calculation formula of the potential score, the weight parameters of the citation change growth rate, contribution score, and field relevance can be adjusted according to specific fields or research objectives to adapt to the requirements of different disciplines for the evaluation of research potential.
[0111] In addition, S300 is not limited to the analysis of a single factor, but also combines multi-dimensional data to comprehensively predict the future citation volume of failed studies. For example, for some failed studies with a low citation volume at the current stage but a high potential score, the potential prediction model can identify the possibility of their becoming future research hotspots and capture their hidden academic value. The prediction results of this model can be used as an important decision-making basis for future research topic selection, resource allocation, and adjustment of scientific research directions.
[0112] This step realizes the accurate capture of the dynamic changes of research results through time series analysis and the construction of citation curves, revealing the academic influence of failed research in different time dimensions. This dynamic analysis makes up for the deficiencies of traditional static evaluation methods, enabling the value of failed research to be gradually explored through the time dimension. Secondly, the multi-factor integration feature of the potential prediction model makes the evaluation more comprehensive and scientific. The citation change rate, contribution score, and field relevance complement each other, avoiding the biases that may occur in a single evaluation indicator and improving the credibility of potential evaluation. Thirdly, this step predicts the future citation trends of failed research, which can help researchers and institutions discover potentially valuable failed research in advance, thus effectively utilizing these research results and promoting the efficient allocation of scientific research resources.
[0113] As Figure 3 shown, the potential score of each failed research is output, which measures the future citation volume change trend, specifically including:
[0114] S310, read the citation data sources of each failed research and successful research in the research result dataset, sort the citation data in chronological order, and create a time series dataset;
[0115] S320, construct the citation curves of each failed research and successful research to dynamically capture the citation volume change trend of research results;
[0116] S330, based on the citation volume change trend, the contribution score of failed research, and semantic similarity, construct a potential prediction model to generate a potential score for failed research:
[0117] Potential score = B 1 · Citation change growth rate + B 2 · Contribution score + B 3 · Field relevance;
[0118] where B 1 、B 2 、B 3 are weight parameters.
[0119] S400, set the citation volume threshold and potential score threshold, screen the failed research whose current citation volume is lower than the citation volume threshold but whose potential score is higher than the potential score threshold, construct a reuse model for failed research, match and analyze the data content in the failed research with the current research needs, and generate a list of available data.
[0120] In this step, the research results with the potential for reuse are identified from the failed research by setting the citation threshold and the potential score threshold as screening conditions. When defining the citation threshold, it is necessary to combine the average citation volume, citation distribution law and time factors of the research results in the current field to ensure that the threshold setting can reasonably distinguish the results with low citation volume, while avoiding excluding the potential research that has not received enough attention in the early stage. The setting of the potential score threshold is based on the overall distribution of the potential score. Through statistical analysis, the key quantiles of the score, such as the median or high quantile interval, are determined to effectively screen out the failed research with high potential value. On this basis, the citation threshold and the potential score threshold are constructed as a screening formula with parallel conditions. This formula ensures that only those failed studies that are temporarily ignored by the academic community but perform well in the potential prediction model can enter the screening data set.
[0121] Next, we sorted out the failed studies that met the screening criteria to form a screening data set, laying the foundation for subsequent analysis. In this process, data sorting not only includes the standardized cleaning of basic information such as research titles, abstracts, and keywords, but also requires structured summarization of detailed contents such as experimental design, data integration, and research background of failed studies. This sorting method can ensure the integrity and consistency of the screening data set and provide high-quality data support for subsequent matching analysis.
[0122] On the basis of screening the data set, the semantic matching algorithm is introduced to construct a failure research reuse model. This model needs to comprehensively evaluate the matching degree between the content of the failure research and the current research needs. In the specific implementation, the similarity of research direction is mainly analyzed by natural language processing (NLP) technology. By analyzing the subject words, research objectives and problem statements of the failure research, and comparing them with the semantic embedding representation of the current research needs, the consistency of the research direction is quantified. The similarity of technical requirements combines the specific details such as experimental methods, technical parameters, variable selection in the failure research, and compares them item by item with the current research technical requirements to evaluate the adaptability of the failure research in technical implementation. The similarity of adaptation requirements focuses on the degree of fit between the failure research and the current research needs in terms of data, resources or experimental conditions, and evaluates its operability and resource efficiency. By calculating the weighted sum of the three similarities, the final matching degree of the reuse of the failure research is obtained, and the feasibility of its reuse is judged based on the matching degree.
[0123] Finally, integrate the research direction similarity, technical requirement similarity, and adaptation requirement similarity obtained from the reuse analysis, and combine the sorting results of the screened dataset to output the final list of available data. The data list should not only include the basic information of the failed research but also append detailed matching analysis results and suggested directions for reuse, such as being used as the theoretical basis for new research, a reference for experimental methods, or a source of data support. This output method ensures that researchers can quickly locate the required resources and clarify their possible application scenarios, improving research efficiency and resource utilization.
[0124] In this step, through the parallel screening formula of the citation threshold and the potential score threshold, S400 can accurately identify those failed studies that may be overlooked in the traditional evaluation system but actually have great scientific potential. This screening method effectively balances the academic influence (citation) and potential value (potential score) of research results, ensuring the scientificity and rationality of the screening results. Secondly, the introduction of the semantic matching algorithm enables the failed research reuse model to comprehensively evaluate failed studies from three dimensions: research direction, technical requirements, and adaptation requirements, avoiding the one-sidedness that may be caused by a single matching index. This multi-dimensional analysis provides higher accuracy and reliability for the reuse of failed studies. In addition, S400 outputs a structured list of available data by integrating the matching analysis results, providing intuitive and operable reuse guidance for researchers.
[0125] As Figure 4 shown, screening the failed studies with the current citation lower than the citation threshold but the potential score higher than the potential score threshold to construct a failed research reuse model, which specifically includes:
[0126] S410, define the citation threshold according to the average citation in the field, and define the potential score threshold according to the distribution status of the potential score;
[0127] S420, use the citation threshold and the potential score threshold as parallel conditions to form a screening formula:
[0128] Citation < citation threshold AND Potential score > potential score threshold;
[0129] S430, screen the failed studies according to the screening formula, obtain all the failed studies that meet the parallel conditions, and organize them into a screened dataset;
[0130] S440, establish a failed research reuse model based on the semantic matching algorithm, compare the content of the failed studies with the research requirements, including: research direction similarity, technical requirement similarity, and adaptation requirement similarity, and add them up to obtain the final research matching degree, and judge the feasibility of reusing the failed research data;
[0131] S450 integrates the feasibility data based on the similarity of research directions, the similarity of technical requirements, and the similarity of adaptation requirements, and outputs a list of available data.
[0132] Figure 6 The block diagram of the big data-supported scientific research achievement tracking system provided by the embodiment of the present invention is as Figure 6 shown, and the system includes:
[0133] A research data extraction module 100, configured to obtain failed research data and successful research data in the relevant field of the current research, fuse and construct a research achievement data set, analyze each failed research, and package and store the analyzed failure reasons, data details, and research background;
[0134] A knowledge graph construction module 200, configured to construct a knowledge graph of failed research and successful research based on the research achievement data set, analyze the dynamic association relationship between failed research and successful research, establish a causal inference model, quantitatively evaluate the contribution of failed research, and generate a contribution score for failed research;
[0135] A potential prediction module 300, configured to perform time series analysis on the citation data of failed research and successful research, establish a citation curve for each research, identify the change trend of the citation volume of each failed research, and establish a potential prediction model based on the citation curve, contribution score, and change trend, and output a potential score for each failed research to measure the future change trend of the citation volume;
[0136] A reuse analysis module 400, configured to set a citation volume threshold and a potential score threshold, screen failed research with a current citation volume lower than the citation volume threshold but a potential score higher than the potential score threshold, construct a failed research reuse model, match and analyze the data content in the failed research with the current research requirements, and generate a list of available data.
[0137] As Figure 7 shown, the research data extraction module 100 includes:
[0138] A research field feature recognition unit 110, configured to recognize the field features of the current research, and retrieve and identify the existing failed research data and successful research data in the associated field;
[0139] A research data classification unit 120, configured to classify the retrieved data according to the success or failure status, field, and theme of the research;
[0140] A research achievement data set fusion unit 130, configured to fuse the failed and successful research data into a structured research achievement data set, and classify and store the data according to the classification results;
[0141] The causal chain construction unit 140 is used to extract the reasons for failure, identify the key factors of failure according to the descriptive fields in the data, extract the experimental design content at the same time, construct the causal chain of the experimental process, and identify the research background field, and synchronously package and store the key factors of failure, the causal chain of the experimental process and the research background field in the research result dataset.
[0142] As Figure 8 shown, the knowledge graph construction module 200 includes:
[0143] The semantic similarity calculation unit 210 is used to perform text content analysis on the failed research results and the successful research results by using natural language processing, and calculate the semantic similarity between the failed research results and the successful research results;
[0144] The semantic association matrix construction unit 220 is used to construct a semantic association matrix representing the exercise intensity between the failed research and the successful research based on the semantic similarity;
[0145] The experimental process overlap analysis unit 230 is used to compare the causal chain of the experimental process of the failed research with the experimental process in the successful research, identify the overlapping parts in the experimental design and variable selection between the two, mark the experimental steps in the failed research that are used in the subsequent successful research, and generate a heuristic association model based on this;
[0146] The background field matching unit 240 is used to match the background fields of the failed research and the successful research, identify the potential inspiration of the failed research for subsequent research in the same background field, and dynamically update the heuristic association model;
[0147] The knowledge graph construction unit 250 is used to construct a knowledge graph of the failed research and the successful research with the failed research results and the successful research results as nodes and the outputs of the semantic association matrix and the heuristic association model as edges;
[0148] The contribution score quantification unit 260 is used to extract the mutually related nodes from the technical routes of the failed research and the successful research based on the experimental causal chain, establish an inference model, analyze whether the failed research constitutes a causal relationship with the successful research, and quantify the contribution degree of the failed research node, calculate the influence value of the failed research on the successful research, and calculate the contribution score.
[0149] As Figure 9 shown, the potential prediction module 300 includes:
[0150] The time series dataset creation unit 310 is used to read the citation data sources of each failed research and successful research in the research result dataset, sort the citation data in chronological order, and create a time series dataset;
[0151] The dynamic trend capture unit 320 is used to construct the citation curves of each failed study and successful study, and dynamically capture the changing trend of the citation volume of research results;
[0152] The potential score generation unit 330 is used to construct a potential prediction model based on the changing trend of citation volume, the contribution score of failed studies, and semantic similarity, and generate potential scores for failed studies.
[0153] As Figure 10 shown, the reuse analysis module 400 includes:
[0154] The potential score threshold definition unit 410 is used to define the citation volume threshold according to the field average citation volume, and define the potential score threshold according to the distribution state of potential scores;
[0155] The failed study screening unit 420 is used to combine the citation volume threshold and the potential score threshold as parallel conditions to form a screening formula;
[0156] The screened dataset sorting unit 430 is used to screen failed studies according to the screening formula, obtain all failed studies that meet the parallel conditions, and sort them into a screened dataset;
[0157] The matching degree calculation unit 440 is used to establish a reuse model for failed studies based on a semantic matching algorithm, compare the content of failed studies with research requirements, including: research direction similarity, technical requirement similarity, and adaptation requirement similarity, and add them up to obtain the final research matching degree to judge the feasibility of reusing failed study data;
[0158] The available data list integration unit 450 is used to integrate the feasible data according to the research direction similarity, technical requirement similarity, and adaptation requirement similarity respectively, and output an available data list.
[0159] It should be understood that although the steps in the flowcharts of the embodiments of the present invention are shown in sequence according to the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps does not have a strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in each embodiment may include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily have to be executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages does not necessarily have to be sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.
[0160] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0161] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0162] The above embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.
[0163] The above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for tracking scientific research results supported by big data, characterized in that: The method comprises: Obtain the failed research data and successful research data in the current research field, integrate them to build a research results data set, analyze each failed research, and package and store the failure reasons, data details and research background obtained from the analysis; Based on the research results dataset, we build a knowledge graph of failed research and successful research, analyze the dynamic correlation between failed research and successful research, establish a causal inference model, quantitatively evaluate the contribution of failed research, and generate a contribution score for failed research; Conduct time series analysis on the citation data of failed and successful studies, establish a citation curve for each study, identify the citation trend of each failed study, and establish a potential prediction model based on the citation curve, contribution score and change trend, output the potential score of each failed study, and measure the future citation trend; Set citation volume threshold and potential score threshold, screen out failed studies whose current citation volume is lower than the citation volume threshold but whose potential score is higher than the potential score threshold, build a failed research reuse model, match and analyze the data content in the failed studies with current research needs, and generate a list of available data.
2. The method according to claim 1, characterized in that The acquisition of failed research data and successful research data in related fields and the integration of them to construct a research results dataset specifically include: Identify the characteristics of the current research field, retrieve and identify the failed research data and successful research data in the related field; Categorize the retrieved data based on the success and failure status, domains, and themes of the research; Merge the failed and successful research data into a structured research results dataset, and store the data in categories based on the classification results; Based on the descriptive fields in the data, the causes of failure are extracted and the key factors of failure are identified. At the same time, the experimental design content is extracted, the causal chain of the experimental process is constructed, and the research background field is identified. The key factors of failure, the causal chain of the experimental process, and the research background field are synchronously packaged and stored in the research results dataset.
3. The method according to claim 2, characterized in that The contribution of the failed research is quantitatively evaluated to generate a failure research contribution score, which specifically includes: Use natural language processing to analyze the text content of failed research results and successful research results, and calculate the semantic similarity between failed research results and successful research results; Based on semantic similarity, a semantic correlation matrix is constructed to represent the intensity of practice between failed studies and successful studies. Compare the causal chain of the experimental process of the failed study with the experimental process of the successful study, identify the overlap between the two in experimental design and variable selection, mark the experimental steps in the failed study that were used in the subsequent successful study, and generate an heuristic association model based on this; Match the background fields of failed research and successful research, identify the potential inspiration of failed research to subsequent research in the same background field, and dynamically update the inspiration association model; The knowledge graph of failed research and successful research is constructed with failed research results and successful research results as nodes and the output of the semantic association matrix and the heuristic association model as edges. Based on the experimental causal chain, the nodes that are interrelated between the failed research and the successful research are extracted from the technical routes of the two, and a used inference model is established to analyze whether the failed research has a causal relationship with the successful research. The contribution of the failed research nodes is quantified, the impact of the failed research on the successful research is calculated, and the contribution score is calculated: Contribution score = A1·semantic similarity weight + A2·experimental causal chain coverage + A3·citation impact factor; Among them, A1, A2, and A3 are weight parameters.
4. The method according to claim 3, characterized in that The output is a potential score for each failed study, measuring the trend of future citation changes, including: Read the citation data source of each failed study and successful study in the research results dataset, sort the citation data in chronological order, and create a time series dataset; Construct citation curves for each failed and successful study to dynamically capture the changing trend of citation volume of research results; Based on the trend of citations, contribution scores of failed studies and semantic similarity, a potential prediction model is constructed to generate potential scores for failed studies: Potential score = B1·Citation change growth rate + B2·Contribution score + B3·Field relevance; Among them, B1, B2, and B3 are weight parameters.
5. The method according to claim 4, characterized in that The method of screening failed studies whose current citation volume is lower than the citation volume threshold but whose potential score is higher than the potential score threshold, and constructing a failed study reuse model specifically includes: Define the citation threshold based on the average citation volume in the field, and define the potential score threshold based on the distribution of potential scores; The citation threshold and potential score threshold are used as parallel conditions to form a screening formula: Citations < Citations Threshold AND Potential Score > Potential Score Threshold; The failed studies are screened according to the screening formula, all failed studies that meet the parallel conditions are obtained, and the data are organized into a screening data set; A failure research reuse model is established based on a semantic matching algorithm to compare the content of the failure research with the research needs, including: research direction similarity, technical requirements similarity, and adaptation requirements similarity, and the final research matching degree is obtained by adding them together to determine the feasibility of reusing the failure research data; The feasibility data is integrated according to the similarity of research directions, technical requirements and adaptation requirements, and a list of available data is output.
6. The scientific research results tracking system supported by big data is characterized by: The system comprises: The research data extraction module is used to obtain the failed research data and successful research data in the current research field, integrate them to build a research results data set, analyze each failed research, and package and store the failure reasons, data details and research background obtained from the analysis; The knowledge graph construction module is used to construct a knowledge graph of failed research and successful research based on the research results dataset, analyze the dynamic correlation between failed research and successful research, establish a causal inference model, quantitatively evaluate the contribution of failed research, and generate a contribution score for failed research; Potential prediction module, which is used to perform time series analysis on the citation data of failed and successful studies, establish a citation curve for each study, identify the citation trend of each failed study, and establish a potential prediction model based on the citation curve, contribution score and change trend, output the potential score of each failed study, and measure the future citation trend; The reuse analysis module is used to set the citation volume threshold and potential score threshold, screen out failed studies whose current citation volume is lower than the citation volume threshold but whose potential score is higher than the potential score threshold, build a failed research reuse model, match and analyze the data content in the failed studies with the current research needs, and generate a list of available data.
7. The system according to claim 6, characterized in that The research data extraction module includes: A research field feature recognition unit, which is used to recognize the field features of the current research, retrieve and recognize the failed research data and successful research data existing in the related field; Research data classification unit, which is used to classify the retrieved data according to the success and failure status, fields and themes of the research; A research result data set fusion unit is used to fuse the failed and successful research data into a structured research result data set, and classify and store the data according to the classification results; The causal chain construction unit is used to extract the causes of failure and identify the key factors of failure based on the descriptive fields in the data. At the same time, it extracts the experimental design content, constructs the causal chain of the experimental process, and identifies the research background field. The key factors of failure, the causal chain of the experimental process, and the research background field are packaged and stored in the research results dataset.
8. The system according to claim 7, characterized in that The knowledge graph construction module includes: A semantic similarity calculation unit, used for performing text content analysis on failed research results and successful research results by using natural language processing, and calculating the semantic similarity between the failed research results and the successful research results; A semantic association matrix construction unit, used to construct a semantic association matrix representing the intensity of practice between failed studies and successful studies based on semantic similarity; The experimental process overlap analysis unit is used to compare the causal chain of the experimental process in the failed study with the experimental process in the successful study, identify the overlap between the two in experimental design and variable selection, mark the experimental steps in the failed study that are used in the subsequent successful study, and generate an heuristic association model based on this; Background domain matching unit, used to match the background domains of failed research and successful research, identify the potential inspiration of failed research to subsequent research in the same background domain, and dynamically update the inspiration association model; A knowledge graph construction unit is used to construct a knowledge graph of failed research and successful research using failed research results and successful research results as nodes and the output of the semantic association matrix and the heuristic association model as edges; The contribution score quantification unit is used to extract the interrelated nodes of the failed research and the successful research from the technical routes based on the experimental causal chain, establish an inference model, analyze whether the failed research has a causal relationship with the successful research, and quantify the contribution of the failed research nodes, calculate the impact value of the failed research on the successful research, and calculate the contribution score.
9. The system according to claim 8, characterized in that The potential prediction module comprises: A time series data set creation unit is used to read the reference data source of each failed study and successful study in the research results data set, sort the reference data in chronological order, and create a time series data set; Dynamic trend capture unit, used to construct citation curves for each failed and successful study, and dynamically capture the changing trend of citation volume of research results; The potential score generation unit is used to build a potential prediction model based on the citation volume change trend, the contribution score of the failed research and the semantic similarity, and generate a potential score for the failed research.
10. The system according to claim 9, characterized in that The reuse analysis module comprises: A potential score threshold definition unit is used to define a citation threshold based on the average citation volume in the field and to define a potential score threshold based on the distribution status of the potential score; The failed research screening unit is used to combine the citation volume threshold and the potential score threshold as parallel conditions into a screening formula; A screening data set sorting unit is used to screen the failed studies according to the screening formula, obtain all failed studies that meet the parallel conditions, and sort them into a screening data set; The matching degree calculation unit is used to establish a failure research reuse model based on the semantic matching algorithm, compare the content of the failure research with the research requirements, including: research direction similarity, technical requirement similarity and adaptation requirement similarity, and add them together to obtain the final research matching degree to determine the feasibility of reusing the failure research data; The available data list integration unit is used to integrate the feasibility data according to the similarity of research directions, the similarity of technical requirements and the similarity of adaptation requirements, and output a list of available data.