Scientific and technological achievement multi-source information fusion evaluation method and system
By optimizing data source discovery strategies and machine learning-driven crawling technology, the problem of insufficient data sources in the multi-source information fusion evaluation system for water conservancy scientific and technological achievements has been solved, achieving more comprehensive data coverage and timely evaluation results. It also provides targeted strategies for the transformation of achievements, supporting decision-makers in making scientific decisions.
Patent Information
- Application Number
- CN202411596048.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-11
- Publication Date
- 2026-02-03
AI Technical Summary
The existing multi-source information fusion evaluation system for water conservancy scientific and technological achievements may not have sufficient data sources to comprehensively cover all information related to scientific and technological achievements, leading to one-sided evaluation results.
By continuously optimizing the data source discovery strategy, we use machine learning-driven crawling technology to crawl data from newly discovered data sources, utilize natural language processing technology to convert data from different sources into a unified format, and build an evaluation model for comprehensive assessment and dynamic analysis.
Ensuring comprehensive data sources, covering a wider range of information, reducing the bias of evaluation results, providing timeliness and dynamism, improving data processing efficiency and accuracy, providing targeted strategies and suggestions for results transformation, and supporting decision-makers in making informed decisions.
Smart Images

Figure CN121458077A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of scientific and technological achievement analysis technology, and in particular to a method and system for multi-source information fusion evaluation of scientific and technological achievements. Background Technology
[0002] Effective evaluation and application of scientific and technological achievements are crucial for resource management and decision support. In existing technologies, various data analysis and evaluation models are usually used to process and evaluate scientific and technological achievements. However, the efficiency of automatic data collection and processing is insufficient, and it often relies on a large amount of manual intervention, resulting in long data processing time and high costs.
[0003] The existing publication number CN118657418A discloses a multi-source information fusion evaluation system for water conservancy scientific and technological achievements. This system includes a data collection module capable of acquiring relevant data on water conservancy scientific and technological achievements from multiple data sources, including but not limited to publicly available literature, patent databases, and government data; a data processing module for cleaning, standardizing, integrating, and storing the collected data, ensuring data consistency, usability, and security; a dynamic evaluation model for comprehensively assessing and dynamically analyzing the input-output efficiency of water conservancy scientific and technological achievements; and an achievement transformation suggestion module for proposing targeted achievement transformation strategies and suggestions based on evaluation results and analysis data. The comprehensive evaluation and dynamic analysis tools provided by the system help decision-makers accurately assess the efficiency and productivity of water conservancy scientific and technological achievements. This precise analysis provides a clear quantitative basis for the input and output of scientific and technological achievements, thereby optimizing resource allocation and enhancing decision support. The system can efficiently collect and process information from multiple data sources, ensuring data timeliness and completeness. This automation reduces labor costs and improves the speed and quality of data processing.
[0004] However, the data sources of the aforementioned multi-source information fusion evaluation system for water conservancy scientific and technological achievements may not be sufficient to fully cover all information related to scientific and technological achievements, leading to a one-sided evaluation result. Summary of the Invention
[0005] The purpose of this invention is to provide a method and system for evaluating scientific and technological achievements through multi-source information fusion, which solves the problem that the data sources of existing evaluation systems for water conservancy scientific and technological achievements may not be sufficient to fully cover the information related to scientific and technological achievements, leading to the one-sidedness of the evaluation results.
[0006] To achieve the above objectives, this invention provides a method for evaluating scientific and technological achievements through multi-source information fusion, comprising the following steps:
[0007] Continuously optimize data source discovery strategies, including regularly evaluating and updating the list of data sources, exploring new data channels, and using machine learning-driven crawlers to crawl data from newly discovered data sources.
[0008] Data is processed using natural language processing technology to convert data from different sources into a unified format;
[0009] Construct an evaluation model to conduct a comprehensive assessment and dynamic analysis of scientific and technological achievements based on the processed data;
[0010] Based on the evaluation results and data analysis, we propose targeted strategies and suggestions for the transformation of research findings.
[0011] This includes continuously optimizing the data source discovery strategy, including regularly evaluating and updating the data source list, exploring new data channels, and using machine learning-driven crawlers to crawl data from newly discovered data sources. The steps also include:
[0012] Determine the initial list of data sources, categorize and label each data source, and clarify its attributes;
[0013] Set an evaluation cycle to assess the effectiveness of existing data sources, and optimize or eliminate data sources based on the evaluation results;
[0014] Learn about new data sources through various channels, use web crawling technology to pre-scan potential data sources, assess their data quality and relevance, and conduct preliminary data crawling tests on newly discovered data sources to verify their feasibility and effectiveness.
[0015] This includes continuously optimizing the data source discovery strategy, including regularly evaluating and updating the data source list, exploring new data channels, and using machine learning-driven crawlers to crawl data from newly discovered data sources. The steps also include:
[0016] Collect sample data, train machine learning models to identify and classify content in data sources, develop machine learning models, integrate machine learning models into crawler systems, and achieve intelligent data crawling and filtering.
[0017] Use machine learning-driven crawlers to scrape data from newly discovered data sources and store the scraped data in a database or data warehouse.
[0018] Establish a data source monitoring mechanism to track changes in data sources and data quality in real time, and adjust the data source discovery strategy and optimize crawler performance based on monitoring and feedback results;
[0019] Regularly update the data source list, add new data sources, remove invalid or low-quality data sources, and maintain the metadata information of the data sources.
[0020] The step of constructing an evaluation model and conducting a comprehensive assessment and dynamic analysis of scientific and technological achievements based on the processed data also includes:
[0021] Determine the key indicators for evaluating scientific and technological achievements;
[0022] Construct an evaluation model based on the selected indicators;
[0023] Integrate data from different sources and formats into the evaluation model to provide comprehensive analysis;
[0024] Test and adjust the evaluation model through historical data or expert verification to ensure its accuracy and reliability.
[0025] Among them, according to the evaluation results and analysis data, targeted achievement transformation strategies and suggestions are proposed. The steps also include:
[0026] Deeply analyze the evaluation results and identify the advantages and disadvantages of scientific and technological achievements;
[0027] Based on the analysis results, formulate transformation strategies for scientific and technological achievements;
[0028] Provide specific action suggestions to decision-makers;
[0029] Establish a feedback mechanism to adjust the evaluation model and transformation strategies according to the implementation results.
[0030] A multi-source information fusion evaluation system for scientific and technological achievements, including a collection module, a processing module, an evaluation module and a suggestion module. The collection module is connected to the processing module, the evaluation module is connected to the processing module, and the suggestion module is connected to the evaluation module;
[0031] The collection module is used to continuously optimize the data source discovery strategy, including regularly evaluating and updating the data source list, exploring new data channels, and using machine learning-driven crawlers to capture data from newly discovered data sources;
[0032] The processing module is used to process data through natural language processing technology and convert data from different sources into a unified format;
[0033] The evaluation module is used to construct an evaluation model and conduct comprehensive evaluation and dynamic analysis of scientific and technological achievements based on the processed data;
[0034] The suggestion module is used to propose targeted achievement transformation strategies and suggestions according to the evaluation results and analysis data.
[0035] This invention discloses a multi-source information fusion evaluation method and system for scientific and technological achievements. By continuously optimizing the data source discovery strategy, it ensures the comprehensiveness of data sources, covers a wider range of information, and reduces the bias of evaluation results. Employing machine learning-driven crawling technology, it can capture information from new data sources in real time, making the evaluation results more timely and dynamic. Natural language processing technology is used to convert data from different sources into a unified format, facilitating analysis and comparison. Integrating machine learning models into the crawling system enables intelligent data capture and filtering, improving the efficiency and accuracy of data processing. The evaluation model is tested and adjusted using historical data or expert verification to ensure the accuracy and reliability of the evaluation results. The fusion of data from different sources and formats provides comprehensive analysis, making the evaluation results more integrated and in-depth. Based on the evaluation results and analyzed data, targeted achievement transformation strategies and suggestions are proposed, providing customized decision support for decision-makers. A feedback mechanism is established to adjust the evaluation model and transformation strategy based on implementation results, enabling the system to adapt to rapidly changing environments and technological developments. The metadata information of the data sources is maintained to ensure the compliance and security of the data sources, while the data source list is updated regularly to maintain the freshness and relevance of the data. The provided strategies and suggestions for commercializing scientific and technological achievements can help decision-makers better understand the potential value and market applications of these achievements, thereby enabling them to make more informed decisions. Attached Figure Description
[0036] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.
[0037] Figure 1 This is a flowchart of the steps of the multi-source information fusion evaluation method for scientific and technological achievements according to the first embodiment of the present invention.
[0038] Figure 2 This is a flowchart illustrating the steps of a continuous optimization data source discovery strategy according to the first embodiment of the present invention, including regularly evaluating and updating the data source list, exploring new data channels, and using machine learning-driven crawlers to crawl data from newly discovered data sources.
[0039] Figure 3 This is a principle block diagram of the multi-source information fusion evaluation system for scientific and technological achievements according to the second embodiment of the present invention.
[0040] In the diagram: 201 - Acquisition module, 202 - Processing module, 203 - Evaluation module, 204 - Suggestion module. Detailed Implementation
[0041] The embodiments of the present invention are described in detail below. Examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, but should not be construed as limiting the present invention.
[0042] The first embodiment of this application is as follows:
[0043] Please see Figure 1 and Figure 2 ,in, Figure 1 This is a flowchart of the steps of the multi-source information fusion evaluation method for scientific and technological achievements according to the first embodiment of the present invention. Figure 2 This is a flowchart illustrating the steps of a continuous optimization data source discovery strategy according to the first embodiment of the present invention, including regularly evaluating and updating the data source list, exploring new data channels, and using machine learning-driven crawlers to crawl data from newly discovered data sources.
[0044] This invention provides a method for evaluating scientific and technological achievements through multi-source information fusion, comprising the following steps:
[0045] S101: Continuously optimize data source discovery strategies, including regularly evaluating and updating the list of data sources, exploring new data channels, and using machine learning-driven crawlers to crawl data from newly discovered data sources.
[0046] S1011: Determine the initial list of data sources, classify and label each data source, and clarify its attributes;
[0047] S1012: Set an evaluation cycle to evaluate the effectiveness of existing data sources, and optimize or eliminate data sources based on the evaluation results;
[0048] S1013: Learn about new data sources through various channels, use web crawling technology to pre-scan potential data sources, evaluate their data quality and relevance, and conduct preliminary data crawling tests on newly discovered data sources to verify their feasibility and effectiveness.
[0049] S1014: Collect sample data, train machine learning models to identify and classify content in data sources, develop machine learning models, integrate machine learning models into crawler systems, and achieve intelligent data crawling and filtering.
[0050] S1015: Use machine learning-driven crawlers to scrape data from newly discovered data sources and store the scraped data in a database or data warehouse;
[0051] S1016: Establish a data source monitoring mechanism to track changes in data sources and data quality in real time, and adjust the data source discovery strategy and optimize crawler performance based on monitoring and feedback results;
[0052] S1017: Regularly update the data source list, add new data sources, remove invalid or low-quality data sources, and maintain the metadata information of data sources.
[0053] Specifically, an initial list of data sources is established, and each data source is meticulously categorized and labeled, clearly defining its attributes such as data type, update frequency, and reliability. This step is fundamental to the entire process, providing a starting point for subsequent data evaluation and optimization. Based on the initial data source list, an evaluation cycle is set to periodically assess the effectiveness of existing data sources. The evaluation cycle takes into account the update frequency and importance of the data sources, using a formula... To identify and ensure timely response to any changes in data sources, new data sources should be identified through various channels, and potential data sources should be pre-scanned using web crawling techniques to assess their data quality and relevance. Preliminary data crawling tests should be conducted on newly discovered data sources to verify their feasibility and effectiveness. This step is crucial for ensuring the continuous updating and expansion of data sources. Some key channels include: 1. Academic databases and journals: including but not limited to IEEE Xplore, Science Direct, Springer Link, Wiley Online Library, etc., which provide a wealth of academic papers and research results. 2. Government and public institution websites: Government websites frequently publish statistics, policy documents, and research reports, which are important channels for understanding public data sources. 3. Professional forums and communities: such as Research Gate, LinkedIn Groups, and professional forums, discussions and sharing in these communities can reveal new data sources. 4. Social media platforms: Hashtags, discussion groups, and official accounts on social media platforms such as Twitter, Facebook, and Weibo can provide real-time data and information. 5. Industry reports and market research: Industry analysis reports and market research reports typically contain the latest industry data and trend analysis. 6. Patent Databases: Such as the United States Patent and Trademark Office (USPTO), the European Patent Office (EPO), and the China National Intellectual Property Administration (CNIPA), providing patent information and technological innovation data. 7. Open Data Platforms: Open datasets provided by governments and non-governmental organizations, such as data.gov and Google Public Data Explorer. 8. Conference Proceedings and Academic Conferences: Proceedings and proceedings of academic and industry conferences, often the venues for the release of the latest research findings. 9. News Media and Press Releases: News websites and press release platforms, such as PRNewswire and Business Wire, providing the latest news reports and company updates. 10. Technical Blogs and Personal Websites: Technical blogs by domain experts and enthusiasts, as well as websites of individual researchers, frequently sharing valuable data and insights. 11. Online Courses and Educational Platforms: Online educational platforms such as Coursera and edX, providing course materials and research data. 12. Libraries and Information Centers: Databases provided by libraries and resources in information centers, which typically offer professional information collection and organization services. Data quality and relevance assessment: Relevance analysis and the data quality scoring formula Q = α·accuracy + β·completeness + γ·consistency are used. To improve the intelligence of data crawling, sample data is collected and machine learning models are trained to identify and classify content in the data sources.Developing and integrating machine learning models into a web crawler system enables intelligent data crawling and filtering, significantly improving the efficiency and accuracy of data processing. The process involves three main steps: Step 1: Collecting Sample Data. Sufficient sample data is collected from identified data sources to train the machine learning model. This includes defining data collection rules based on the data source's attributes and the required data type; implementing data crawling using a basic web crawler according to the rules; and cleaning the crawled data by removing irrelevant content and noise, retaining valuable information. Step 2: Preprocessing and Feature Extraction. The collected sample data is preprocessed, and features for training are extracted. This includes text preprocessing (word segmentation, stop word removal, stemming, etc.), feature extraction (using methods such as TF-IDF, Word2Vec, and BERT), and feature selection (using correlation analysis and mutual information to select the most informative features). Step 3: Training the Machine Learning Model. A suitable machine learning algorithm is selected to train the model, and parameters are optimized. This involves choosing a suitable model based on the data characteristics, such as Support Vector Machine (SVM), Random Forest, or Neural Networks. Training and Test Set Splitting: Divide the dataset into training and test sets, typically in a 70% training and 30% test ratio. Model Training: Train the model using the training set data. Parameter Optimization: Optimize model parameters using methods such as cross-validation and grid search. Step 4: Model Evaluation: Evaluate the model's performance to ensure its accuracy and generalization ability. Performance Evaluation: Evaluate model performance using metrics such as accuracy, recall, and F1 score. Confusion Matrix: Generate a confusion matrix to visualize model performance. Model Tuning: Adjust the model structure or parameters based on the evaluation results. Step 5: Integration into the Web Crawler System: Integrate the trained model into the web crawler system to achieve intelligent data crawling and filtering. Model Deployment: Deploy the model into the web crawler system, enabling it to classify and filter real-time crawled data. Real-time Prediction: After the crawler crawls data, the model makes real-time predictions to determine whether to store or further process the data. Feedback Loop: Continuously adjust and optimize the model based on prediction results and subsequent human review feedback. Step 6: Continuous Optimization: Continuously optimize the model and crawler strategy based on the actual operation of the web crawler system. Performance Monitoring: Monitor the real-time performance of the model, such as accuracy and response time. Model Updates: Regularly update the model with newly collected data to adapt to changes in data distribution. Strategy Adjustment: Adjust the crawling strategy based on performance monitoring results and business needs. Use the trained model to classify and filter new data sources, store the crawled data in a database (SQL or NoSQL), and use data quality monitoring tools such as Apache Kafka for real-time data stream monitoring. Based on the monitoring results, adjust the data source discovery strategy using optimization algorithms.Update the data source list using the formula N = N - low quality sources + new sources, updating the metadata information of each data source, such as update time and data format.
[0054] S102: Data processing is performed using natural language processing technology to convert data from different sources into a unified format;
[0055] Specifically, Step 1: Extract raw text data from multiple data sources and perform data cleaning to remove useless information such as HTML tags, special characters, blank lines, etc., and identify and process noise in the data. This includes data cleaning: using regular expressions and text processing libraries (such as Python's Beautiful Soup) to remove HTML tags and special characters. Noise processing: identifying outliers through statistical analysis (such as box plots) and processing them. Step 2: Segment the collected text data into manageable units, such as sentences or paragraphs. This includes sentence segmentation: using sentence segmentation tools in natural language processing libraries (such as NLTK or spaCy) to segment sentences based on punctuation marks such as periods and question marks. Paragraph segmentation: segmenting the text into paragraphs based on line breaks or specific delimiters. Step 3: Identify entities in the text, such as names, locations, and organizations, and label the identified entities with category labels. This includes entity recognition: using named entity recognition (NER) models, such as the BiLSTM-CRF model, to identify entities in the text. Step 4: Tag the words in the text with predefined category labels, such as PER (person name), LOC (location), and ORG (organization). This includes part-of-speech tagging (POS) of each word, such as noun, verb, adjective, etc. This involves using a pre-trained POS tagging model, such as HMM or CRF, to tag words with POS. Step 5: Analyze the grammatical dependencies between words in the sentence. This involves using dependency parsing tools, such as the Stanford dependency parser, to construct the dependency parsing tree of the sentence. Step 6: Identify the semantic roles of each component in the sentence, such as agent, patient, instrument, etc. This involves using semantic role labeling models, such as frame-based models, to identify the semantic roles of sentence components. Step 7: Convert the data in the text to a standard format, such as date, currency, and unit of measurement. This involves data normalization, using regular expressions and custom conversion rules to convert non-standard format data to a standard format. Step 8: Evaluate the sentiment tendency in the text, such as positive, negative, or neutral. The process includes: Step 9: Sentiment analysis: Using sentiment analysis models, such as dictionary-based methods or deep learning models, to assess the sentiment tendency of the text. Step 10: Using the LDA algorithm to identify topics in the text and classify the text data into predefined topic categories. This includes topic modeling: Using the LDA algorithm, by setting the number of topics (K) and the number of iterations, to identify the topic distribution in the text. Step 11: Classifying the text data into predefined topic categories. This includes integrating the processed data into a unified data model. This includes data fusion: Integrating data from different sources and formats into a unified data model, such as a relational database or a NoSQL database. Step 12: Verifying the accuracy and consistency of the processed data. This includes data validation: Using data quality assessment tools, such as data integrity checks and consistency checks, to verify the accuracy and consistency of the data.Step 12: Store the processed and transformed data in a database or data warehouse. This includes data storage: storing the cleaned, labeled, and transformed data in a database or data warehouse for subsequent analysis and application. These steps ensure that data from different sources is effectively processed and transformed, enabling subsequent evaluation models to accurately conduct comprehensive assessments and dynamic analyses of scientific and technological achievements.
[0056] S103: Construct an evaluation model to conduct a comprehensive evaluation and dynamic analysis of scientific and technological achievements based on the processed data;
[0057] Specifically, key indicators for evaluating scientific and technological achievements are identified through literature review, expert consultation, policy analysis, and industry standards. These indicators may include, but are not limited to, technological advancement, market potential, social impact, environmental impact, and economic feasibility. Weights are assigned to each indicator to reflect its relative importance in the evaluation system. Weights can be determined using methods such as expert scoring, the Analytic Hierarchy Process (AHP), or entropy weighting. Since different indicators may have different dimensions and magnitudes, they need to be standardized to ensure comparability within the model. An appropriate evaluation model is selected based on the evaluation purpose and indicator characteristics. Common models include Multi-Criterion Decision Analysis (MCDM), Data Envelopment Analysis (DEA), and Artificial Neural Networks (ANN). The model's inputs (i.e., the selected indicators) and outputs (i.e., the evaluation results) are defined. Model construction may involve parameter estimation, algorithm design, and software implementation. Sensitivity analysis is performed on the model to assess the impact of changes in different indicators on the evaluation results, ensuring the model's robustness. Data fusion techniques, such as Principal Component Analysis (PCA), factor analysis, or machine learning algorithms, are used to integrate multi-source data into the evaluation model. Ensure the merged data is logically consistent, free of contradictions or redundant information. Backtest the model using historical data to evaluate its predictive power and accuracy. This may involve methods such as cross-validation and model fit testing. Verify the model's logical rationality and practical application value through expert review. Experts may provide feedback on model structure, parameter settings, and result interpretation. Adjust model parameters or structure based on test results and expert feedback to improve the model's accuracy and reliability. This may involve re-estimating parameters, adjusting weights, or optimizing the algorithm. Through these steps, a scientific, reasonable, and reliable evaluation model can be constructed, enabling comprehensive evaluation and dynamic analysis of scientific and technological achievements. This not only helps decision-makers understand the comprehensive value of scientific and technological achievements but also provides a scientific basis for their further development and application.
[0058] S104: Based on the evaluation results and analysis data, propose targeted strategies and suggestions for the transformation of research results.
[0059] Specifically, the evaluation model outputs are interpreted in detail to identify the performance of scientific and technological achievements across various evaluation indicators. Through comparative analysis, the relative advantages of the achievements are identified, such as technological innovation and market competitiveness. Potential risks and shortcomings are identified, such as technological defects and market entry barriers. A comprehensive SWOT analysis is employed to systematically analyze the strengths, weaknesses, opportunities, and threats of the achievements. Based on these strengths and weaknesses, specific transformation strategies are developed, such as technology improvement plans and market development strategies. The resources required for implementing these strategies, including funding, human resources, and technology, are determined and allocated rationally. Potential risks during the transformation process are assessed, and corresponding risk management measures and response strategies are developed. Based on the analysis results and transformation strategies, specific action recommendations are provided to decision-makers, such as R&D investment and market promotion. Decision support materials, including data analysis reports, strategic plans, and risk assessment reports, are provided to help decision-makers make informed decisions. Effective communication with decision-makers ensures that the recommendations are understood and accepted, and resources are coordinated to support their implementation. Establish a monitoring mechanism to track the implementation and effectiveness of the conversion strategy. Collect feedback information during implementation, including successful experiences, existing problems, and improvement suggestions. Adjust the evaluation model based on feedback and new data to improve its accuracy and applicability. Optimize the conversion strategy based on implementation results and market changes to adapt to the ever-changing environment.
[0060] By continuously optimizing data source discovery strategies, we ensure the comprehensiveness of data sources, covering a wider range of information and reducing the bias of evaluation results. Employing machine learning-driven web crawling technology, we can capture information from new data sources in real time, making evaluation results more timely and dynamic. Natural language processing technology is used to convert data from different sources into a unified format, facilitating analysis and comparison. Integrating machine learning models into the web crawling system enables intelligent data capture and filtering, improving the efficiency and accuracy of data processing. We test and adjust evaluation models using historical data or expert validation to ensure the accuracy and reliability of evaluation results. By fusing data from different sources and formats, we provide comprehensive analysis, making evaluation results more integrated and in-depth. Based on evaluation results and analyzed data, we propose targeted technology transfer strategies and recommendations, providing customized decision support for decision-makers. We establish a feedback mechanism to adjust evaluation models and transfer strategies based on implementation results, enabling the system to adapt to rapidly changing environments and technological developments. We maintain the metadata information of data sources to ensure compliance and security, while regularly updating the data source list to keep the data fresh and relevant. The provided technology transfer strategies and recommendations help decision-makers better understand the potential value and market applications of scientific and technological achievements, thereby making more informed decisions.
[0061] The second embodiment of this application is as follows:
[0062] Based on the first embodiment, please refer to Figure 3 ,in, Figure 3 This is a principle block diagram of the multi-source information fusion evaluation system for scientific and technological achievements according to the second embodiment of the present invention.
[0063] This embodiment of a multi-source information fusion evaluation system for scientific and technological achievements includes a data acquisition module 201, a processing module 202, an evaluation module 203, and a suggestion module 204. The data acquisition module 201 is connected to the processing module 202, the evaluation module 203 is connected to the processing module 202, and the suggestion module 204 is connected to the evaluation module 203.
[0064] The acquisition module 201 is used to continuously optimize the data source discovery strategy, including regularly evaluating and updating the data source list, exploring new data channels, and using machine learning-driven crawlers to crawl data from newly discovered data sources.
[0065] The processing module 202 is used to process data using natural language processing technology and convert data from different sources into a unified format.
[0066] The evaluation module 203 is used to construct an evaluation model and to conduct a comprehensive evaluation and dynamic analysis of scientific and technological achievements based on the processed data.
[0067] The suggestion module 204 is used to propose targeted results transformation strategies and suggestions based on the evaluation results and analysis data.
[0068] Using the multi-source information fusion evaluation system for scientific and technological achievements in this embodiment, when the system starts running, the acquisition module 201 first starts, executes the continuous optimization strategy of the data source, regularly evaluates and updates the data source list, explores new data channels, and automatically crawls data from newly discovered data sources using machine learning-driven crawling technology to ensure the comprehensiveness and timeliness of the data. The acquired data is transmitted to the processing module 202, where natural language processing technology is used to clean, format, and standardize the data, converting it into a unified format to facilitate subsequent evaluation and analysis. The processed data enters the evaluation module 203, where an evaluation model is constructed, and a comprehensive evaluation and dynamic analysis of the scientific and technological achievements are performed based on the processed data. The evaluation results provide the performance of the scientific and technological achievements on multiple key indicators. The evaluation results and analysis data are sent to the suggestion module 204, which proposes targeted achievement transformation strategies and suggestions based on the evaluation results. These suggestions aim to help decision-makers understand the potential value of scientific and technological achievements and guide the actual transformation and application of scientific and technological achievements.
[0069] The system, through multi-source information fusion, can comprehensively evaluate multiple aspects of scientific and technological achievements, providing more comprehensive analytical results. The system can update data sources in real time and dynamically analyze the latest developments in scientific and technological achievements, maintaining the timeliness of evaluation results. The system employs machine learning technology, improving the intelligence level of data collection and processing and reducing human intervention. The application of natural language processing technology improves the accuracy of data processing, while a scientific evaluation model ensures the reliability of evaluation results. The specific action suggestions provided by the system offer strong decision support to decision-makers, helping them make more informed decisions. The system can adjust the evaluation model and transformation strategy based on feedback and market changes, exhibiting excellent flexibility and adaptability. The system's automated processes improve the efficiency of data processing and evaluation, reducing the need for human resources. The system design allows for continuous optimization of data sources and evaluation models, ensuring performance and effectiveness in long-term operation.
[0070] The above-disclosed embodiments are merely one or more preferred embodiments of this application and should not be construed as limiting the scope of this application. Those skilled in the art can understand that all or part of the processes for implementing the above embodiments and equivalent changes made in accordance with the claims of this application still fall within the scope of this application.
Claims
1. A method for evaluating scientific and technological achievements through multi-source information fusion, characterized in that, Includes the following steps: Continuously optimize data source discovery strategies, including regularly evaluating and updating the list of data sources, exploring new data channels, and using machine learning-driven crawlers to crawl data from newly discovered data sources. Data is processed using natural language processing technology to convert data from different sources into a unified format; Construct an evaluation model to conduct a comprehensive assessment and dynamic analysis of scientific and technological achievements based on the processed data; Based on the evaluation results and data analysis, we propose targeted strategies and suggestions for the transformation of research findings.
2. The method for multi-source information fusion evaluation of scientific and technological achievements as described in claim 1, characterized in that, Continuously optimize the data source discovery strategy, including regularly evaluating and updating the data source list, exploring new data channels, and using machine learning-driven crawlers to crawl data from newly discovered data sources. The steps also include: Determine the initial list of data sources, categorize and label each data source, and clarify its attributes; Set an evaluation cycle to assess the effectiveness of existing data sources, and optimize or eliminate data sources based on the evaluation results; Learn about new data sources through various channels, use web crawling technology to pre-scan potential data sources, assess their data quality and relevance, and conduct preliminary data crawling tests on newly discovered data sources to verify their feasibility and effectiveness.
3. The method for evaluating scientific and technological achievements through multi-source information fusion as described in claim 2, characterized in that, Continuously optimize the data source discovery strategy, including regularly evaluating and updating the data source list, exploring new data channels, and using machine learning-driven crawlers to crawl data from newly discovered data sources. The steps also include: Collect sample data, train machine learning models to identify and classify content in data sources, develop machine learning models, integrate machine learning models into crawler systems, and achieve intelligent data crawling and filtering. Use machine learning-driven crawlers to scrape data from newly discovered data sources and store the scraped data in a database or data warehouse. Establish a data source monitoring mechanism to track changes in data sources and data quality in real time, and adjust the data source discovery strategy and optimize crawler performance based on monitoring and feedback results; Regularly update the data source list, add new data sources, remove invalid or low-quality data sources, and maintain the metadata information of the data sources.
4. The multi-source information fusion evaluation method for scientific and technological achievements as described in claim 3, characterized in that, The steps of processing data using natural language processing technology to convert data from different sources into a unified format also include: Collect raw text data from multiple data sources, perform data cleaning, remove useless information, and identify and process noise in the data; The collected text data is divided into manageable units; Identify entities in the text and label the identified entities with category tags; Perform part-of-speech tagging on each word in the text; Analyze the grammatical dependencies between words in a sentence; Identify the semantic roles of each component in a sentence; Convert the data in the text to a standard format; Assess the sentiment in the text; The LDA algorithm is used to identify topics in text and classify text data into predefined topic categories; The processed data is integrated into a unified data model; Verify the accuracy and consistency of the processed data, and store the processed and transformed data in a database or data warehouse.
5. The multi-source information fusion evaluation method for scientific and technological achievements as described in claim 4, characterized in that, The process of constructing an evaluation model and conducting a comprehensive assessment and dynamic analysis of scientific and technological achievements based on the processed data also includes: Determine the key indicators for evaluating scientific and technological achievements; Construct an evaluation model based on the selected indicators; Integrate data from different sources and formats into the evaluation model to provide a comprehensive analysis; Test and adjust the evaluation model through historical data or expert verification to ensure its accuracy and reliability.
6. The method for evaluating scientific and technological achievements through multi-source information fusion as described in claim 5, characterized in that, According to the evaluation results and analysis data, propose targeted achievement transformation strategies and suggestions. The steps also include: Deeply analyze the evaluation results to identify the advantages and disadvantages of scientific and technological achievements; Based on the analysis results, formulate transformation strategies for scientific and technological achievements; Provide specific action suggestions to decision-makers; Establish a feedback mechanism to adjust the evaluation model and transformation strategies according to the implementation results.
7. A multi-source information fusion evaluation system for scientific and technological achievements, applicable to the multi-source information fusion evaluation method for scientific and technological achievements as described in any one of claims 1 to 6, characterized in that it includes a collection module, a processing module, an evaluation module and a suggestion module. The collection module is connected to the processing module, the evaluation module is connected to the processing module, and the suggestion module is connected to the evaluation module; the collection module is used to continuously optimize the data source discovery strategy, including regularly evaluating and updating the data source list, exploring new data channels, and using machine learning-driven crawlers to capture data from newly discovered data sources; the processing module is used to process data through natural language processing technology and convert data from different sources into a unified format; the evaluation module is used to construct an evaluation model and conduct comprehensive evaluation and dynamic analysis of scientific and technological achievements based on the processed data; the suggestion module is used to propose targeted achievement transformation strategies and suggestions according to the evaluation results and analysis data.