An intelligent auxiliary method, system and device for scientific research and innovation topic selection

By classifying and preprocessing scientific literature by field, extracting keywords using the TextRank algorithm, and displaying a list of recent popular topics, this approach solves the problem of intelligent and personalized selection of research topics in existing technologies, thereby improving the efficiency and quality of scientific research innovation.

CN119537593BActive Publication Date: 2026-02-03STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411371501.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-29
Publication Date
2026-02-03
Estimated Expiration
2044-09-29

AI Technical Summary

Technical Problem

Existing technologies are insufficient to provide personalized and intelligent research topic selection, making it difficult for researchers to accurately identify valuable information. Furthermore, existing retrieval technologies cannot meet personalized needs, which can easily lead to resource waste and misjudgment.

Method used

By acquiring scientific and technological literature, classifying and preprocessing it by field, extracting keywords using the TextRank algorithm, and combining it with visualization technology to display a list of recent popular topics, users can be assisted in selecting research and innovation topics.

Benefits of technology

It enables personalized research topic selection, improves the efficiency and quality of scientific research innovation, and helps researchers quickly identify the most promising research directions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119537593B_ABST
    Figure CN119537593B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of intelligent auxiliary method, system and device for scientific research innovation topic selection, including obtaining scientific and technological literature, scientific and technological literature is handled, using TextRank algorithm is extracted to the data set after pre-processing Keyword, according to keyword paper / patent standard and so on scientific and technological literature is classified, then the quantity of paper / patent standard and so on scientific and technological literature under each category is calculated, according to quantity, get hot word new word list, according to hot word new word list, get recent popular theme list, this table is shown to user, intelligent auxiliary user carries out scientific research innovation topic selection.The present application plays the auxiliary decision-making role of artificial intelligence and big data technology in scientific research topic, from mass scientific research literature, data and trend, quickly filters out the most potential and prospect topic, so that scientific research personnel and innovation team can more targetedly select research direction, avoid blind exploration, improve scientific and technological innovation efficiency and achievement quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of natural language processing and data mining technology, specifically relating to an intelligent auxiliary method, system, and device for selecting research and innovation topics. Background Technology

[0002] With the development of the times and the progress of science and technology, the division of labor in society has become increasingly refined, and everyone is undertaking increasingly heavy labor in order to survive and develop. Therefore, science and technology have become an indispensable tool for contemporary human production and life.

[0003] Research topic selection is the primary issue in scientific and technological innovation. In scientific research practice, accurately identifying and tackling research topics that align with national development strategies and future trends in scientific and technological innovation is a crucial indicator of achieving high-level scientific and technological self-reliance. To take the initiative and be proactive in scientific and technological innovation, it is essential to establish a scientific, independent, and open research topic selection mechanism, improve research topic selection capabilities, and enhance the planning and implementation of major research projects.

[0004] Currently, my country's economic and social development, energy system optimization and upgrading, and the construction of a new power system face many practical problems that need to be addressed. This requires accelerating the development of technologies that can be quickly mastered and solve problems promptly, while strategic technologies requiring sustained effort should be planned in advance. Power grid companies need to increase investment in basic and applied research, strengthen industry-academia-research cooperation, establish a research topic selection mechanism that transforms major practical problems into scientific and technological issues, identify scientific and technological problems from the needs of industrial innovation, and break through a number of fundamental principles and key core technologies.

[0005] The emergence and development of science are inseparable from the division of disciplines. With the development of society and the economy, and the vigorous promotion of modern information technology and computer networks, effective communication platforms and technical support have been provided for scientific research. At the same time, the pace of knowledge updating in today's society is constantly accelerating, leading to a gradual decline in people's in-depth understanding of knowledge, while their breadth of knowledge is increasing. This has resulted in professionals reaching a certain peak in their research within their own fields; to achieve further success, they must venture into new technological fields. However, faced with a complex knowledge system and a large amount of research results, researchers often don't know where to begin to find suitable solutions. In the actual application process of the applicant of this invention, it was found that the current solutions to the above problems mainly rely on human experience summaries or automatic recommendations from search software to determine a relatively suitable research direction. However, these methods still have certain limitations, mainly reflected in the following aspects: (1) These solutions are all automatically given by programs designed by humans, without considering personal interests and concerns, and cannot meet personalized needs; (2) Most of the current technical solutions only provide analysis and screening based on existing results, which cannot directly reflect the user's true intentions, and can only provide some reference opinions, and cannot play an intelligent guiding role; (3) The information provided by the existing technical solutions is mostly obtained through simple text reading, which makes it easy for operators to ignore some key information points or even mistakenly regard relevant information as important, thus causing judgment errors, ultimately affecting the progress of subsequent research or causing unnecessary waste of resources. In addition, although the search technology used in the existing technology can achieve fast data retrieval and display of relevant content, it is difficult to accurately identify truly valuable information and make effective use of it. Due to the large differences between different fields, the same problem may have different solutions in different situations; and there may be multiple potential solutions available for the same problem. In summary, there is currently a lack of a comprehensive data query service system that can be automated and has intelligent features. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of the existing technology and provide an intelligent auxiliary method, system and device for selecting scientific research and innovation topics.

[0007] The objective of this invention can be achieved through the following technical solutions:

[0008] This invention provides an intelligent auxiliary method for selecting research and innovation topics, comprising the following steps:

[0009] Obtain scientific and technological literature and their publication dates as the raw dataset;

[0010] The original dataset is classified into three categories based on literature: paper dataset, patent dataset, and technical standard dataset.

[0011] The paper datasets were classified by domain according to the Chinese Library Classification, resulting in paper datasets for each domain.

[0012] The patent datasets are classified by domain according to patent type to obtain patent datasets for each domain.

[0013] The technical standard datasets are classified by domain according to standard type to obtain technical standard datasets for each domain.

[0014] The paper datasets, patent datasets, and technical standard datasets from various fields are merged by field category to obtain the datasets for each field.

[0015] Extract titles and abstracts from paper and patent datasets in various fields, and extract titles and scopes from technical standard datasets. Use the extracted data as the dataset to be processed in that field.

[0016] Preprocess the dataset to be processed;

[0017] The TextRank algorithm was used to extract keywords from the preprocessed dataset.

[0018] The extracted keywords were used as the research topics of their corresponding scientific and technological literature.

[0019] Statistical analysis was conducted on all topics in each field based on publication time and research topic to obtain a list of hot words and a list of new words for each field. The list of hot words includes the topics of literature published within the first preset time period and their publication frequency, while the list of new words includes the topics of literature published only within the second preset time period and their publication frequency.

[0020] The lists of trending words and new words are sorted according to their publication frequency;

[0021] Obtain the user's area of ​​interest (as input);

[0022] Extract the first preset value position of the literature topics and their publication frequency in the list of new words in the field of interest, find the publication frequency of their literature topics in the list of hot words in the field of interest, add the frequency to the extracted list, and obtain a preliminary list of recent hot topics.

[0023] The initial list of recent popular topics is sorted by the frequency of publication within a first preset time period to obtain the current list of popular topics.

[0024] Visualization technology is used to display a list of recent popular topics to users, intelligently assisting them in selecting research and innovation topics.

[0025] Furthermore, the scientific and technological literature includes: academic journals, academic papers, conference papers, scientific and technological achievements, patent documents, and technical standard documents.

[0026] Furthermore, the paper dataset includes academic journals, academic papers, conference papers, and scientific and technological achievements; the patent dataset includes patent documents; and the technical standard dataset includes technical standard documents.

[0027] Furthermore, the data to be processed includes unstructured data such as Word and PDF files.

[0028] Furthermore, the preprocessing of the dataset to be processed includes the following steps:

[0029] Read the data to be processed;

[0030] Convert unstructured data into structured text using a Python preprocessing script;

[0031] Remove blank text fields and individual abnormal text fields generated during the extraction process from the structured data text. Use a basic stop word library to supplement some stop words.

[0032] Furthermore, the keyword extraction from the preprocessed dataset using the TextRank algorithm includes the following steps:

[0033] Text preprocessing includes word segmentation, stop word removal, and part-of-speech tagging, converting the text into a format suitable for algorithm processing;

[0034] Building a graph model: Using words or sentences in the processed text as nodes, a graph model is built based on the relationships between words or sentences;

[0035] Calculate node weights: Calculate the weight of each node in the graph using the iterative approach of the PageRank algorithm;

[0036] Sorting and Extraction: Sort nodes according to their weights, and select the second-to-last node after sorting as the keyword of the text.

[0037] Furthermore, the relationships between the nodes include co-occurrence relationships and semantic similarity, and the weight of a node is determined by the weights of other nodes and the strength of the relationships between them.

[0038] Furthermore, the first preset time period is longer than the second preset time period, the first preset time period is the past five years, and the second preset time period is the past year.

[0039] Another aspect of this invention provides an intelligent auxiliary system for selecting research and innovation topics, comprising: a data acquisition module, a data classification module, a data extraction module, a data preprocessing module, a keyword extraction module, a hot word and new word processing module, an input module, and a visualization module. The keyword extraction module is used to extract keywords from the preprocessed dataset using the TextRank algorithm. The hot word and new word processing module is used to perform statistical analysis on all topics in each field according to publication time and research topic to obtain a list of hot words and a list of new words in each field. The input module is used to obtain the user's input field of interest and send the field of interest to the visualization module. The visualization module is used to receive the field of interest sent by the input module and display a list of recent popular topics in the corresponding field to the user using visualization technology.

[0040] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the intelligent assistance method for selecting scientific research and innovation topics as described above.

[0041] Compared with the prior art, the present invention has the following advantages:

[0042] This invention extracts keywords from scientific literature using the TextRank algorithm and generates a list of recent popular topics by statistically ranking the keywords. This method assists researchers in selecting research and innovation topics, leveraging the role of artificial intelligence and big data technologies in assisting decision-making in research topic selection. It provides an effective tool for scientific and technological innovation, quickly filtering out the most promising topics from massive amounts of scientific literature, data, and trends. This enables researchers and innovation teams to select research directions more effectively, avoid blind exploration, and improve the efficiency and quality of scientific and technological innovation results. Attached Figure Description

[0043] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0044] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0045] Example 1:

[0046] This embodiment provides an intelligent auxiliary method for selecting research and innovation topics, such as... Figure 1 As shown, it includes the following steps:

[0047] The original dataset consists of scientific and technological literature and its publication date. The scientific and technological literature includes: academic journals, academic papers, conference papers, scientific and technological achievements, patent documents and technical standard documents.

[0048] The original dataset is classified into three categories based on literature: paper dataset, patent dataset, and technical standard dataset.

[0049] The paper datasets were classified by domain according to the Chinese Library Classification, resulting in paper datasets for each domain.

[0050] The patent datasets are classified by domain according to patent type to obtain patent datasets for each domain.

[0051] The technical standard datasets are classified by domain according to standard type to obtain technical standard datasets for each domain.

[0052] The paper dataset includes academic journals, academic papers, conference papers, and scientific and technological achievements; the patent dataset includes patent documents; and the technical standard dataset includes technical standard documents.

[0053] The paper datasets, patent datasets, and technical standard datasets from various fields are merged by field category to obtain the datasets for each field.

[0054] Extract titles and abstracts from paper and patent datasets in various fields, and extract titles and scopes from technical standard datasets. Use the extracted data as the dataset to be processed in that field.

[0055] The dataset to be processed is preprocessed, including unstructured data such as Word and PDF files;

[0056] The preprocessing of the dataset to be processed includes the following steps:

[0057] Write relevant Python preprocessing scripts;

[0058] Read the data to be processed;

[0059] Convert unstructured data into structured text using a Python preprocessing script;

[0060] Remove blank text fields and individual abnormal text fields generated during the extraction process from the structured data text. Use a basic stop word library to supplement some stop words.

[0061] Keyword extraction from the preprocessed dataset using the TextRank algorithm includes the following steps:

[0062] Text preprocessing includes word segmentation, stop word removal, and part-of-speech tagging, converting the text into a format suitable for algorithm processing;

[0063] Building a graph model: Using words or sentences in the processed text as nodes, a graph model is built based on the relationships between words or sentences;

[0064] Calculate node weights: Using the iterative approach of the PageRank algorithm, calculate the weight of each node in the graph. The relationships between the nodes include co-occurrence and semantic similarity. The weight of a node is determined by the weights of other nodes and the strength of the relationships between them.

[0065] Sorting and Extraction: Sort nodes according to their weights, and select the second-to-last node after sorting as the keyword of the text.

[0066] The extracted keywords were used as the research topics of their corresponding scientific and technological literature.

[0067] Statistical analysis was conducted on all topics in each field based on publication time and research topic to obtain a list of hot words and a list of new words for each field. The list of hot words includes the topics of literature published within a first preset time period and their publication frequency, while the list of new words includes the topics of literature published only within a second preset time period and their publication frequency. The first preset time period is longer than the second preset time period, which is the past five years and the second preset time period is the past year.

[0068] The lists of trending words and new words are sorted according to their publication frequency;

[0069] Obtain the user's area of ​​interest (as input);

[0070] Extract the first preset value position of the literature topics and their publication frequency in the list of new words in the field of interest, find the publication frequency of their literature topics in the list of hot words in the field of interest, add the frequency to the extracted list, and obtain a preliminary list of recent hot topics.

[0071] The initial list of recent popular topics is sorted by the frequency of publication within a first preset time period to obtain the current list of popular topics.

[0072] Visualization technology is used to display a list of recent popular topics to users, intelligently assisting them in selecting research and innovation topics.

[0073] This embodiment mainly describes the specific implementation steps of a research innovation topic selection intelligent assistance method based on text mining and intelligent analysis. Step 1: Data Collection and Classification. First, paper datasets, patent datasets, and technical standard datasets for each field are obtained through domain classification. These datasets can be obtained from publicly available literature databases, patent databases, and technical standard databases. Then, these datasets are further classified according to patent type and standard type to obtain more refined datasets. Finally, these datasets are merged by domain category to obtain datasets for each field. Step 2: Data Preprocessing. The datasets to be processed are preprocessed, including converting them into structured text data, removing blank and abnormal text fields, and adding stop words. This step mainly uses the Python programming language and some natural language processing (NLP) libraries, such as NLTK and TextBlob. Specifically, the steps to convert unstructured data into structured text data include word segmentation, stop word removal, and part-of-speech tagging. The steps to remove blank and abnormal text fields are mainly implemented through regular expressions and string manipulation. The step of adding stop words mainly involves adding specific stop words from the existing stop word library according to the specific field and needs. Step 3: Keyword Extraction. The TextRank algorithm is used to extract keywords from the preprocessed dataset. TextRank is a graph-based algorithm that effectively captures keywords in text. This step primarily uses the Python programming language and graph algorithm libraries such as NetworkX. Specifically, words or sentences in the processed text are first treated as nodes, and a graph model is constructed based on the relationships between words or sentences. Then, the iterative approach of the PageRank algorithm is used to calculate the weight of each node in the graph. Finally, the nodes are sorted according to their weights, and the nodes with the second-to-last preset weight are selected as the keywords of the text. Step 4: Topic Statistics and Analysis. Statistical analysis is performed on all topics in each field based on publication time and research topic to obtain lists of hot words and new words for each field. This step primarily uses the Python programming language and data analysis libraries such as Pandas and Matplotlib. Specifically, the datasets for each field are first sorted by publication time, and then the topics of each article are categorized and statistically analyzed according to research topic. Finally, lists of hot words and new words for each field are obtained. Step 5: Visualization. The list of recently popular topics is displayed to users using visualization technology. This step primarily utilizes the Python programming language and visualization libraries such as Matplotlib and Seaborn. Specifically, it first sorts the list of recent trending topics by publication frequency, and then presents it to users using visualization techniques such as pie charts and word clouds. Step Six: User Interaction.The process involves obtaining the user's input of their area of ​​interest, extracting the first preset value of literature topics and their publication frequencies from the list of new words in that area, and then finding their publication frequencies in the hot word list to obtain a preliminary list of recent popular topics. This preliminary list is then sorted to obtain a final list of recent popular topics. This step is primarily implemented through a user interface (UI), allowing users to easily input their areas of interest and view the corresponding results. Step Seven: Intelligent Assistance. Intelligent analysis technology is used to assist in the selection of research and innovation topics. This step mainly uses the Python programming language and some machine learning libraries, such as Scikit-learn and TensorFlow. Specifically, the user's input of their area of ​​interest and the list of new words are taken as input, and then a machine learning model is used for prediction and analysis to obtain recommended research directions and topics. These are the specific implementation steps of this embodiment. It should be noted that the specific parameters, tools, and libraries mentioned in this embodiment are exemplary and can be adjusted and replaced according to specific needs.

[0074] Example 2:

[0075] This embodiment also provides an intelligent auxiliary system for selecting research and innovation topics, including: a data acquisition module, a data classification module, a data extraction module, a data preprocessing module, a keyword extraction module, a hot word and new word processing module, an input module, and a visualization module. The keyword extraction module is used to extract keywords from the preprocessed dataset using the TextRank algorithm. The hot word and new word processing module is used to perform statistical analysis on all topics in each field according to publication time and research topic to obtain a list of hot words and new words in each field. The input module is used to obtain the user's input field of interest and send the field of interest to the visualization module. The visualization module is used to receive the field of interest sent by the input module and display a list of recent popular topics in the corresponding field to the user using visualization technology.

[0076] This embodiment also provides an electronic device, including a processor and a memory, wherein the processor is used to execute a processing program for tables in a version document stored in the memory, so as to implement the text element deduplication extraction method described above.

[0077] This embodiment also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the method described above.

[0078] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0079] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. An intelligent auxiliary method for selecting research and innovation topics, characterized in that, Includes the following steps: Obtain scientific and technological literature and their publication dates as the raw dataset; The original dataset is classified into three categories based on literature: paper dataset, patent dataset, and technical standard dataset. The paper datasets were classified by domain according to the Chinese Library Classification, resulting in paper datasets for each domain. The patent datasets are classified by domain according to patent type to obtain patent datasets for each domain. The technical standard datasets are classified by domain according to standard type to obtain technical standard datasets for each domain. The paper datasets, patent datasets, and technical standard datasets from various fields are merged by field category to obtain the datasets for each field. Extract titles and abstracts from paper and patent datasets in various fields, and extract titles and scopes from technical standard datasets. Use the extracted data as the dataset to be processed in that field. Preprocess the dataset to be processed; The TextRank algorithm was used to extract keywords from the preprocessed dataset. The extracted keywords were used as the research topics of their corresponding scientific and technological literature. Statistical analysis was conducted on all topics in each field based on publication time and research topic to obtain a list of hot words and a list of new words for each field. The list of hot words includes the topics of literature published within the first preset time period and their publication frequency, while the list of new words includes the topics of literature published only within the second preset time period and their publication frequency. The lists of trending words and new words are sorted according to their publication frequency; Obtain the user's area of ​​interest (as input); Extract the first preset value position of the literature topics and their publication frequency in the list of new words in the field of interest, find the publication frequency of their literature topics in the list of hot words in the field of interest, add the frequency to the extracted list, and obtain a preliminary list of recent hot topics. The initial list of recent popular topics is sorted by the frequency of publication within a first preset time period to obtain the current list of popular topics. Visualization technology is used to display a list of recent popular topics to users, intelligently assisting them in selecting research and innovation topics. The process of extracting keywords from the preprocessed dataset using the TextRank algorithm includes the following steps: Text preprocessing includes word segmentation, stop word removal, and part-of-speech tagging, converting the text into a format suitable for algorithm processing; Building a graph model: Using words or sentences in the processed text as nodes, a graph model is built based on the relationships between words or sentences; Calculate node weights: Calculate the weight of each node in the graph using the iterative approach of the PageRank algorithm; Sorting and Extraction: Sort nodes according to their weights, and select the second-to-last node after sorting as the keyword of the text.

2. The intelligent auxiliary method for selecting scientific research and innovation topics according to claim 1, characterized in that, The scientific and technological documents include: academic journals, academic papers, conference papers, scientific and technological achievements, patent documents, and technical standard documents.

3. The intelligent auxiliary method for selecting scientific research and innovation topics according to claim 1, characterized in that, The paper dataset includes academic journals, academic papers, conference papers, and scientific and technological achievements; the patent dataset includes patent documents; and the technical standard dataset includes technical standard documents.

4. The intelligent auxiliary method for selecting scientific research and innovation topics according to claim 1, characterized in that, The data to be processed is unstructured data, including Word and PDF files.

5. The intelligent auxiliary method for selecting scientific research and innovation topics according to claim 1 or 4, characterized in that, The preprocessing of the dataset to be processed includes the following steps: Read the data to be processed; Convert unstructured data into structured text using a Python preprocessing script; Remove blank text fields and individual abnormal text fields generated during the extraction process from the structured data text. Use a basic stop word library to supplement some stop words.

6. The intelligent auxiliary method for selecting scientific research and innovation topics according to claim 1, characterized in that, The relationships between the nodes include co-occurrence relationships and semantic similarity. The weight of a node is determined by the weights of other nodes and the strength of the relationships between them.

7. The intelligent auxiliary method for selecting scientific research and innovation topics according to claim 1, characterized in that, The first preset time period is longer than the second preset time period. The first preset time period is the past five years, and the second preset time period is the past year.

8. A system for the intelligent auxiliary method for selecting scientific research and innovation topics as described in any one of claims 1 to 7, characterized in that, include: The system comprises a data acquisition module, a data classification module, a data extraction module, a data preprocessing module, a keyword extraction module, a hot word and new word processing module, an input module, and a visualization module. The keyword extraction module extracts keywords from the preprocessed dataset using the TextRank algorithm. The hot word and new word processing module performs statistical analysis on all topics in each field based on publication time and research topic to obtain a list of hot words and new words for each field. The input module obtains the user's input field of interest and sends it to the visualization module. The visualization module receives the field of interest sent by the input module and displays a list of recent popular topics in the corresponding field to the user using visualization technology.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the intelligent auxiliary method for selecting scientific research and innovation topics as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Metering acquisition research direction trend analysis method based on natural language processing

    CN110598972A

  • Scientific and technical literature review automatic generation method and device

    CN118278365A