Innovative theme identification method and device, equipment, storage medium and program product
By screening interdisciplinary feature indicators and calculating multi-dimensional indicators, the problem of incomplete feature measurement and insufficient semantic mining in existing technologies for identifying groundbreaking scientific innovation topics has been solved, achieving efficient and accurate identification and visualization analysis of innovation topics.
Patent Information
- Application Number
- CN202511572675.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-30
- Publication Date
- 2026-03-03
AI Technical Summary
Existing technologies, when identifying groundbreaking scientific innovations, suffer from insufficient feature measurement and semantic mining, resulting in inaccurate and poorly interpretable identification results.
Potential paper collections are screened using interdisciplinary characteristic indicators, optimal topic groups are determined by combining the silhouette coefficient, topic clustering and dimensionality reduction visualization are used to calculate novelty, content transformativeness, mutation and impact indicators, and the target innovative topics are determined by allocating weights using the entropy weight method.
It accurately identifies groundbreaking innovation topics, improves data processing efficiency and topic interpretability, and provides an objective and systematic basis for scientific research resource allocation and industrial development planning.
Smart Images

Figure CN121597830A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of innovative subject recognition technology, and in particular to an innovative subject recognition method, apparatus, device, storage medium, and program product. Background Technology
[0002] In the context of science and technology innovation management, scientific research resource allocation, and industrial development planning, accurately identifying groundbreaking scientific innovation themes is key to opening up new tracks in science and technology and achieving leapfrog development. For example, in cutting-edge fields such as agricultural robots, it is necessary to accurately identify groundbreaking innovation directions to provide a basis for scientific research decisions in order to break through core technology bottlenecks.
[0003] In existing technologies, the identification of groundbreaking scientific innovation topics mainly includes two types: qualitative analysis and quantitative analysis. Qualitative analysis relies primarily on the wisdom and experience of experts. Quantitative methods, typically categorized as bibliometrics, citation analysis, co-occurrence network analysis, and text mining, analyze text content or network structures formed based on relationships to quantify the characteristics of groundbreaking innovations, thereby identifying the research topic.
[0004] However, the above scheme has obvious shortcomings: on the one hand, the measurement features and dimensions of groundbreaking scientific innovation are relatively limited and not comprehensive enough, making it difficult to fully reflect the innovation potential; on the other hand, the identification methods lack depth in semantic mining of scientific texts, traditional models are unable to accurately capture professional terms and contextual semantic relationships, and lack the ability to mine and visualize thematic structures through multi-method collaboration, resulting in poor interpretability and insufficient accuracy of the identification results. Summary of the Invention
[0005] This invention provides an innovative topic identification method, apparatus, device, storage medium, and program product to address the problems of incomplete feature measurement and insufficient semantic mining in existing methods. It can objectively screen potential papers, accurately cluster topics, and efficiently identify groundbreaking innovative topics in specific fields, providing a basis for scientific research decision-making.
[0006] This invention provides an innovative topic identification method, comprising: acquiring paper data in a target field; screening a set of potential papers whose scores meet a preset threshold based on interdisciplinary feature indicators; extracting text information from the potential paper set to generate text vectors; determining the optimal number of topic groups by combining silhouette coefficients; and obtaining multiple research topics through topic clustering and dimensionality reduction visualization processing; calculating the novelty index, content transformation index, mutation index, and influence index of each research topic; assigning weights to all indicators using the entropy weight method; calculating a comprehensive score; and determining the target innovative topic based on the comprehensive score.
[0007] This invention also provides an innovative topic identification device, comprising the following modules: an acquisition module and a processing module; the acquisition module is used to acquire paper data in a target field; the processing module is used to screen a set of potential papers whose scores meet a preset threshold based on interdisciplinary feature indicators; extract text information from the potential paper set to generate text vectors, determine the optimal number of topic groups by combining contour coefficients, and obtain multiple research topics through topic clustering and dimensionality reduction visualization processing; calculate the novelty index, content transformation index, mutation index, and influence index of each research topic, assign weights to all indicators using the entropy weight method, calculate a comprehensive score, and determine the target innovative topic based on the comprehensive score.
[0008] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the innovative topic identification method as described above.
[0009] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the innovative topic identification method as described above.
[0010] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the innovative topic identification method as described above.
[0011] The innovative topic identification method, apparatus, equipment, storage medium, and program products provided by this invention can filter high-potential research literature through interdisciplinary characteristic indicators, thus eliminating papers with single disciplines and low innovation potential at the data source, reducing the amount of data required for subsequent topic modeling and indicator calculation, and improving data processing efficiency. Because it can clearly and accurately divide multiple research topics, and each topic's connotation can be clearly defined through clustering results and visualization features, it can significantly improve topic interpretability, overcoming the shortcomings of existing topic-based feature-based methods in terms of poor topic interpretability. Because it can comprehensively characterize the innovative potential of a topic from multiple dimensions and objectively allocate weights, avoiding subjective judgment bias, it can more objectively and systematically identify groundbreaking innovative topics, providing a reliable basis for science and technology innovation management, scientific research resource allocation, and industrial development planning. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0013] Figure 1 This is one of the flowcharts illustrating the innovative topic identification method provided by this invention; Figure 2 This is the second flowchart illustrating the innovative topic identification method provided by this invention; Figure 3 This is a schematic diagram of the structure of the innovative subject recognition device provided by the present invention; Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0014] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0015] It should be noted that in the embodiments of this application, the words "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the words "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0016] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0017] To facilitate a clear description of the technical solutions of the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish the same or similar items with essentially the same function and effect. Those skilled in the art can understand that the terms "first" and "second" are not intended to limit the quantity or execution order.
[0018] This application describes some exemplary embodiments for illustrative purposes. It should be understood that this application may be implemented in other ways not specifically shown in the accompanying drawings.
[0019] like Figure 1 As shown, this application provides an innovative topic identification method, which can be applied to an innovative topic identification device. The innovative topic identification method may include steps S101-S103: S101. An innovative topic identification device acquires paper data in the target field and selects potential paper sets whose scores meet preset thresholds based on interdisciplinary feature indicators.
[0020] Specifically, such as Figure 2 As shown, the data first comes from two core data sources: the Web of Science (WOS) database and the Journal of Journals (JCR) database. Original paper data in the target field is obtained, and references and subject classification information of corresponding journals are obtained simultaneously. Then, preprocessing operations are performed on the original paper data, including data item extraction, data deduplication, and deletion of meaningless items. A domain thesaurus and stop word list are constructed in combination with professional knowledge of the target field. Finally, after word segmentation, merging, and stemming, the data is stored in the data pool.
[0021] Based on the preprocessed paper data, and according to the journal-subject mapping table provided by the JCR database, two mapping relationships are established: "target paper-subject" and "references-subject," clarifying the subject classification of each paper and its references. Then, using the Derwent data analysis tool, a subject co-occurrence matrix is generated based on the above mapping relationship. In this matrix, the elements represent the frequency of two subjects appearing together in the references of the same paper. At the same time, the Leydesdorff global subject distance matrix is introduced to obtain similarity data between different subjects. Finally, the subject co-occurrence matrix and subject similarity data are substituted into the subject cross-disciplinary feature index calculation formula to calculate the subject cross-disciplinary feature index for each target paper, thereby quantifying the comprehensive subject cross-disciplinary degree of the paper.
[0022] Optionally, the expression for the interdisciplinary characteristic index is: ; in, This indicates the interdisciplinary nature of the paper. The higher the value, the greater the degree of interdisciplinary integration in the paper. Documents i and literature j Discipline similarity, and Representing subject categories i and subject categories j The proportion of the number of references in the total number of references.
[0023] After that, for all target papers Perform statistical analysis on the values and plot them. Value distribution histograms are used to observe data distribution characteristics; preset thresholds are then determined based on these characteristics. Taking agricultural robotics as an example, statistics show that approximately 82% of papers... The values are concentrated in the range of 1-11.5, with values above 11.5 indicating highly interdisciplinary papers. Therefore, 11.5 can be set as the preset threshold; finally, the following papers were selected. Papers that meet the preset threshold are selected to form a potential collection of papers in the target field.
[0024] It should be noted that, on the one hand, by using interdisciplinary feature indicators for screening, papers with a single discipline and low innovation potential can be eliminated from the source, avoiding interference from irrelevant data in subsequent topic modeling and indicator calculation, and ensuring that subsequent analysis is based on high-value research literature; on the other hand, since subsequent operations are only performed on the selected potential paper set, the amount of data processing can be significantly reduced, improving computational efficiency in large-scale data scenarios and solving the problem of insufficient computational efficiency of existing topic-based feature-based methods; furthermore, because The indicators comprehensively consider the multi-dimensional characteristics of interdisciplinary studies, and can accurately identify papers that contain the integration of diverse knowledge. Therefore, they can provide clear analytical objects for subsequent accurate identification of innovative themes.
[0025] S102. The innovative topic recognition device extracts the text information of the potential paper collection to generate text vectors, combines the contour coefficient to determine the optimal number of topic groups, and obtains multiple research topics through topic clustering and dimensionality reduction visualization processing.
[0026] Optionally, the step of extracting text information from the potential paper collection to generate text vectors, determining the optimal number of topic groups based on silhouette coefficients, and obtaining multiple research topics through topic clustering and dimensionality reduction visualization processing includes: using the SciBERT model to vectorize the text of the potential paper collection to generate high-dimensional text vectors; calculating the silhouette coefficients corresponding to the number of topic groups based on the high-dimensional text vectors, and selecting the number of topic groups with the largest silhouette coefficients as the optimal number of topic groups; using the K-means unsupervised clustering algorithm to cluster the high-dimensional text vectors according to the optimal number of topic groups, obtaining topic clusters consistent with the optimal number of topic groups; extracting the core topic words of each topic cluster, using the UMAP dimensionality reduction algorithm to perform dimensionality reduction on the high-dimensional text vectors, drawing a topic clustering visualization based on the dimensionality reduction results, and determining multiple research topics by combining the core topic words and the topic clustering visualization.
[0027] Specifically, such as Figure 2 As shown, firstly, for the potential paper collection selected from S101, the title, abstract, and keywords of each paper are extracted as text corpus. Text cleaning is then performed again to remove stop words and merge synonyms. The cleaned text is then input into the SciBERT pre-trained model to generate a high-dimensional text vector matrix with dimensions of "number of papers × 768". Next, the silhouette coefficients corresponding to different topic grouping numbers are calculated using Python's scikit-learn library. Then, the high-dimensional text vector matrix is input into the K-means unsupervised clustering algorithm. For example, if K=5 is used for clustering, five topic clusters and the "paper-topic cluster" correspondence can be obtained. Next, the UMAP algorithm is used to reduce the 768-dimensional vector to a 2-dimensional space. Based on the dimensionality reduction results, a topic clustering visualization is drawn, with different colors used to label the five topic clusters and the proportion of papers in each cluster. Finally, for the text within each topic cluster, the TF-IDF algorithm is used to calculate the lexical importance, and the top 15-20 highly important words are extracted as core topic words. Finally, the research topic name is determined by combining the core topic words with the visualization.
[0028] Optionally, the expression for the contour coefficient is: ; in, This represents the intra-topic dissimilarity, which is the average dissimilarity between the i-th vector and all other vectors within the same topic. The smaller the value, the more semantically consistent the papers within the topic. This represents the dissimilarity between topics, which is the minimum average dissimilarity between the i-th vector and all other topics. The larger the value, the more significant the differences between topics. Indicates the first Contour coefficients of individual objects, all objects The mean is the overall profile coefficient corresponding to the K value. The larger the mean, the better the theme grouping effect.
[0029] It should be noted that text vectorization using the SciBERT model avoids the problem of ordinary text models' insufficient understanding of scientific terminology, ensuring that high-dimensional vectors accurately reflect the semantics of the paper and laying the foundation for the accuracy of subsequent clustering. Determining the optimal number of topic groups based on the silhouette coefficient avoids the subjectivity of manually setting the K value, ensuring that the topic division is neither too broad nor too fragmented. The combination of K-means clustering and UMAP dimensionality reduction visualization not only achieves accurate grouping of potential papers but also presents the topic distribution intuitively through visualization, making it easier for researchers to quickly grasp the research landscape of the field. The extraction of core topic words solves the problem of poor interpretability of traditional topic clustering results, making the research content of each topic clear and specific. The entire process focuses on the set of potential papers, which further improves computational efficiency compared to topic modeling of all papers. At the same time, the accuracy of topic division provides high-quality analytical units for subsequent calculation of indicators such as novelty and content transformativeness, avoiding the deviation of indicator calculation due to topic mixing.
[0030] S103. The innovative theme identification device calculates the novelty index, content transformation index, mutation index, and impact index of each research theme, assigns weights to all indicators using the entropy weight method, calculates the comprehensive score, and determines the target innovative theme based on the comprehensive score.
[0031] Optionally, such as Figure 2 As shown, the novelty index is used to measure the degree of innovation of the research topic in the time dimension, including the novelty of the knowledge base and the novelty of the knowledge; the content transformation index is used to measure the ability of the research topic to change the existing knowledge structure of the field; the mutation index is used to measure the explosiveness of the research topic in the development pace; and the influence index is used to measure the importance of the research topic in the field value, including degree centrality, betweenness centrality and proximity centrality.
[0032] (1) Novelty Novelty is the foundation of groundbreaking scientific innovation. In this application, it is mainly reflected in the novelty of the time dimension, including the novelty of the knowledge base and the novelty of knowledge.
[0033] 1) Novelty of the Knowledge Base. References form the knowledge base of a scientific paper. The more recent the publication date of the references, the more the paper reflects the research progress. Therefore, the proportion of references from the past five years in the topic can be used to measure the novelty of the knowledge base.
[0034] ; in, It is the first The novelty of the knowledge base of each topic These are references in the topic.
[0035] 2) Novelty of knowledge. The novelty of knowledge can be measured by the average publication year of the thematic collection of papers.
[0036] ; in, It is the first The novelty of the topic For the first The publication year of the document The number of documents in the topic.
[0037] (2) Content Transformation Content transformativity primarily characterizes the degree of change in knowledge structure and is a metric based on citation networks. The degree of content change can be measured by the average disruptiveness index of a thematic collection of papers. The formula is as follows: ; in, It refers to the number of times that only the main paper is cited without citing its references. It is the number of times the focus paper and any of its references are cited simultaneously.
[0038] (3) Mutability In the process of scientific development, the abrupt breaks in network structures are key nodes for topic mutations and important locations for measuring mutability. The classic Kleinberg burst word detection algorithm can be used as an indicator of the mutability dimension of knowledge structures. ; ; ; in, Indicator sudden value, Indicator Frequency of occurrence within time period k Indicator The average frequency of occurrence over the entire study period. Indicator The standard deviation of the frequency of occurrence over the entire study period.
[0039] (4) Influence The impact of breakthrough scientific innovation can be divided into its own influence and its influence within the field. Its own influence is measured by the number of citations of papers in a collection under a specific theme, representing the absolute influence of the theme formed by the papers.
[0040] ; in, To demonstrate one's influence. Indicate the theme, It refers to the number of topics.
[0041] Domain influence is used to measure the status of a topic within a network. In network analysis, network centrality is often used to measure the influence of a node in the network. Related indicators include proximity centrality, betweenness centrality, and degree centrality. This application can calculate the centrality by weighting degree centrality, betweenness centrality, and proximity centrality.
[0042] 1) Degree centrality measures the number of direct connections a node has with other nodes in a social network. It reflects a node's direct influence and activity level. The formula is as follows: ; in, It is a node Degree centrality, It is a node The degree, that is, the number of nodes connected to it. It represents the total number of nodes in the network.
[0043] 2) Proximity centrality measures the reciprocal of the sum of the shortest path lengths from a node to all other nodes. It reflects the central position of a node in the network and the speed at which information propagates from that node to other nodes. The formula is as follows: ; in, It is a node Proximity to centrality It is the total number of papers on the internet. It is a node To the node The shortest path length.
[0044] 3) Betweenness centrality is the proportion of the number of shortest paths through a certain topic node in the topic network to the total number of shortest paths. The higher the betweenness centrality value of a node, the stronger the node's control over information and ability to utilize resources compared to the non-connected nodes at both ends of the node, reflecting its domain influence.
[0045] ; in, It is a thesis The centrality of the middle, and Indicates and Other papers that connect to this topic, Indicates connection to the paper and thesis And after the paper The number of shortest paths, Indicates the connection node and nodes The number of shortest paths.
[0046] Optionally, the entropy weight method assigns weights based on the amount of information contained in each indicator; the greater the information content, the higher the weight. First, the original values of all indicators are converted to dimensionless values in the [0,1] interval using the maximum-minimum standardization method to eliminate dimensional differences; then, the formula is applied... Calculate the entropy value of each indicator using the formula. Calculate the indicator weights; finally, multiply the standardized indicator values of each theme by their corresponding weights and sum them to obtain the comprehensive score. Represents entropy value. Represents weight, Indicates the first The sample at the th The proportion of each indicator It is a constant. It refers to the number of samples.
[0047] It should be noted that the multi-dimensional indicators cover four dimensions: time, structure, rhythm, and value, avoiding the one-sidedness of existing technologies that judge innovation potential from only a single dimension. The introduction of the entropy weight method ensures that the weight allocation is objective and avoids the influence of human subjective preferences on the results. By combining network analysis with indicator calculation, both the individual characteristics of the paper and the macro network position of the theme are taken into account. This solves the shortcomings of existing document-level methods that ignore the overall trend and theme-level methods that lack micro support, and achieves accurate identification of the target innovation theme.
[0048] In this embodiment, because high-potential research literature can be screened using interdisciplinary characteristic indicators, papers with a single discipline and low innovation potential can be eliminated from the data source, reducing the amount of data required for subsequent topic modeling and indicator calculation, and improving data processing efficiency. Since multiple research topics can be clearly and accurately divided, and the connotation of each topic can be clearly defined through clustering results and visualization features, topic interpretability can be significantly improved, overcoming the shortcomings of existing topic-based feature-based methods in terms of poor topic interpretability. Because the innovation potential of a topic can be comprehensively characterized from multiple dimensions and weights can be objectively allocated, avoiding subjective judgment bias, breakthrough innovation topics can be identified more objectively and systematically, providing a reliable basis for science and technology innovation management, scientific research resource allocation, and industrial development planning.
[0049] The foregoing mainly describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the above functions, it includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, in conjunction with the units and algorithm steps of the various examples described in the embodiments disclosed herein, the embodiments of this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0050] The innovation topic identification method provided in this application can be executed by an innovation topic identification device or a control module for innovation topic identification within that device. This application uses an innovation topic identification device executing the innovation topic identification method as an example to illustrate the innovation topic identification device provided in this application.
[0051] It should be noted that the embodiments of this application can divide the innovative theme identification device into functional modules according to the above method examples. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing module. The integrated modules can be implemented in hardware or as software functional modules. Optionally, the module division in the embodiments of this application is illustrative and is only a logical functional division; other division methods may be used in actual implementation.
[0052] like Figure 3 As shown in the figure, this application embodiment provides an innovative topic identification device 300. The innovative topic identification device 300 includes: an acquisition module 301 and a processing module 302. The acquisition module 301 can be used to acquire paper data in a target field; the processing module 302 can be used to screen potential paper sets whose scores meet preset thresholds based on interdisciplinary feature indicators; extract text information from the potential paper sets to generate text vectors, determine the optimal number of topic groups by combining silhouette coefficients, and obtain multiple research topics through topic clustering and dimensionality reduction visualization processing; calculate the novelty index, content transformation index, mutation index, and influence index of each research topic, assign weights to all indicators using the entropy weight method, calculate a comprehensive score, and determine the target innovative topic based on the comprehensive score.
[0053] Optionally, the expression for the interdisciplinary characteristic index is: ; in, This indicates the interdisciplinary nature of the paper. The higher the value, the greater the degree of interdisciplinary integration in the paper. Documents i and literature j Discipline similarity, and Representing subject categories i and subject categories j The proportion of the number of references in the total number of references.
[0054] Optionally, the step of extracting text information from the potential paper collection to generate text vectors, determining the optimal number of topic groups based on silhouette coefficients, and obtaining multiple research topics through topic clustering and dimensionality reduction visualization processing includes: using the SciBERT model to vectorize the text of the potential paper collection to generate high-dimensional text vectors; calculating the silhouette coefficients corresponding to the number of topic groups based on the high-dimensional text vectors, and selecting the number of topic groups with the largest silhouette coefficients as the optimal number of topic groups; using the K-means unsupervised clustering algorithm to cluster the high-dimensional text vectors according to the optimal number of topic groups, obtaining topic clusters consistent with the optimal number of topic groups; extracting the core topic words of each topic cluster, using the UMAP dimensionality reduction algorithm to perform dimensionality reduction on the high-dimensional text vectors, drawing a topic clustering visualization based on the dimensionality reduction results, and determining multiple research topics by combining the core topic words and the topic clustering visualization.
[0055] Optionally, the expression for the contour coefficient is: ; in, Indicates the degree of dissimilarity within the topic. Indicates the dissimilarity between topics. Indicates the first Contour coefficients of each object.
[0056] Optionally, the novelty index is used to measure the degree of innovation of the research topic in the time dimension; the content transformation index is used to measure the ability of the research topic to change the existing knowledge structure of the field; the mutation index is used to measure the explosiveness of the research topic in the development pace; and the influence index is used to measure the importance of the research topic in the field value.
[0057] In this embodiment, because high-potential research literature can be screened using interdisciplinary characteristic indicators, papers with a single discipline and low innovation potential can be eliminated from the data source, reducing the amount of data required for subsequent topic modeling and indicator calculation, and improving data processing efficiency. Since multiple research topics can be clearly and accurately divided, and the connotation of each topic can be clearly defined through clustering results and visualization features, topic interpretability can be significantly improved, overcoming the shortcomings of existing topic-based feature-based methods in terms of poor topic interpretability. Because the innovation potential of a topic can be comprehensively characterized from multiple dimensions and weights can be objectively allocated, avoiding subjective judgment bias, breakthrough innovation topics can be identified more objectively and systematically, providing a reliable basis for science and technology innovation management, scientific research resource allocation, and industrial development planning.
[0058] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4 As shown, the electronic device may include a processor 410, a communications interface 420, a memory 430, and a communication bus 440, wherein the processor 410, communications interface 420, and memory 430 communicate with each other via the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute an innovative topic identification method. This method includes: acquiring paper data in the target field; screening a set of potential papers whose scores meet a preset threshold based on interdisciplinary feature indicators; extracting text information from the potential paper set to generate text vectors; determining the optimal number of topic groups by combining silhouette coefficients; obtaining multiple research topics through topic clustering and dimensionality reduction visualization processing; calculating the novelty index, content transformative index, mutation index, and impact index of each research topic; assigning weights to all indicators using the entropy weight method; calculating a comprehensive score; and determining the target innovative topic based on the comprehensive score.
[0059] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0060] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the innovative topic identification method provided by the above methods. The method includes: acquiring paper data in the target field, screening a set of potential papers whose scores meet a preset threshold based on interdisciplinary feature indicators; extracting text information from the potential paper set to generate text vectors, determining the optimal number of topic groups by combining silhouette coefficients, and obtaining multiple research topics through topic clustering and dimensionality reduction visualization processing; calculating the novelty index, content transformation index, mutation index, and influence index of each research topic, assigning weights to all indicators using the entropy weight method, calculating a comprehensive score, and determining the target innovative topic based on the comprehensive score.
[0061] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the innovative topic identification method provided by the above methods. The method includes: acquiring paper data in a target field; screening a set of potential papers whose scores meet a preset threshold based on interdisciplinary feature indicators; extracting text information from the potential paper set to generate text vectors; determining the optimal number of topic groups by combining silhouette coefficients; and obtaining multiple research topics through topic clustering and dimensionality reduction visualization processing; calculating the novelty index, content transformation index, mutation index, and influence index of each research topic; assigning weights to all indicators using the entropy weight method; calculating a comprehensive score; and determining the target innovative topic based on the comprehensive score.
[0062] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0063] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0064] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An innovative topic identification method, characterized in that, include: Acquire paper data in the target field and filter potential paper sets whose scores meet preset thresholds based on interdisciplinary characteristic indicators; Text information of the potential paper collection is extracted to generate text vectors. The optimal number of topic groups is determined by combining the silhouette coefficient. Multiple research topics are obtained through topic clustering and dimensionality reduction visualization. For each research topic, calculate the novelty index, content transformation index, mutation index, and impact index. Assign weights to all indicators using the entropy weight method, calculate the comprehensive score, and determine the target innovation topic based on the comprehensive score.
2. The innovative topic identification method according to claim 1, characterized in that, The expression for the interdisciplinary characteristic index is as follows: ; in, This indicates the interdisciplinary nature of the paper. The higher the value, the greater the degree of interdisciplinary integration in the paper. Documents i and literature j Discipline similarity, and Representing subject categories i and subject categories j The proportion of the number of references in the total number of references.
3. The innovative topic identification method according to claim 1, characterized in that, The process involves extracting textual information from the potential paper collection to generate text vectors, combining this with silhouette coefficients to determine the optimal number of topic groups, and then using topic clustering and dimensionality reduction visualization to obtain multiple research topics, including: The SciBERT model was used to vectorize the text of the potential paper collection, generating high-dimensional text vectors. The contour coefficients corresponding to the number of topic groups are calculated based on the high-dimensional text vectors, and the number of topic groups with the largest contour coefficients is selected as the optimal number of topic groups. Based on the optimal number of topic groups, the high-dimensional text vectors are clustered using the K-means unsupervised clustering algorithm to obtain topic clusters that are consistent with the optimal number of topic groups. The core keywords of each topic cluster are extracted, and the high-dimensional text vector is reduced using the UMAP dimensionality reduction algorithm. A topic clustering visualization is drawn based on the dimensionality reduction results. Multiple research topics are determined by combining the core keywords and the topic clustering visualization.
4. The innovative topic identification method according to claim 3, characterized in that, The expression for the contour coefficient is: ; in, Indicates the degree of dissimilarity within the topic. Indicates the dissimilarity between topics. Indicates the first Contour coefficients of each object.
5. The innovative topic identification method according to claim 1, characterized in that, The novelty index is used to measure the degree of innovation of the research topic over time. The content transformation index is used to measure the ability of a research topic to change the existing knowledge structure of the domain. The mutation rate is used to measure the explosiveness of the research topic in terms of its development pace; The influence metric is used to measure the importance of a research topic in terms of its value within the field.
6. An innovative subject recognition device, characterized in that, include: Acquisition module and processing module; The acquisition module is used to acquire paper data in the target field; The processing module is used to screen potential paper collections whose scores meet preset thresholds based on interdisciplinary feature indicators; extract text information from the potential paper collections to generate text vectors, determine the optimal number of topic groups by combining the silhouette coefficient, and obtain multiple research topics through topic clustering and dimensionality reduction visualization; calculate the novelty index, content transformation index, mutation index, and impact index of each research topic, assign weights to all indicators using the entropy weight method, calculate the comprehensive score, and determine the target innovative topic based on the comprehensive score.
7. The innovative theme recognition device according to claim 6, characterized in that, The expression for the interdisciplinary characteristic index is as follows: ; in, This indicates the interdisciplinary nature of the paper. The higher the value, the greater the degree of interdisciplinary integration in the paper. Documents i and literature j Discipline similarity, and Representing subject categories i and subject categories j The proportion of the number of references in the total number of references.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the innovative topic recognition method as described in any one of claims 1 to 5.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the innovative topic identification method as described in any one of claims 1 to 5.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the innovative topic identification method as described in any one of claims 1 to 5.