Subject cross domain frontier identification method, system and program product
By introducing the concepts of paper families and basic papers in the interdisciplinary fields, combining text similarity detection algorithms and interdisciplinary indicators, we identify cutting-edge research in interdisciplinary fields, solving the problems of time lag, abnormal citation interference and lack of core papers in the existing technology, and achieving more efficient and accurate cutting-edge recognition.
Patent Information
- Application Number
- CN202411968392.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-23
AI Technical Summary
In the interdisciplinary fields, it is difficult for the existing technology to effectively identify cutting-edge research, and there are problems such as time lag, abnormal citation interference and lack of core papers.
By introducing the concepts of paper families and basic papers, the text similarity detection algorithm is used to extract the paper families, and the basic papers and alternative frontier topics are screened based on the paper node degree and interdisciplinary indicators Variety, and clustered again to obtain the frontier names.
This method can more accurately identify cutting-edge research in interdisciplinary fields, solve the problems of time lag, abnormal citation interference and lack of core papers, and improve the accuracy and efficiency of cutting-edge recognition.
Smart Images

Figure CN120030143A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of document analysis, and more specifically, to a method, system and program product for identifying the frontier of an interdisciplinary field. Background Art
[0002] In theory, the screening of frontiers in academic fields should rely on the judgment of experts in the field. However, since the number of disciplinary papers is usually large and experts generally have their own research fields, it is difficult to screen the frontiers in academic fields from a macro perspective.
[0003] At present, research on frontier identification methods has made some progress, mainly including research frontier identification methods based on citation analysis, research frontier identification methods based on topic analysis, and research frontier identification methods based on multiple data sources. Citation-based research frontier identification requires citations between papers before clustering can be performed, and it takes a certain amount of time for papers to be published and cited, so there is a time lag in citation clustering. In addition, there is also interference from abnormal citations. Research frontier identification based on topic analysis can solve the problem of time lag, but topic clustering is clustered by the similarity between words, and papers in different fields are often clustered together, so there is no clear evolutionary context and the topic is not focused enough. Since the 1990s, there has been no new breakthrough in identification methods. There are still few studies combining the ever-changing publication and citation characteristics of academic papers, and further research is needed on research frontier identification methods in the interdisciplinary field of basic research.
[0004] Therefore, the identification of frontiers in interdisciplinary fields of basic research still faces technical problems such as time lag, abnormal citation interference and missing core papers in the process of frontier identification, which need to be solved urgently. Summary of the invention
[0005] In view of the above-mentioned defects of the prior art, the present invention provides a method, system and program product for identifying the frontiers in interdisciplinary fields, and proposes the concepts of paper families and basic papers. By identifying paper families and basic papers, the theoretical basis and development context of the frontiers can be effectively analyzed, thereby efficiently and accurately identifying the frontiers in interdisciplinary fields, and solving the problems of time lag, abnormal citation interference and missing core papers in the frontier identification process.
[0006] To achieve the above objectives, in a first aspect, the present invention provides a method for identifying frontiers in interdisciplinary fields, characterized in that it comprises the following steps:
[0007] Step S1, search the subject to obtain and confirm the scope of the data source;
[0008] Step S2: Use a text similarity detection algorithm to calculate the similarity of the papers, and through continuous iteration, extract paper families; each paper family contains 3 to 100 papers.
[0009] Step S3: According to the paper node degree and the interdisciplinary index Variety, extract the basic papers and alternative frontier topics in the paper families.
[0010] Step S4: Use topic clustering or co-citation clustering algorithm or expert judgment to cluster the alternative frontier topics again to obtain the frontier names.
[0011] A paper family refers to a cluster of papers that have the same research topic or use the same research method to conduct in-depth exploration or extended research on research questions. Basic papers are the most influential papers in a paper family and are of great value for in-depth analysis and revelation of the mechanisms and characteristics of specific fields. Basic papers can represent the core research content of a paper family. Compared with previous clustering methods, by identifying paper families and basic papers, clustering themes can be better condensed, so that each clustering can reveal a certain scientific problem. Basic papers can be selected based on different indicators according to disciplinary characteristics. The present invention adopts the identification of frontiers in interdisciplinary fields based on paper families and basic papers, which can better link frontier identification with the disciplinary development context, and strengthen the interdisciplinary attributes through clustering screening. By combining the paper family size limit and re-clustering of frontier topics, the purpose of balancing the recognition accuracy and algorithm overhead is achieved, thus solving the problems existing in the prior art.
[0012] Further, in the step S2, the steps for extracting paper families are as follows:
[0013] Step S21: Extract the core words of each paper to construct a data set S i ;
[0014] Step S22: Use the Simhash algorithm to calculate the similarity N between every two papers.
[0015] Step S23: Select the papers with similarity N < n (n is the similarity threshold) to construct a paper similarity network clustering; each connected subgraph in the paper similarity network is a paper family; for different fields, different n values can be set.
[0016] Further, the Simhash algorithm is a Simhash algorithm based on information entropy weighting.
[0017] Further, the core words of the papers come from the titles, abstracts and keywords of the papers.
[0018] Furthermore, if the size of a paper similarity network is greater than 100, it is pruned starting from the edge with the lowest similarity N until the size of each cluster is between 3 and 100.
[0019] Furthermore, in step S3, the papers whose node degree and variety rank in the top 10% are selected to obtain the intersection, so as to obtain basic papers and alternative frontier topics.
[0020] In a second aspect, the present invention provides a frontier identification system for interdisciplinary fields, characterized by comprising:
[0021] The data collection module searches for subject topics and obtains the scope of data sources;
[0022] The frontier clustering module uses a text similarity detection algorithm to calculate the similarity of papers and extract paper families through continuous iteration; each paper family contains 3 to 100 papers;
[0023] The clustering screening module extracts the basic papers and candidate frontier topics in the paper family based on the paper node degree and interdisciplinary indicator Variety;
[0024] The frontier interpretation module uses thematic clustering or co-citation clustering algorithms or expert judgment to cluster the candidate frontier topics again and extract the frontier names.
[0025] Furthermore, in the cluster screening module, the papers whose node degree and variety rank in the top 10% are selected to take the intersection, so as to obtain basic papers and alternative frontier topics.
[0026] In a third aspect, the present invention provides a computer program product, characterized in that when the computer program product is run on a computer, the computer is enabled to execute the method for identifying frontiers in interdisciplinary fields as described above.
[0027] Compared with the prior art, the present invention has the following technical effects:
[0028] (1) The present invention adopts interdisciplinary frontier identification based on paper families and basic papers, which can better link frontier identification with the development context of disciplines and improve the accuracy of frontier identification.
[0029] (2) This invention strengthens the interdisciplinary attributes of identifying frontiers through cluster screening standards (paper node degree and interdisciplinary indicator variety).
[0030] (3) The present invention combines paper family size restrictions with frontier topic clustering to ensure recognition accuracy while minimizing algorithm overhead. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 Flowchart of the method for identifying the frontiers in the interdisciplinary field according to an embodiment of the present invention;
[0032] Figure 2 Schematic diagram of the paper family according to an embodiment of the present invention;
[0033] Figure 3 Schematic diagram of the evolution of the basic papers according to an embodiment of the present invention.
[0034] Figure 4 Schematic diagram of the clustering of alternative frontier themes according to an embodiment of the present invention. Detailed implementation manners
[0035] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, but it is not intended to limit the present invention.
[0036] In the following detailed description, many specific details are set forth to provide a more thorough understanding of the present invention. However, it is obvious to those skilled in the art that well-known algorithms and models (such as the Simhash algorithm) are not shown in detail to avoid obscuring the gist of the present invention.
[0037] In addition, for the execution order of the actions and steps in the devices and methods shown in the claims, the description and the drawings, as long as there is no specific limitation on the order, and as long as the output of the previous process is not used in the subsequent process, they can be implemented in any order.
[0038] Embodiment 1
[0039] Refer to Figure 1 , this embodiment provides a method for identifying the frontiers in the interdisciplinary field, including the following steps:
[0040] Step S1: Retrieve the disciplinary themes and obtain and confirm the data source scope.
[0041] Step S2: Use the text similarity detection algorithm to calculate the similarity of the papers, and extract the paper families through continuous iteration; each paper family contains 3 to 100 papers. More specifically, the example of the paper family extraction step is as follows:
[0042] Step S21: Extract the core words of each paper and construct the data set S i ;
[0043] Step S22: Use the Simhash algorithm to calculate the similarity N between two papers;
[0044] Step S23: Select the papers with the similarity N < n (n is the similarity threshold) to construct the paper similarity network clustering; each connected subgraph in the paper similarity network is a paper family (refer to Figure 2, one color in the figure represents one paper family); different n values can be set for different fields.
[0045] The algorithm of the extraction process of the paper family is implemented by Python. Taking the topic of "Isogeometric Analysis" as an example, the search formula is Isogeometric*, and the papers in the core collection of Web of Science are searched. There are 4953 papers in total, and the core words of the papers come from the title, abstract and keywords.
[0046] There are many ways to calculate text similarity. Cosine similarity is a more accurate calculation method, but the efficiency is low. Simhash is faster when processing large-scale data. Simhash is the effect of feature vector hash superposition, which is locally sensitive. Ordinary hash algorithms will produce completely different hash values for strings that are not much different, but the Hamming distance of hash values calculated by Simhash will be very similar. The Simhash algorithm is used to calculate the similarity of two papers. The lower the weight (or similarity), the more similar the two papers are. Furthermore, the Simhash algorithm based on information entropy weighting introduces TF-IDF and information entropy. By increasing text distribution information, the weight and threshold calculation in the Simhash algorithm are optimized, and the correlation between fingerprint information and weight is analyzed. The deduplication rate, recall rate, F value, etc. of the Simhash algorithm used for text deduplication are better than those of the traditional Simhash algorithm.
[0047] As shown in Table 1, Paper 1 and Paper 21 are both introductions to data analysis using the equigeometry method, and Paper 2 and Paper 22 are both introductions to the dynamics of fluid-driven tubing in a pump and the calculation method of residence time. Because they can be considered to belong to the same paper family, Papers 3 and 23, as well as Papers 6 and 26 are two parts of the same topic, so they can be considered to belong to the same paper family. By analyzing the papers, it can be found that the Simhash algorithm can better identify similar papers, and a cluster of similar papers forms a paper family.
[0048] surface Similarity of papers in the field of geometric analysis
[0049]
[0050]
[0051]
[0052]
[0053] If the size of a paper similarity network is greater than 100, it will be pruned from the edge with the lowest similarity N until each cluster size is between 3 and 100. The limitation of network size makes the clustering theme more focused. If the size is too large, each cluster may involve multiple topics. If the size is too small, with only 1 or 2 papers, the representativeness and coverage of the research topics interpreted will be insufficient. The average number of references for global core papers from 2016 to 2022 is 43.8. If each cluster consists of at least 3 papers, the upper limit of the cluster is selected as 100 papers.
[0054] Step S3, extract the basic papers and alternative frontier topics in the paper family according to the paper node degree and the interdisciplinary indicator Variety. Node degree refers to the number of edges associated with the node; if a paper has a relatively high node degree, it can be considered that it is closely connected to multiple other papers or has a relatively high connection strength with certain papers, then its position in the knowledge flow network is relatively important. Variety is an indicator that describes interdisciplinarity and can more intuitively reflect the intersection of disciplines. More specifically, select the papers that rank in the top 10% in terms of node degree and variety and take the intersection to obtain the basic papers and alternative frontier topics in the paper family.
[0055] Basic papers are papers that play an important role in the development of a field. Basic papers will promote the further development of a field. The number of basic papers in each paper family varies, but there will be at least one. Figure 3 As shown in the figure, the red, yellow and blue circles are the basic papers of a paper family. Basic papers will change dynamically. In a paper family, some papers are defined as basic papers quickly, and some papers are proven to be basic papers after a period of silence. Papers that have been identified as basic papers may be proven to be unable to solve scientific problems well in subsequent developments, and thus withdraw from the ranks of basic papers. Therefore, both paper families and basic papers are dynamic and time-sensitive.
[0056] Step S4, cluster the candidate frontier topics again using thematic clustering or co-citation clustering algorithm or expert judgment to obtain the frontier name. Due to the limitation of the size of the paper family, the papers will be divided, so the frontiers of the same topic will be divided into several clusters. Therefore, when interpreting the frontier, clusters of the same topic need to be merged and irrelevant topics need to be deleted. When confirming the name of the frontier, it needs to be named in combination with the basic paper content. The frontier name needs to reflect the core content of the basic paper. When the number of papers is small, the expert judgment method can also be used to refine and interpret the frontier.
[0057] The following uses a paper in the field of civil engineering as an example to demonstrate the process of identifying frontiers in this interdisciplinary field.
[0058] (1) Data collection
[0059] Data collection is performed according to the fields required for subsequent analysis. Since references are used for cluster analysis, the downloaded papers contain at least title, abstract, keywords, author, institution, citation frequency, year, CNCI, and reference list information. The data comes from Web of Science. Data from 1975 to 2022 are downloaded for analysis, and a total of 431,010 papers are obtained. Among them, the frontiers of 65,918 papers from 2021 to 2022 are extracted, and data from other years are used for verification, such as the frontiers extracted from data from 2011 to 2015 for verification.
[0060] (2) Frontier Clustering
[0061] By using the Simhash algorithm to calculate the similarity of two papers, and clustering the papers through continuous iteration, new papers can be clustered into a category more quickly. Each cluster contains 3-100 papers. A total of 2,858 candidate hotspot clusters were obtained for the papers in 2021-2022.
[0062] (3) Cluster screening
[0063] Filter the basic papers in each cluster. This example uses node degree and interdisciplinary index Variety as the indicators for filtering basic papers. Filter the top 10% cluster set A based on the interdisciplinary index Variety; filter the basic papers based on node degree and select the top 10% cluster set B. Take the intersection of A and B, see Figure 4 , 39 candidate frontier clusters were obtained.
[0064] (4) Frontier interpretation
[0065] By interpreting the 39 frontiers, the names of the 39 frontiers are obtained, namely (1) Application of artificial intelligence and data mining (AI&DM) technology in runoff simulation, (2) Application of deep learning hybrid model in runoff prediction, (3) Water system evaluation and optimization based on wavelet analysis and neural network, (4) Application of machine learning in drought prediction, (5) Large-scale structural health monitoring based on artificial intelligence, (6) River runoff modeling based on machine learning, (7) Impact of global warming on urban heating and cooling degree days, (8) Estimation of reservoir evaporation loss, and (9) Building energy consumption simulation, (10) urban energy performance and indoor and outdoor thermal comfort, (11) thermal mitigation of urban green space infrastructure, (12) alkali activation and geopolymer materials, (13) performance of fiber-reinforced geopolymer composites, (14) automatic posture ergonomic risk assessment based on machine learning, (15) the role of artificial intelligence in construction engineering and management, (16) durability of alkali-activated concrete materials, (17) structural damage detection based on artificial intelligence and computer vision, (18) building structure design and performance evaluation based on machine learning, (19) water quality impact of extreme events, (2 0) Performance evaluation of fiber concrete, (21) Water resistance and luminescent thermal stability of self-luminous cement-based materials, (22) Performance evaluation of concrete containing calcined clay, (23) Sensitivity analysis of runoff simulation, (24) Concrete damage assessment based on machine learning, (25) Nano-modified geopolymer grouting method, (26) Runoff model based on machine learning, (27) Outdoor thermal comfort of spray cooling system, (28) Turbulence model, (29) Automatic detection of fastener defects in any direction of high-speed railway, (30) Reliability method for structural seismic elastic evaluation, (31) Ice plasticity of building materials Model numerical simulation, (32) Flood risk assessment based on artificial intelligence model, (33) Dynamic monitoring of super-high-rise structures and multi-span bridges based on GNSS-RTK technology, (34) Spatial elastic modulus of coarse-grained soil in the pavement subbase and base, (35) Cement paste backfill using flue gas desulfurization gypsum and phosphogypsum and its acid resistance, (36) Key factors for predicting lake surface water temperature using machine learning, (37) Damage mechanics of reinforced concrete components under post-explosion fire scenarios, (38) Prediction of compressive strength of geopolymer concrete based on machine learning, (39) Water vapor permeability of shale.
[0066] As can be seen from the above names, several AI-based runoff simulations have been split and need to be merged when clustering again, and some irrelevant topics need to be deleted.
[0067] Example 2
[0068] This embodiment provides a frontier identification system for interdisciplinary fields, which is used to implement the identification method described in Embodiment 1, including:
[0069] The data collection module searches for subject topics and obtains the scope of data sources;
[0070] The frontier clustering module uses a text similarity detection algorithm to calculate the similarity of papers and extract paper families through continuous iteration; each paper family contains 3 to 100 papers;
[0071] The clustering screening module extracts the basic papers and candidate frontier topics in the paper family based on the paper node degree and interdisciplinary indicator Variety;
[0072] The frontier interpretation module uses thematic clustering or co-citation clustering algorithms or expert judgment to cluster the candidate frontier topics again and extract the frontier names.
[0073] If the above-mentioned frontier identification method of the interdisciplinary field is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Therefore, the essence of this technical solution or the part that contributes to the prior art or the part of this technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions to enable an electronic device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0074] In summary, the present invention provides a method, system and program product for identifying frontiers in interdisciplinary fields, which solve the problems of time lag, abnormal citation interference and missing core papers in the frontier identification process. The identification method includes data collection, extracting paper families using a text similarity detection algorithm, extracting basic papers and alternative frontier topics in the paper family based on paper node degree and interdisciplinary indicator Variety, and clustering the alternative frontier topics again using topic clustering or co-citation clustering algorithm or expert judgment to obtain frontier names. The present invention uses interdisciplinary frontier identification based on paper families and basic papers to better link frontier identification with the development context of disciplines, and strengthens interdisciplinary attributes through clustering screening. It balances the purpose of identification accuracy and algorithm overhead by combining paper family size restrictions and re-clustering of frontier topics.
[0075] Those skilled in the art should understand that those skilled in the art can implement variations by combining the prior art and the above embodiments, which will not be described in detail here. Such variations do not affect the essential content of the present invention, and will not be described in detail here.
[0076] The above describes the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the above-mentioned specific embodiments, and the devices and structures that are not described in detail should be understood to be implemented in a common manner in the art; any technician familiar with the art can use the above-disclosed methods and technical contents to make many possible changes and modifications to the technical solutions of the present invention without departing from the scope of the technical solutions of the present invention, or modify them into equivalent embodiments of equivalent changes, which does not affect the essential content of the present invention. Therefore, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solutions of the present invention are still within the scope of protection of the technical solutions of the present invention.
[0077] Those skilled in the art will appreciate that the units, i.e., algorithm steps, of the various examples described in conjunction with this embodiment can be implemented in electronic hardware or in a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
Claims
1. A method for identifying frontiers in interdisciplinary fields, characterized in that: It includes the following steps: Step S1: Retrieve the subject theme, and obtain and confirm the data source scope; Step S2: Use the text similarity detection algorithm to calculate the similarity of papers, and through continuous iteration, extract paper families; each paper family contains 3 to 100 papers; Step S3: According to the paper node degree and the interdisciplinary index Variety, extract the basic papers and alternative frontier themes in the paper families; Step S4: Use the theme clustering or co-citation clustering algorithm or expert judgment to cluster the alternative frontier themes again to obtain the frontier names.
2. The method for identifying frontiers in interdisciplinary fields according to claim 1, characterized in that: In the said Step S2, the paper family extraction steps are as follows: Step S21: Extract the core words of each paper and construct the data set S i ; Step S22: Use the Simhash algorithm to calculate the similarity N between every two papers; Step S23: Select the papers with similarity N < n (n is the similarity threshold) to construct a paper similarity network clustering; each connected subgraph in the paper similarity network is a paper family; For different fields, different n values can be set.
3. The method for identifying frontiers in interdisciplinary fields according to claim 2, characterized in that: The said Simhash algorithm is the Simhash algorithm based on information entropy weighting.
4. The method for identifying frontiers in interdisciplinary fields according to claim 2, characterized in that: The core words of the said papers come from the titles, abstracts and keywords of the papers.
5. The method for identifying frontiers in interdisciplinary fields according to claim 2, characterized in that: If the scale of a paper similarity network is greater than 100, start pruning from the edge with the lowest similarity N until the size of each cluster is between 3 and 100.
6. A method for identifying frontiers in interdisciplinary fields according to claim 1 or 2, characterized in that: In the said Step S3, select the intersection of the papers ranked in the top 10% in terms of node degree and Variety to obtain the basic papers and alternative frontier themes.
7. A frontier identification system for interdisciplinary fields, characterized by: It includes: A data collection module that retrieves the subject theme and obtains the data source scope; A frontier clustering module that uses the text similarity detection algorithm to calculate the similarity of papers, and through continuous iteration, extracts paper families; each paper family contains 3 to 100 papers; A clustering screening module that extracts the basic papers and alternative frontier themes in the paper families according to the paper node degree and the interdisciplinary index Variety; A frontier interpretation module that uses the theme clustering or co-citation clustering algorithm or expert judgment to cluster the alternative frontier themes again to obtain the frontier names.
8. The system for identifying frontiers in interdisciplinary fields according to claim 7, characterized in that: In the said clustering screening module, select the intersection of the papers ranked in the top 10% in terms of node degree and Variety to obtain the basic papers and alternative frontier themes.
9. A computer program product, characterized in that When the said computer program product runs on a computer, it enables the computer to execute the identification method for the frontier in the interdisciplinary field as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Document citation network visualization and document recommendation method and system
CN105589948A
Literature clustering method based on reference network and text similarity network
CN110083703A
Aviation scientific research paper classification method based on deep learning
CN110516064A
Article clustering method and apparatus, electronic device and storage medium
CN110888978A
Clustering method of science and technology documents
CN111460154A