Method and system for co-citation of CNKI documents by filling reference data

By converting format and completing references to CNKI databases, the problem that CNKI database cannot perform co-citation analysis is solved, and the co-citation analysis of CNKI database documents is realized, research efficiency is improved, scholars can research channels, and new analytical ideas are provided.

CN116467289BActive Publication Date: 2025-08-15WUHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310127590.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-15
Publication Date
2025-08-15
Estimated Expiration
2043-02-15

AI Technical Summary

Technical Problem

In the prior art, the literature data of the CNKI database cannot be co-cited and analyzed, resulting in the inability of relevant scholars to find highly cited papers, highly cited authors and highly cited journals in the massive literature resources of China National Knowledge Infrastructure, hindering scholars' identification of field discipline communities and exploration of hot topics and cutting-edge research.

Method used

By filling reference data, the CNKI tool is used to convert the reference format, convert the reference format of the CNKI database into the text format exported by the WOS database, and perform field completion, rewrite the reference document, match and write the reference fields in the Refworks format document, and realize the co-citation analysis of the CNKI database document.

Benefits of technology

It has realized the co-citation analysis of CNKI database literature, expanded the research channels of scholars, improved research efficiency, and provided new co-citation analysis ideas, helping scholars to more widely analyze the development of discipline knowledge fields and their research hotspots and frontiers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116467289B_ABST
    Figure CN116467289B_ABST
Patent Text Reader

Abstract

In order to solve the limitation that when using CiteSpace to analyze CNKI database documents, only very basic analysis can be performed, and co-citation analysis cannot be performed like the WOS database and CSSCI database, the present invention proposes a CNKI document co-citation method and system by filling reference data, crawling the references of the documents in the CNKI database through a self-written crawler program, and writing the crawled references into the CNKI documents converted from CiteSpace through a self-written Python program, thereby realizing co-citation analysis, author co-citation analysis and journal co-citation analysis of CNKI database documents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text processing, and in particular to a method and system for co-citing documents in China National Knowledge Infrastructure (CNKI) by filling in reference data. Background Art

[0002] In 1973, American information scientist Small first proposed the concept of co-citation as a research method for measuring the degree of relationships between documents. As research progressed, White and Griffith expanded co-citation to the author and journal levels in 1981, developing the methods of author co-citation analysis (ACA) and journal co-citation analysis. With the rise of scientific knowledge graphs, they have become a key research method and tool in scientometrics and knowledge measurement. Through continuous research, co-citation analysis has been combined with scientific knowledge graphs, and the results have gradually been visualized. A scientific knowledge graph is a diagram that uses a knowledge domain as its object, displaying the development process and structural relationships of scientific knowledge. The concept of the scientific knowledge graph originated from a 2003 workshop held by the U.S. National Academy of Sciences. With the development of science and technology, scholars have integrated the concept of the scientific knowledge graph with technology, resulting in the development of various knowledge graph creation tools. Among the many visualization software programs, CiteSpace, developed by Chaomei Chen of Drexel University in the United States, has become popular due to its rich and beautiful graphs, which offer scholars research perspectives from multiple perspectives. With the widespread adoption of CiteSpace, numerous academic papers have been published both domestically and internationally on its application of CiteSpace and its knowledge graph. For example, Jayantha et al. used the Scopus database to search for relevant literature on hedonic price models from 1970 to 2019 and then used CiteSpace to analyze and visualize the data. Rawat et al. used CiteSpace to analyze the scientometric characteristics of literature published on the use of ICT in education from 2011 to 2020. Widziewicz-Rzonca et al. used the WOS database to search for relevant literature on PM-bound water from 1996 to 2018 and used CiteSpace to identify past trends and potential future directions in measuring aerosol-bound water.Domestic scholars such as Chen Xiaoling and others used the three core databases of SCI-E, SSCI, and CPSI in the WOS database as data sources, and used CiteSpace to analyze the research hotspots and discipline trends in the three northeastern provinces from 2012 to 2016; Hua Longxue and others used the relevant literature in the field of "process mining" included in China National Knowledge Infrastructure as samples, and used VOSviewer and CiteSpace software to analyze the literature characteristics, hot topics and cutting-edge trends; Li Lingzhi and others used CiteSpace to conduct quantitative analysis such as literature distribution, co-occurrence analysis, and co-citation analysis on the core literature on infrastructure resilience assessment in the WOS database and drew conclusions.

[0003] The inventors of this application have found through analysis of existing research that most scholars at home and abroad use CiteSpace to conduct co-citation analysis on papers in databases such as the WOS and CSSCI databases, while only keyword analysis is performed on the CNKI database. CiteSpace is rarely used to conduct document co-citation analysis, author co-citation analysis, and journal co-citation analysis on CNKI database documents. Even if a very small number of scholars conduct relevant research, they do so by manually downloading references and importing them into the corresponding articles. By downloading the latest version of CiteSpace 6.2.3 and conducting an analysis, it was found that the current version can perform co-citation analysis, keyword co-occurrence analysis, author coupling analysis, and institution coupling analysis on document data downloaded from databases such as the WOS and CSSCI databases, helping relevant researchers explore the research hotspots, research frontiers, knowledge base, major authors and institutions, etc. in a certain research field, and predict the future development direction of a certain research field. However, when the literature data exported from the CNKI database is imported into CiteSpace for analysis, only keyword co-occurrence analysis, author co-occurrence analysis, and institution co-occurrence analysis can be performed, and literature co-citation analysis cannot be performed. As a result, relevant scholars are unable to find highly cited papers, highly cited journals, and highly cited authors in the massive literature resources of CNKI. To a certain extent, this hinders relevant scholars from identifying the disciplinary community in the field and is not conducive to scholars summarizing the disciplinary paradigms in related fields. Summary of the Invention

[0004] The present invention proposes a method and system for co-citation of CNKI documents by filling reference data, so as to solve the technical problem that the prior art can only perform basic analysis on CNKI documents but cannot perform co-citation analysis.

[0005] In order to achieve the above-mentioned object, the first aspect of the present invention provides a method for co-citing documents in CNKI by filling reference data, comprising:

[0006] S1: Search CNKI documents according to preset search conditions, and obtain corresponding references based on the retrieved CNKI documents;

[0007] S2: Use CNKI tools to convert the format of the references obtained, and complete the fields in the converted text format according to the text format exported from the WOS database;

[0008] S3: rewriting the reference according to the format after field completion to obtain a rewritten reference document, wherein the file name of the rewritten reference document is the reference title;

[0009] S4: Match the file name of the rewritten reference document with the reference title in the RefWorks format document exported by the CNKI tool. If the match is successful, write the reference of the rewritten reference document into the reference field CR of the corresponding document in the RefWorks format document to obtain the reference text filled with reference data;

[0010] S5: Co-citation analysis of CNKI documents based on the reference texts filled with reference data.

[0011] In one embodiment, step S1 includes:

[0012] Perform CNKI document search based on preset search conditions to obtain a CNKI document list;

[0013] Click on a document in the CNKI document list and determine whether there are any references based on the CNKI document's citation network;

[0014] If so, the reference information will be crawled, and further judgment will be made as to whether there are multiple pages of references. If so, page-turning crawling will be performed. If not, the next document in the CNKI document list will be clicked until all the references of the documents are crawled.

[0015] In one embodiment, determining whether there is a reference based on the citation network of the CNKI document includes:

[0016] Determine whether the corresponding Chinese HowNet document has references based on the "Journal" field in the citation network.

[0017] In one embodiment, further determining whether there are multiple pages of references includes:

[0018] Use the Len() function to determine whether there are multiple pages of references.

[0019] In one embodiment, in step S2, the converted text format is completed according to the text format exported from the WOS database, including:

[0020] The converted text format is converted into the text format exported from the WOS database, and the reference field CR is completed.

[0021] In one embodiment, the reference field CR includes the author, publication year, journal information, and digital object unique identifier.

[0022] In one embodiment, step S4 includes:

[0023] Based on the reference text filled with reference data, the co-citation of documents, co-citation of authors, and co-citation of journals were analyzed for CNKI documents.

[0024] Based on the same inventive concept, the second aspect of the present invention provides a CNKI document co-citation system by populating reference data, comprising:

[0025] Reference acquisition module, used to search CNKI documents according to preset search conditions and obtain corresponding references based on the retrieved CNKI documents;

[0026] The text construction module is used to convert the format of the obtained references using CNKI tools, and complete the fields of the converted text format according to the text format exported from the WOS database;

[0027] A reference rewriting module is used to rewrite the reference according to the format after field completion to obtain a rewritten reference document, wherein the file name of the rewritten reference document is the reference title;

[0028] The reference writing module is used to match the file name of the rewritten reference document with the reference title of the RefWorks format document exported by the CNKI tool. If the match is successful, the reference of the rewritten reference document is written into the reference field CR of the corresponding document in the RefWorks format document to obtain the reference text filled with reference data;

[0029] The co-citation analysis module is used to perform co-citation analysis on CNKI documents based on the reference texts filled with reference data.

[0030] Compared with the prior art, the advantages and beneficial technical effects of the present invention are as follows:

[0031] The present invention proposes a method for co-citing CNKI documents by filling reference data, converting the format of the obtained references using CNKI tools, and completing the fields of the text format after the format conversion according to the text format exported from the WOS database; rewriting the references according to the format after the field completion, and then matching the file name of the rewritten reference document with the reference title of the Refworks format document exported by the CNKI tool. If the match is successful, the reference of the rewritten reference document is written into the reference field CR of the corresponding document in the Refworks format document to obtain the reference text filled with reference data; since the reference data is filled, co-citation analysis of CNKI documents can be realized, solving the technical problem that the existing technology can only perform basic analysis on CNKI documents but cannot perform co-citation analysis. In the actual application process, a new co-citation analysis idea is provided for relevant researchers, thereby broadening the research channels of scholars and improving the research efficiency of relevant scholars, which has certain practical significance. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0033] Figure 1 A flowchart of rewriting and filling references provided by an embodiment of the present invention;

[0034] Figure 2 This is a flow chart obtained from references in the embodiments of the present invention;

[0035] Figure 3 A flowchart constructed from the reference text in the embodiments of the present invention;

[0036] Figure 4 A flowchart rewritten from references in the embodiments of the present invention;

[0037] Figure 5 This is a schematic diagram of a plain text format derived from a WOS database according to an embodiment of the present invention;

[0038] Figure 6 This is a schematic diagram of the data format of the CNKI database after conversion by CiteSpace in an embodiment of the present invention;

[0039] Figure 7 This is a text style diagram of reference documents crawled by a crawler in an embodiment of the present invention;

[0040] Figure 8 Graph showing the text style (rewritten reference document) after rewriting using document data in an embodiment of the present invention;

[0041] Figure 9 A reference document text style diagram after being filled with reference document data in an embodiment of the present invention;

[0042] Figure 10 Schematic diagram of the co-citation analysis results of the literature in the embodiment of the present invention;

[0043] Figure 11 Schematic diagram of the co-citation analysis results of authors in the embodiment of the present invention;

[0044] Figure 12 Schematic diagram of the journal co-citation analysis results in an embodiment of the present invention. DETAILED DESCRIPTION

[0045] In response to the technical problems existing in the prior art, the inventors found through analysis that the bibliographic data exported from the CNKI database does not contain references, and thus cannot be co-citation analyzed. This results in researchers being able to only conduct co-citation analysis of Chinese literature based on the CSSCI database. However, for some natural science disciplines, the amount of data contained in the CSSCI database is limited, the retrieval logic is relatively simple, and the process is cumbersome when exporting data, which not only reduces research efficiency but also fails to produce more accurate research results. Based on this, in order to improve the research efficiency of relevant researchers, explore new channels for co-citation analysis, and more broadly analyze the development of disciplinary knowledge fields and their research hotspots, frontiers, and trends, the present invention proposes a method for implementing co-citation analysis of CNKI documents by filling in reference data, which realizes co-citation analysis of documents, author co-citation analysis, and journal co-citation analysis of the CNKI database, providing relevant researchers with new co-citation analysis ideas, thereby broadening the research channels of scholars, improving the research efficiency of relevant scholars, and having certain practical significance.

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0047] Example 1

[0048] The embodiment of the present invention provides a method for co-citing documents in CNKI by filling reference data, including:

[0049] S1: Search CNKI documents according to preset search conditions, and obtain corresponding references based on the retrieved CNKI documents;

[0050] S2: Use CNKI tools to convert the format of the references obtained, and complete the fields in the converted text format according to the text format exported from the WOS database;

[0051] S3: rewriting the reference according to the format after field completion to obtain a rewritten reference document, wherein the file name of the rewritten reference document is the reference title;

[0052] S4: Match the file name of the rewritten reference document with the reference title of the RefWorks format document exported by the CNKI tool. If the match is successful, write the reference of the rewritten reference document into the reference field CR of the corresponding document in the RefWorks format document to obtain the reference text filled with reference data;

[0053] S5: Co-citation analysis of CNKI documents based on the reference texts filled with reference data.

[0054] In order to improve the research efficiency of relevant researchers, explore new channels for co-citation analysis, and more broadly analyze the development of disciplinary knowledge fields and their research hotspots, frontiers and trends, the present invention proposes a method for co-citation analysis of CNKI documents by filling in reference data. This method realizes co-citation analysis of documents, co-citation analysis of authors and co-citation analysis of journals in the CNKI database, providing relevant researchers with new ideas for co-citation analysis, thereby broadening the research channels of scholars and improving their research efficiency, which has certain practical significance.

[0055] Combining the previous experience of using CiteSpace and reading relevant literature, it can be known that when using CiteSpace to perform keyword analysis on CNKI, the general steps are to search for documents in CNKI, export the Refworks format, and then convert it in CiteSpace, converting it into the plain text form of "Full Record and References" downloaded from the WOS document database. It can be seen that CiteSpace does not restrict the database when performing analysis, but restricts the fields of the data text. As long as the fields of the data text meet the software requirements, the CNKI database text can be co-cited. In order to realize this idea, the present invention needs to solve the following problems: CNKI reference acquisition, reference data text format construction and reference writing method construction. For the specific process, please refer to Figure 1 .

[0056] In one embodiment, step S1 includes:

[0057] Perform CNKI document search based on preset search conditions to obtain a CNKI document list;

[0058] Click on a document in the CNKI document list and determine whether there are any references based on the CNKI document's citation network;

[0059] If so, the reference information will be crawled, and further judgment will be made as to whether there are multiple pages of references. If so, page-turning crawling will be performed. If not, the next document in the CNKI document list will be clicked until all the references of the documents are crawled.

[0060] See Figure 2 , which is a flow chart obtained from references in the embodiments of the present invention.

[0061] In practice, when searching for documents using CNKI, you can begin searching after determining the search criteria. This will then display all documents that meet the search criteria, creating a list of CNKI documents. Clicking on a document reveals its title, abstract, and other information, while the citation network provides detailed information about the reference. Therefore, you can determine whether a reference exists based on the CNKI document's citation network. Specifically, you can click to obtain the URL of the article to which the reference belongs, and then parse the URL to obtain basic information about the reference.

[0062] When a document has references, since CNKI only displays 10 references, the present invention further determines whether there are multiple pages of references. If not, the operation is performed on the next document. Finally, by combining the reference title, document type identifier, author, journal, and journal year, volume, issue, and page information parsed from the URL, the references of papers on a specific topic are crawled and stored in a text document.

[0063] In one embodiment, determining whether there is a reference based on the citation network of the CNKI document includes:

[0064] Determine whether the corresponding Chinese HowNet document has references based on the "Journal" field in the citation network.

[0065] Specifically, when the return result of the journal field judgment is True, the reference information is crawled; if the return result is False, the reference information of the next document is identified and crawled.

[0066] In one embodiment, further determining whether there are multiple pages of references includes:

[0067] Use the Len() function to determine whether there are multiple pages of references.

[0068] In one embodiment, in step S2, the converted text format is completed according to the text format exported from the WOS database, including:

[0069] The converted text format is converted into the text format exported from the WOS database, and the reference field CR is completed.

[0070] Specifically, through reading the literature, it can be seen that scholars in related fields at home and abroad mostly use the WOS database to conduct co-citation analysis of documents. An important factor in this is that the WOS database can export the plain text format of "Full Record and References". Through analysis, it can be seen that the plain text format of "Full Record and References" in the WOS database mainly contains key information such as PT (publication type), AU (document author), AF (author full name), TI (document title), SO (publication name), DT (document type), AB (abstract), C1 (author address) and CR (references). It is precisely because of this information that CiteSpace can perform keyword analysis, co-citation analysis and other operations on relevant document data. When the inventor analyzed the documents in the CNKI database, the most basic operation was to export the document data in the Refworks format from the CNKI database and name the file download_XX so that it can be recognized by CiteSpace. Then, the CNKI data was converted through operations such as Data>Import / Export→CNKI→FormatConversion. Analysis revealed that the format of data downloaded from the CNKI database, after conversion via CiteSpace, is essentially the same as the plain text data format of the "Full Record & References" section of the World Wide Web (WOS). The only difference is that the CR value in the converted CNKI data is empty. This phenomenon occurs because the CNKI database cannot export references for a document, which explains why CiteSpace cannot perform co-citation analysis on data exported from the CNKI database. Therefore, the present invention completes the CR in the converted CNKI data text according to the text format of the CR in WOS, enabling CiteSpace to recognize the CR in the CNKI data text and thus complete the co-citation analysis of the CNKI data text.

[0071] By observing the CR field in the plain text format of WOS "Full Record and References", this article can know that the basic format of its references is "Author, Publication Year, Journal, v, p, DOI (Digital Object Unique Identifier)" and other fields, and each field is followed by a space and a comma with a half-width symbol. At the same time, there is one space in the reference author after the CR field, and three spaces in the other reference authors. At the same time, during the use of CiteSpace, this article found that the display of its visualization results mainly reads the author, publication year and journal information, and the following v, p, DOI can be ignored. However, in order to ensure that CiteSpace can smoothly read the data text format added in this article, the present invention sets the three data of v, p, and DOI to a fixed content. When writing references subsequently, you can write according to the above format.

[0072] See Figure 3 and Figure 4 ,in Figure 3 A flowchart constructed from the reference text in the embodiments of the present invention; Figure 4 This is a flowchart rewritten from the references in the embodiments of the present invention.

[0073] Next is the rewriting of references. Before writing the references into the converted document from CNKI, the references downloaded from CNKI need to be rewritten. This implementation sets the format of the references crawled from CNKI to "[Serial Number] Document Main Responsible Person. Document Title [Document Type Flag]. Serial Publication Title (Other Title Information), Year, Volume (Issue): Page Number.", while the three fields of author, journal, and year are required when rewriting references. At the same time, the inventor observed that the above three fields are all separated by the symbol ".", so the "author" is defined as "Name", and the method of extracting Name is "Name = ref.split('.')[1].split(',')[0].strip()"; the "year" is defined as Year, and the method of extracting Year is "Year = ref.split('.')[-1].split('(')[0].strip()"; finally, the "journal" is defined as "Article", and the method of extracting Article is "Article = ref.split('.')[2].strip()". After completing the extraction of the above data, the "Name", "Year" and "Article" field data of the document are obtained respectively. Finally, combined with the CR format "author, publication year, journal, V6, DOI Figure. 1186 / s40168-018-0470-z" determined in the previous article, the method of extracting Article is "ef1 = name+','+year+','+Article+',V6,DOI 10.1186 / s40168-018-0470-z” to complete string concatenation, thereby achieving automatic rewriting of Python references.

[0074] Then comes the writing of references. Through the foregoing preparatory work, the present invention crawls the references of relevant papers under a specific theme by self-compiled python, determines the reference rewriting format and implements it through self-compiled Python code. After the reference is rewritten, the present invention uses the document title as the file name, and the file content is the rewritten content of all references of the reference, and then reads the document information after all references are rewritten by self-compiled code and returns the dictionary result. After extracting the rewritten reference information, it is necessary to write the extracted reference information into the Refworks document downloaded by CNKI that can be read by CiteSpace. Therefore, the present invention reads this document and extracts all fields in the document, thereby generating a new dictionary. The document titles in the dictionary are then traversed and compared with the titles of the aforementioned document dictionaries. If the title meets the requirements, the CR file under the field information is replaced with a field, thereby completing the writing of the reference.

[0075] It should be noted that the aforementioned "document dictionary" refers to "the present invention uses the document title as the file name, the file content is the content after all references of the reference are rewritten, and the document information of all references after rewriting is read through self-compiled code and the dictionary result returned", that is to say: the references of multiple documents were crawled by crawlers before, and multiple text texts were crawled, and the text title is the document title. When the reference is rewritten, it is necessary to extract the crawled reference information, modify its format, and modify it into a format that can be recognized by CiteSpace software. After the modification, a new file is generated. The file name is the document name of the modified reference itself. When using CiteSpace for document visualization, it is necessary to export a Refworks format document in CNKI. The document also contains a document title. Therefore, the modified reference text title is matched with the document title in the Refworks format document. If the name matches, the modified reference in the text document is written into the CR of the corresponding document in the Refworks format document, thereby realizing the filling of reference data.

[0076] In one embodiment, the reference field CR includes the author, publication year, journal information, and digital object unique identifier.

[0077] In one embodiment, step S4 includes:

[0078] Based on the reference text filled with reference data, the co-citation of documents, co-citation of authors, and co-citation of journals were analyzed for CNKI documents.

[0079] Overall, this paper explores co-citation analysis of CNKI documents from the perspective of document co-citation analysis, innovatively implementing co-citation analysis of CNKI documents by modifying the data text format. Addressing the complexities of data acquisition and document cleaning, this paper employs self-written Python code to rapidly rewrite and re-format reference data, significantly improving research efficiency and providing new insights for scholars to analyze the current state of research in other academic fields based on co-citation analysis.

[0080] The method for filling reference data proposed by the present invention is described in detail below by using specific examples, and mainly includes the following steps:

[0081] Step 1: Reference Acquisition. By combining the reference title, document type identifier, author, journal, and journal year, volume, issue, and page information obtained from the URL, we crawl the references of papers on a specific topic and store them in a text document.

[0082] Step 2: Reference data text construction. Through reading the literature, we know that scholars in related fields at home and abroad often use the WOS database to conduct co-citation analysis of literature. The important factor is that the WOS database can export the plain text format of "full record and references". The main factors of this text format include: Figure 5 shown.

[0083] Depend on Figure 5 It can be seen that the plain text format of "Full Record and References" in the WOS database mainly includes key information such as PT (publication type), AU (author), AF (author's full name), TI (document title), SO (publication name), DT (document type), AB (abstract), C1 (author address) and CR (references). It is precisely because of this information that CiteSpace can perform keyword analysis, co-citation analysis and other operations on relevant document data. When the inventor analyzes the documents in the CNKI database, the most basic operation is to export the document data in the Refworks format in the CNKI database and name the file download_XX so that it can be recognized by CiteSpace. Then, the CNKI data is converted through operations such as Data>Import / Export→CNKI→Format Conversion. The converted text format is as follows: Figure 6 shown.

[0084] Depend on Figure 6 As can be seen, the format of data downloaded from the CNKI database after CiteSpace conversion is essentially the same as the plain text data format of the "Full Record & References" section of the WOS. The only difference is that the CR value of the converted CNKI data is empty. This phenomenon occurs because the CNKI database cannot export references of documents, which also explains why CiteSpace cannot perform co-citation analysis on data exported from the CNKI database. Therefore, the present invention completes the CR in the converted CNKI data text according to the text format of the CR in WOS, allowing CiteSpace to recognize the CR of the CNKI data text, thereby completing the co-citation analysis of the CNKI data text.

[0085] By observing the CR field in the plain text format of WOS "Full Record and References", this article can know that the basic format of its references is "author, publication year, journal, v, p, DOI" and other fields, and each field is followed by a space and a comma with a half-width symbol. At the same time, there is one space in the reference author after the CR field, and three spaces in the other reference authors. At the same time, in the process of using CiteSpace, this article found that the display of its visualization results mainly reads the author, publication year and journal information, and the following v, p, DOI can be ignored. However, in order to ensure that CiteSpace can smoothly read the data text format added in this article, this article sets a fixed content for the three data v, p, DOI. The content set in this experiment is "V6, DOI 10.1186 / s40168-018-0470-z", so the final format of CR is determined to be: "Author, publication year, journal, V6, DOI 10.1186 / s40168-018-0470-z", so when writing references, just follow the above format.

[0086] Step 3: Rewriting references. Before writing references into the converted document on CNKI, it is necessary to rewrite the references downloaded from CNKI. The present invention sets the format of the references crawled from CNKI to "[Serial Number] Document Main Responsible Person. Document Title [Document Type Flag]. Serial Publication Title (Other Title Information), Year, Volume (Issue): Page Number.", while the three fields of author, journal, and year are required when rewriting references. At the same time, the inventor observed that the above three fields are all separated by the symbol ".", so the "author" is defined as "Name", and the method of extracting Name is "Name = ref.split('.')[1].split(',')[0].strip()"; the "year" is defined as Year, and the method of extracting Year is "Year = ref.split('.')[-1].split('(')[0].strip()"; finally, the "journal" is defined as "Article", and the method of extracting Article is "Article = ref.split('.')[2].strip()". After completing the extraction of the above data, the "Name", "Year" and "Article" field data of the document are obtained respectively. Finally, combined with the CR format "author, publication year, journal, V6, DOI 10.1186 / s40168-018-0470-z" determined in the previous article, "ef1 = name+','+year+','+Article+',V6,DOI 10.1186 / s40168-018-0470-z” to complete string concatenation, thereby achieving automatic rewriting of Python references.

[0087] Step 4, reference writing. Through the previous preparatory work, the present invention crawls the references of relevant papers under a specific topic through self-compiled python, determines the reference rewriting format and implements it through self-compiled Python code. After the reference is rewritten, the present invention uses the document title as the file name, and the file content is the rewritten content of all references of the reference, and then reads the document information of all references after rewriting through self-compiled code and returns the dictionary result. After extracting the rewritten reference information, it is necessary to write the above fields into the Refworks document downloaded by CNKI that can be read by CiteSpace. Therefore, the present invention reads this document and extracts all fields in the document, thereby generating a new dictionary. The document titles in the dictionary are then traversed and compared with the titles of the aforementioned document dictionary. If the title meets the requirements, the CR file under the field information is replaced with a field, thereby completing the reference writing.

[0088] Some experimental results are shown as follows Figures 7 to 9 shown.

[0089] ①Crawler data style, such as Figure 7 shown.

[0090] ② Literature data rewrites text style, such as Figure 8 shown.

[0091] ③ The document is written in the book style, such as Figure 9 shown.

[0092] Application Results

[0093] Through the system's "data crawling - document rewriting - document writing - document export" functions, this implementation method obtains CNKI document data that is readable by CiteSpace and contains references, and then imports the document input into CiteSpace for operation. The import process is as follows:

[0094] ① First open the CiteSpace software and click agree;

[0095] ② Since the text written in the reference is the text converted by CiteSpace through data-import-WOS-RemoveDuplicates, only its CR field is completed, so all the written text is integrated into a text document and named "download.CNKI", and then directly placed in the NEW-DataDirectory of CiteSpace for direct document co-citation analysis.

[0096] ③ After the documents are imported, the co-citation analysis of documents, co-citation of authors and co-citation of journals are performed respectively.

[0097] Co-citation refers to two (or more) papers being cited simultaneously by one or more subsequent papers. These two papers are said to be in a co-citation relationship. For example, if document A cites both documents C and D, then documents C and D are in a co-citation relationship. The number of papers that cite these two papers is called the co-citation intensity. In this case, the co-citation intensity is 1, because only document A cites both C and D. Another example: if documents A and B cite documents C, D, and E simultaneously, then documents C, D, and E are in a co-citation relationship, with a co-citation intensity of 2. Co-citation relationships between documents can change over time, and studying co-citation networks can help explore the development and evolution of a discipline.

[0098] Author co-citation occurs when two (or more) authors are simultaneously cited by one or more subsequent papers. These two or more authors are said to be in a co-citation relationship. For example, if document A cites authors C and D, then authors C and D are in a co-citation relationship. The number of papers that cite these two authors is called the co-citation intensity. In this case, the co-citation intensity is 1, because only document A cites C and D. Another example: if documents A and B cite authors C, D, and E, then authors C, D, and E are in a co-citation relationship, with a co-citation intensity of 2. Author co-citation relationships can change over time.

[0099] Journal co-citation refers to the situation where a paper from two journals (or papers from multiple journals) is simultaneously cited by one or more papers from a later journal. These two journals are said to be in a co-citation relationship. For example, if a paper from journal A cites a paper from both journals C and D, then journals C and D are in a co-citation relationship. The number of papers that cite papers from both journals is called the co-citation intensity. In this case, the co-citation intensity is 1, because only the paper from journal A cites both journals C and D. Another example: if papers from journals A and B cite papers from journals C, D, and E, then journals C, D, and E are in a co-citation relationship, with a co-citation intensity of 2.

[0100] The final analysis results are as follows:

[0101] ① Co-citation network of literature, please see Figure 10 .

[0102] Table 1 Information on documents with citation frequency greater than 20

[0103]

[0104]

[0105] from Figure 10As shown in Table 1, within the research literature on "smart healthcare," authors tend to cite papers that review the theoretical evolution of "smart healthcare," clarify the development path of "smart healthcare," and analyze the current state of "smart healthcare" to highlight their understanding of the significance and impact of "smart healthcare" research. Furthermore, the figure and table also show that the most frequently cited papers were published between 2013 and 2016. This phenomenon stems from the fact that during this period, research on "smart healthcare" was transitioning from theoretical to practical research, leading to a rapid increase in the number of published papers. Furthermore, given the substantial theoretical research already underway, many papers with cutting-edge perspectives were published, expanding beyond theoretical research to include "smart healthcare big data" and "artificial intelligence and smart healthcare," leading to this high citation rate. This suggests that the research conducted between 2013 and 2016 provided a strong theoretical foundation for the flourishing development of the "smart healthcare" field.

[0106] ② Author co-citation network, please see Figure 11 .

[0107] Table 2 Information of authors with citation frequency greater than 20

[0108]

[0109]

[0110] Figure 11 The table shows the co-citation network of the literature on “smart healthcare” in the CNKI database; Table 2 lists the information of authors with a citation frequency greater than 20. Author co-citation means that two (or more) authors are cited by one or more subsequent papers at the same time, and these two (or more) authors are said to form a co-citation relationship. The network density is the “actual number of relationships” divided by the “maximum theoretical number of relationships”, that is, in an undirected network with N nodes, the maximum possible number of relationships is Assuming the actual relationship is m, the network density is 2m / [n(n-1)]. The network density of the co-citation network of authors in the field of "Smart Healthcare" is 0.0188, which shows that the network density is relatively dense. Figure 11 The middle nodes represent citation frequency, with node size proportional to citation frequency. The top five authors by citation frequency are Li Jiangong, Xiang Gaoyue, Gong Fangfang, Fang Yuan, and Zuo Xiuran. Among them, Li Jiangong not only has the highest co-citation frequency, but also has a relatively high centrality of 0.17, indicating his high academic standing in the field of smart healthcare. Scholars seeking to understand research related to smart healthcare should prioritize reading Li Jiangong's publications.

[0111] ③ Journal co-citation network, please see Figure 12 .

[0112] Table 3 Information of journals with citation frequency greater than 50

[0113]

[0114] Figure 12 Figure 3 shows the co-citation network of journals in the "smart healthcare" field under the CNKI database; Table 3 lists journals with citations greater than 50. Academic journals are important vehicles for scholarly dissemination, the foundation of academic research, and a symbol of publication quality. As the figure shows, research related to "smart healthcare" encompasses relatively similar areas of study, primarily focusing on medicine, public health, and information management. The size of the node represents the number of co-citations received by the journal. The largest node in the figure is the Journal of Medical Informatics, with 191 citations, demonstrating its propensity to include literature related to "smart healthcare" and its in-depth research in this field. This provides valuable insights for scholars seeking to engage with research related to "smart healthcare" or to understand cutting-edge developments in the field; they should prioritize reading these journals. Five journals, including "Chinese Digital Medicine," "Chinese Hospital Management," "Chinese Journal of Health Information Management," "Chinese General Practice," and "Chinese Hospital," all have citations exceeding 100 and relatively high betweenness centrality, making them high-quality journals for understanding the field of "smart healthcare."

[0115] It should be noted here that since CiteSpace is open source software and is open to users free of charge, users can import downloaded data into the software to visualize the data and use the displayed results for purposes and occasions such as paper publication and academic discussion. Therefore, the use of images generated by CiteSpace by this software and this manual will not give rise to any copyright issues.

[0116] It can be seen from this that the CNKI reference conversion method implemented by this invention can not only realize the co-citation analysis of CNKI document data, clarify the highly cited papers, highly cited authors and highly cited journals in a certain field in the CNKI database, but also analyze whether the connection between the documents is close, thereby providing reference and reference for relevant researchers to carry out academic research. At the same time, the design of this invention has also broadened the channels for domestic scholars to conduct co-citation analysis, innovated the research method of document co-citation analysis, and provided relevant researchers with new co-citation analysis ideas. The development of this invention has certain practical significance in improving the research efficiency of relevant researchers, exploring new channels for co-citation analysis, and analyzing the development of subject knowledge fields and their research hotspots, frontiers and trends more extensively than ever before.

[0117] Example 2

[0118] Based on the same inventive concept, this embodiment provides a CNKI document co-citation system by populating reference data, including:

[0119] Reference acquisition module, used to search CNKI documents according to preset search conditions and obtain corresponding references based on the retrieved CNKI documents;

[0120] The text construction module is used to convert the format of the obtained references using CNKI tools, and complete the fields of the converted text format according to the text format exported from the WOS database;

[0121] A reference rewriting module is used to rewrite the reference according to the format after field completion to obtain a rewritten reference document, wherein the file name of the rewritten reference document is the reference title;

[0122] The reference writing module is used to match the file name of the rewritten reference document with the reference title of the RefWorks format document exported by the CNKI tool. If the match is successful, the reference of the rewritten reference document is written into the reference field CR of the corresponding document in the RefWorks format document to obtain the reference text filled with reference data;

[0123] The co-citation analysis module is used to perform co-citation analysis on CNKI documents based on the reference texts filled with reference data.

[0124] Since the system described in Example 2 of the present invention is the system used to implement the CNKI document co-citation method by filling reference data in Example 1 of the present invention, the specific structure and variations of this system are readily apparent to those skilled in the art based on the method described in Example 1 of the present invention, and thus will not be described in detail here. All systems used in the method of Example 1 of the present invention fall within the scope of protection of the present invention.

[0125] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0126] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0127] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0128] Obviously, those skilled in the art may make various changes and modifications to the embodiments of the present invention without departing from the spirit and scope of the embodiments of the present invention. Thus, if such changes and modifications of the embodiments of the present invention fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. The method of co-citation of CNKI documents by filling reference data is characterized by: include: S1: Search CNKI documents according to preset search conditions, and obtain corresponding references based on the retrieved CNKI documents; S2: Use CNKI tools to convert the format of the references obtained, and complete the fields in the converted text format according to the text format exported from the WOS database; S3: rewriting the reference according to the format after field completion to obtain a rewritten reference document, wherein the file name of the rewritten reference document is the reference title; S4: Match the file name of the rewritten reference document with the reference title in the RefWorks format document exported by the CNKI tool. If the match is successful, write the reference of the rewritten reference document into the reference field CR of the corresponding document in the RefWorks format document to obtain the reference text filled with reference data; S5: Co-citation analysis of CNKI documents based on the reference texts filled with reference data.

2. The CNKI document co-citation method by filling reference data according to claim 1, characterized in that: Step S1 includes: Perform CNKI document search based on preset search conditions to obtain a CNKI document list; Click on a document in the CNKI document list and determine whether there are any references based on the CNKI document's citation network; If so, the reference information will be crawled, and further judgment will be made as to whether there are multiple pages of references. If so, page-turning crawling will be performed. If not, the next document in the CNKI document list will be clicked until all the references of the documents are crawled.

3. The CNKI document co-citation method by filling reference data according to claim 2, characterized in that: Determine whether there are references based on the citation network of the CNKI document, including: Determine whether the corresponding Chinese HowNet document has references based on the "Journal" field in the citation network.

4. The CNKI document co-citation method by filling reference data according to claim 1, characterized in that: Further checks are performed to determine if there are multiple pages of references, including: Use the Len() function to determine whether there are multiple pages of references.

5. The CNKI document co-citation method by filling reference data according to claim 1, characterized in that: In step S2, the converted text format is completed according to the text format exported from the WOS database, including: The converted text format is converted into the text format exported from the WOS database, and the reference field CR is completed.

6. The CNKI document co-citation method by filling reference data according to claim 5, characterized in that: The reference field CR includes the author, publication year, journal information, and digital object unique identifier.

7. The CNKI document co-citation method by filling reference data according to claim 1, characterized in that: Step S4 includes: Based on the reference text filled with reference data, the co-citation of documents, co-citation of authors, and co-citation of journals were analyzed for CNKI documents.

8. The CNKI document co-citation system is populated with reference data, which is characterized by: include: Reference acquisition module, used to search CNKI documents according to preset search conditions and obtain corresponding references based on the retrieved CNKI documents; The text construction module is used to convert the format of the obtained references using CNKI tools, and complete the fields of the converted text format according to the text format exported from the WOS database; A reference rewriting module is used to rewrite the reference according to the format after field completion to obtain a rewritten reference document, wherein the file name of the rewritten reference document is the reference title; The reference writing module is used to match the file name of the rewritten reference document with the reference title of the RefWorks format document exported by the CNKI tool. If the match is successful, the reference of the rewritten reference document is written into the reference field CR of the corresponding document in the RefWorks format document to obtain the reference text filled with reference data; The co-citation analysis module is used to perform co-citation analysis on CNKI documents based on the reference texts filled with reference data.

Citation Information

Patent Citations

  • Research front visual analysis method based on literature co-citation clustering

    CN108509481A

  • Automatic construction system of references

    KR1020160064306A