Document metering method and device based on knowledge graph

Through bibliographic methods based on knowledge graphs, statistics and clustering topic words, the problem that the current bibliographic methods ignore text problems, making it difficult to highlight core topics due to the analysis results, and the effect of effectively determining the core subjects of the text and research hotspots is achieved.

CN119988575APending Publication Date: 2025-05-13BEIJING INST OF AEROSPACE INFORMATION & INFORMATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411902529.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Current bibliographic methods tend to ignore text problems in the literature, making it difficult for the analysis results to highlight the core topics and research hotspots in related fields.

Method used

The bibliographic method based on knowledge graph is used to calculate the word frequency of subject words, calculate the inverse text frequency index, sort it in descending order, obtain the text feature value, and use the clustering algorithm and cosine distance function to cluster the subject words to determine the core theme of the text.

Benefits of technology

The core subjects of the text and related fields have been effectively identified, and the problem that the current bibliographic method is difficult to highlight the core topics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988575A_ABST
    Figure CN119988575A_ABST
Patent Text Reader

Abstract

The invention discloses a literature metering method and device based on a knowledge graph, and the method comprises the steps: carrying out the statistics of the word frequency of each subject term in a literature data set, determining an inverse text frequency index, and obtaining an inverse text frequency weight value of each subject term; carrying out descending sorting on the inverse text frequency weight value of each subject term to obtain an inverse text frequency weight value after descending sorting; obtaining a text feature value according to a preset screening condition and the inverse text frequency weight value after descending; and clustering the subject terms by adopting a clustering algorithm and a cosine distance function to obtain a text core subject. According to the literature metering method based on the knowledge graph, the word frequency matrix is constructed through co-word analysis of high-frequency words, correlation between keywords is mined, the distance between subject terms is calculated through text clustering, a research subject is clarified, and a text core subject and research hotspots in related fields are effectively determined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bibliometrics, and in particular to a knowledge graph-based bibliometrics method and device. Background Art

[0002] In recent years, bibliometric analysis methods have taken the full text of the document as the research object, using mathematical, statistical and other quantitative methods, combined with the distribution structure, quantitative relationship, change law and quantitative evaluation methods of the knowledge units within the document. At present, bibliometric research mainly focuses on the following aspects: citation content analysis, analyzing the position, intensity and context of citations and citations; knowledge unit measurement, starting from the word level, sentence level and discourse level, and deeply measuring knowledge entities through the full text; full text measurement indicators, combining full text features, semantics and traditional measurement indicators. However, the current bibliometric methods tend to ignore the text problems in the document, resulting in the analysis results not being able to highlight the core themes and research hotspots in related fields. Summary of the invention

[0003] The present invention proposes a knowledge graph-based bibliometric method and device to solve the problem that current bibliometric methods tend to ignore text issues in documents, resulting in the difficulty in highlighting core topics in analysis results.

[0004] The present invention provides the following technical solutions:

[0005] In the first aspect, this specification provides a bibliometric method based on knowledge graph, including:

[0006] Count the frequency of each subject word in the document data set, determine the inverse text frequency index, and obtain the inverse text frequency weight value of each subject word;

[0007] Sorting the inverse text frequency weight value of each subject word in descending order to obtain the inverse text frequency weight value in descending order;

[0008] Obtaining a text feature value according to a preset screening condition and the descending inverse text frequency weight value;

[0009] Clustering algorithm and cosine distance function are used to cluster the subject words and obtain the core theme of the text.

[0010] In a second aspect, the present invention provides a knowledge graph-based bibliometrics device, comprising:

[0011] The inverse text frequency determination module is used to count the frequency of each subject word in the document data set, determine the inverse text frequency index, and obtain the inverse text frequency weight value of each subject word;

[0012] An inverse text frequency descending module is used to sort the inverse text frequency weight value of each subject word in descending order to obtain the inverse text frequency weight value after descending order;

[0013] A text feature value determination module, used to obtain a text feature value according to a preset screening condition and the descending inverse text frequency weight value;

[0014] The text core theme determination module is used to cluster the subject words using a clustering algorithm and a cosine distance function to obtain the text core theme.

[0015] In a third aspect, the present invention provides an electronic device comprising a memory and a processor, wherein the memory is used to store a computer program, the computer program comprises program instructions, and the processor is configured to call the program instructions to execute a bibliometric method based on a knowledge graph.

[0016] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes a bibliometric method based on a knowledge graph.

[0017] The knowledge graph-based bibliometric method and device provided in the embodiment of the present invention construct a word frequency matrix through high-frequency word co-word analysis, mine the mutual correlation between keywords, calculate the distance between subject words by using text clustering, clarify the research topic, and thus effectively determine the core body of the text and research hotspots in related fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 Schematic diagram of a process of a knowledge graph-based bibliometric method in an embodiment of the present invention.

[0019] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solutions and advantages of this specification more clear, the technical solutions of this specification will be clearly and completely described below in combination with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this specification.

[0021] As described herein, the term “including” and various variations thereof may be understood as open-ended terms meaning “including but not limited to,” and the term “one embodiment” may be understood as “at least one embodiment.”

[0022] The inventors have found that current bibliometric methods tend to ignore textual issues in documents, resulting in the difficulty in highlighting core themes in the analysis results. In view of this, in an embodiment of the present invention, text clustering is used to calculate the distance between subject terms, thereby effectively determining the core body of the text and research hotspots in related fields.

[0023] Embodiment 1

[0024] like Figure 1 As shown, an embodiment of the present invention provides a literature measurement method based on a knowledge graph, which is executed by a processor and includes the following steps:

[0025] Step 1: Count the frequency of each subject word in the document data set, determine the inverse text frequency index, and obtain the inverse text frequency weight value of each subject word.

[0026] In specific implementation, the inverse text frequency weight value is calculated according to the following formula:

[0027]

[0028] Among them, TFIDF is the inverse text frequency weight value, IDF is the inverse text frequency index, TF is the frequency of the subject word, X is the number of times the subject word appears, N is the total number of words in the text, and D A is the total number of document samples, and D is the number of texts in which the keyword appears.

[0029] Step 2: sort the inverse text frequency weight value of each subject word in descending order to obtain the inverse text frequency weight value in descending order.

[0030] Step 3: Obtain text feature values ​​according to the preset screening conditions and the descending inverse text frequency weight values.

[0031] Step 4: Use clustering algorithm and cosine distance function to cluster the keywords and obtain the core theme of the text.

[0032] When implementing it, follow the steps below:

[0033] For the literature dataset, set the random seed number and target cluster K value;

[0034] Select K texts from the document dataset as the center points of the text cluster;

[0035] Calculate the distance between each object in the document data set and each initial cluster center point of the text set, and the cosine similarity between the subject words, and match each object in the document data set to the cluster center with the closest distance;

[0036] The text set cluster center point is recalculated, and each object in the document data set is matched to the nearest cluster center again until the cluster center is stable, and the core theme of the text is obtained.

[0037] Optionally, the cosine similarity between topic words is calculated according to the following formula:

[0038]

[0039] Among them, cos(s i ,c j ) is the cosine similarity between the subject words, s i is the first keyword data point, c j is the second keyword data point, ||...|| represents a vector.

[0040] The above embodiment constructs a word frequency matrix through high-frequency word co-word analysis, mines the mutual correlation between keywords, uses text clustering to calculate the distance between subject words, clarifies the research topic, and effectively determines the core body of the text and research hotspots in related fields, thereby solving the problem that the current bibliometric method easily ignores text problems in the literature, resulting in the difficulty of highlighting the core theme in the analysis results.

[0041] Embodiment 2

[0042] Based on the same technical concept, an embodiment of the present invention also provides a literature measurement device based on knowledge graph. Since the principle of solving the problem by the above-mentioned device is similar to the literature measurement method based on knowledge graph, the implementation of the above-mentioned device can refer to the implementation of the method, and the repeated parts will not be repeated.

[0043] The embodiment of the present invention provides a knowledge graph-based bibliometrics device, including:

[0044] The inverse text frequency determination module is used to count the frequency of each subject word in the document data set, determine the inverse text frequency index, and obtain the inverse text frequency weight value of each subject word;

[0045] An inverse text frequency descending module is used to sort the inverse text frequency weight value of each subject word in descending order to obtain the inverse text frequency weight value after descending order;

[0046] A text feature value determination module, used to obtain a text feature value according to a preset screening condition and the descending inverse text frequency weight value;

[0047] The text core theme determination module is used to cluster the subject words using a clustering algorithm and a cosine distance function to obtain the text core theme.

[0048] Embodiment 3

[0049] Based on the same technical concept, an embodiment of the present invention also provides an electronic device, including a memory and a processor, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute a bibliometric method based on a knowledge graph.

[0050] Embodiment 4

[0051] Based on the same technical concept, an embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program. The computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes a bibliometric method based on a knowledge graph.

[0052] In summary, the embodiment of the present invention constructs a word frequency matrix through high-frequency word co-word analysis, mines the mutual correlation between keywords, uses text clustering to calculate the distance between subject words, clarifies the research topic, and effectively determines the core body of the text and research hotspots in related fields, thereby solving the problem that the current bibliometric method easily ignores text problems in the literature, resulting in the difficulty of highlighting the core theme in the analysis results.

[0053] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0054] Those skilled in the art know that, in addition to implementing the system and its various subsystems, modules, and units provided by the present invention in a purely computer-readable program code, it is entirely possible to implement the same functions of the system and its various subsystems, modules, and units provided by the present invention in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, etc. by performing logic programming on the method steps. Therefore, the system and its various subsystems, modules, and units provided by the present invention can be regarded as structures within hardware components or as software modules for implementing the method.

[0055] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A bibliometric method based on knowledge graph, characterized in that: include: Count the frequency of each subject word in the document data set, determine the inverse text frequency index, and obtain the inverse text frequency weight value of each subject word; Sorting the inverse text frequency weight value of each subject word in descending order to obtain the inverse text frequency weight value in descending order; Obtaining a text feature value according to a preset screening condition and the descending inverse text frequency weight value; Clustering algorithm and cosine distance function are used to cluster the subject words and obtain the core theme of the text.

2. The method according to claim 1, characterized in that Clustering algorithm and cosine distance function are used to cluster the subject words and obtain the core topics of the text, including: For the literature dataset, set the random seed number and target cluster K value; Select K texts from the document dataset as the center points of the text cluster; Calculate the distance between each object in the document data set and each initial cluster center point of the text set, and the cosine similarity between the subject words, and match each object in the document data set to the cluster center with the closest distance; The text set cluster center point is recalculated, and each object in the document data set is matched to the nearest cluster center again until the cluster center is stable, and the core theme of the text is obtained.

3. The method according to claim 1, characterized in that The inverse text frequency weight value is calculated according to the following formula: Among them, TFIDF is the inverse text frequency weight value, IDF is the inverse text frequency index, TF is the frequency of the subject word, X is the number of times the subject word appears, N is the total number of words in the text, and D A is the total number of document samples, and D is the number of texts in which the keyword appears.

4. The method according to claim 1, characterized in that The cosine similarity between topic words is calculated according to the following formula: Among them, cos(s i ,c j ) is the cosine similarity between the subject words, s i is the first keyword data point, c j is the second keyword data point, ||...|| represents a vector.

5. A device for executing the knowledge graph-based bibliometric method according to any one of claims 1 to 4, characterized in that: include: The inverse text frequency determination module is used to count the frequency of each subject word in the document data set, determine the inverse text frequency index, and obtain the inverse text frequency weight value of each subject word; An inverse text frequency descending module is used to sort the inverse text frequency weight value of each subject word in descending order to obtain the inverse text frequency weight value after descending order; A text feature value determination module, used to obtain a text feature value according to a preset screening condition and the descending inverse text frequency weight value; The text core theme determination module is used to cluster the subject words using a clustering algorithm and a cosine distance function to obtain the text core theme.

6. The method according to claim 5, characterized in that Clustering algorithm and cosine distance function are used to cluster the subject words and obtain the core topics of the text, including: For the literature dataset, set the random seed number and target cluster K value; Select K texts from the document dataset as the center points of the text cluster; Calculate the distance between each object in the document data set and each initial cluster center point of the text set, and the cosine similarity between the subject words, and match each object in the document data set to the cluster center with the closest distance; The text set cluster center point is recalculated, and each object in the document data set is matched to the nearest cluster center again until the cluster center is stable, and the core theme of the text is obtained.

7. The method according to claim 5, characterized in that The inverse text frequency weight value is calculated according to the following formula: Among them, TFIDF is the inverse text frequency weight value, IDF is the inverse text frequency index, TF is the frequency of the subject word, X is the number of times the subject word appears, N is the total number of words in the text, and D A is the total number of document samples, and D is the number of texts in which the keyword appears.

8. The method according to claim 1, characterized in that The cosine similarity between topic words is calculated according to the following formula: Among them, cos(s i ,c j ) is the cosine similarity between the subject words, s i is the first keyword data point, c j is the second keyword data point, ||...|| represents a vector.

9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory is used to store a computer program, the computer program comprises program instructions, and the processor is configured to call the program instructions to execute the method according to any one of claims 1 to 4.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 4.