A method for analyzing the relationship between knowledge resources and author identities

By calculating the similarity of the text content and name scarcity of resource resources, the Siamese algorithm and TextRank method are used to solve the identity recognition problem of the author resources of the same name, and the accurate identification of the author identity is achieved in the absence of detailed information.

CN115905466BActive Publication Date: 2025-07-22PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211369322.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-28
Publication Date
2025-07-22
Estimated Expiration
2042-10-28

AI Technical Summary

Technical Problem

In the absence of detailed author information, it is difficult to accurately distinguish the resource content of the author of the same name, resulting in the inability to effectively identify the author's identity.

Method used

By calculating the content similarity and name scarcity of resource text, the text summary is extracted using the Siamese algorithm and TextRank method, and the name scarcity is calculated in combination with the name library to determine the identity relationship of the resource author.

Benefits of technology

In the case of insufficient information, accurately identifying the resource content of the author of the same name, realizing the effective correlation between the resource and the author's identity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115905466B_ABST
    Figure CN115905466B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for analyzing the relationship between knowledge resources and author identities, belonging to the field of information processing of knowledge resources (such as books, papers, periodicals, lecture audio and video). The present invention analyzes the relationship between knowledge resources and author identities by calculating text content similarity and name scarcity, and accurately obtains the author identities of the resource content. The present invention realizes that only the resource content and the author's name are needed to accurately obtain the relationship between the knowledge resource and the author identity in the case of less available information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of text information processing, and focuses on the information processing of knowledge resources (such as books, papers, periodicals, lecture audio and video). Background Art

[0002] There are many homonymous phenomena in Chinese, especially some common names, such as: Liu Yang, Zhang Wei, etc. In many scenarios, we need to accurately identify the author and distinguish him / her from other authors with the same name, but lack detailed information about this author (such as: work unit, email, native place, etc.), and can only obtain data of this author's works. For example: many books and audio-visual lectures only give the author's name. In this case, it is necessary to identify this author based on the author's creative content (to distinguish from other authors with the same name in other resources).

[0003] In addition, the name itself is also information that can be utilized. If the names of the homonymous authors in two resources are very rare and the probability of having the same name is small, then unless the contents of the two resources are too obviously different, their authors should be the same person. And when the names of the homonymous authors in two resources are very common, unless the contents of the two resources are highly similar, their authors should not be the same person. Therefore, a method for analyzing the dependency relationship between resource content and author identity needs to be provided. Summary of the Invention

[0004] In view of the above problems, the present invention provides a method for analyzing the relationship between knowledge resources and author identity, and accurately obtains the author identity of resource content.

[0005] The technical solution provided by the present invention is as follows:

[0006] A method for analyzing the relationship between knowledge resources and author identity, the steps of which include:

[0007] 1) For resource texts z1 and z2 in different media and fields, respectively extract resource summaries T(z1) and T(z2);

[0008] 2) If the authors of resource texts z1 and z2 have the same name, calculate the rarity of this name:

[0009]

[0010] where z1 is the aforementioned resource text; au(z1) represents the author of z1; au_x(z1) represents the surname of the author of z1; au_m(z1) represents the given name of the author of z1; au_m i (z1) represents the i-th character of the given name of the author of z1;

[0011] F(au_x(z1)) represents the probability that the surname of the author of z1 is used as the surname; F(au_m i (z1)|au_x(z1)) represents the probability that the first character of the given name is au_m1(z1) when the surname is au_x(z1); F(au_m i (z1)|au_m i-1 (z1)) represents the probability that the i-th character of the given name is au_m i-1 (z1) when the (i - 1)-th character of the given name is au_m i (z1); num(au_m(z1)) is the number of characters of au_m(z1).

[0012] The probability of the surname is calculated by the following formula:

[0013]

[0014] count(au_x(z1)) represents the number of times au_x(z1) is used as the surname; namebase represents the name database, and size(namebase) represents the size of namebase, that is, the number of names contained in namebase;

[0015]

[0016] count(au_x(z1), au_m1(z1)) represents the number of times the surname is au_x(z1) and the first character of the given name is au_m1(z1). count(au_x(z1)) represents the number of times au_x(z1) is used as the surname;

[0017]

[0018] count(au_m i-1 (z1), au_m i (z1)) represents the number of times the (i - 1)-th character of the given name is au_m i-1 (z1) and the i-th character of the given name is au_m i (z1), and count(au_m i-1 (z1)) represents the number of times the (i - 1)-th character of the given name is au_m i-1 (z1).

[0019] 3) Obtain the similarity sim(T(z1), T(z2)) of the text content summaries of the resources based on the siamese algorithm;

[0020] where z1 and z2 are the aforementioned resource texts; T(z1), T(z2) represent the text content summaries of z1 and z2;

[0021] 4) Calculate the weighted content similarity of names R(z1, z2);

[0022]

[0023] Among them, scarcity(au(z1)) is the scarcity of the name of the author of the aforementioned resource z1; represents the logarithmic operation of scarcity(au(z1)) (base 10); sim(T(z1), T(z2)) is the similarity of the text content summaries of the aforementioned z1 and z2. If R(z1, z2) > the threshold L, then it is considered that the authors of resources z1 and z2 are the same person.

[0024] Further, perform preliminary processing on the resource text to remove the garbled or irrelevant content in the text.

[0025] Further, for knowledge resources such as books, papers, and periodicals that contain abstracts, directly obtain the abstract content and the title to jointly form the resource abstract. For text type knowledge resources that do not contain abstracts, use the TextRank method to obtain the abstract content and the title to jointly form the resource abstract. For knowledge resources of PPT and PDF types, use the OpenXML and pdf2word tools to parse out the text content, and then use the TextRank method to obtain the abstract content and the title to jointly form the resource abstract. For knowledge resources of audio and video types, use speech recognition tools to obtain the text content, and use the TextRank method to obtain the abstract content and the title to jointly form the resource abstract.

[0026] The advantages and positive effects of the present invention are as follows:

[0027] The purpose of the present invention is to accurately analyze the author identity of the resource content. The principle is to comprehensively utilize the text content similarity and name scarcity: that is, if the names of the same-name authors of two resources are very rare and the probability of having the same name is small, then unless the contents of these two resources are too obviously different, their authors should be the same person. And when the names of the same-name authors of two resources are very common, unless the content similarity of these two resources is very high, then their authors should not be the same person, so as to accurately obtain the author identity of the resource content. The present invention realizes that only the resource content and the author name are needed to accurately obtain the relationship between the knowledge resource and the author identity under the condition of less available information. Description of the Drawings

[0028] Figure 1 The flowchart of the specific embodiment of the present invention is shown. Specific Embodiments

[0029] The present invention takes book and video resources as examples, as Figure 1 shown, the specific steps include:

[0030] (1) Obtain the text content of the resource

[0031] Resource z1: The book "Introduction to and Practice of Kunpeng Architecture", series name: "Series on Computer Technology Development and Application", author: "Zhao Jiani", abstract: "This book first explains the origin of the Kunpeng architecture and the composition of the Kunpeng ecosystem, and builds a Kunpeng development environment. Then it details the reasons, methods, and the Kunpeng development kit for assisting in the migration of applications from the x86 architecture to the Kunpeng architecture. Finally, it introduces Kunpeng certification and how to pass the compatibility certification of Kunpeng cloud services. This book is targeted at beginners who hope to understand what the Kunpeng architecture is and those interested in Kunpeng. It is also of great reference value to developers with a certain technical foundation who hope to use the Kunpeng architecture in their work, architects who design Kunpeng architecture systems, and testers and developers responsible for migrating products to the Kunpeng platform."

[0032] Resource z2: Video, title: "Overview of the Android System", author: "Zhao Jiani".

[0033] For text resources such as books and periodicals, directly obtain their text content as the basic content for analysis.

[0034] For knowledge resources in the PPT and PDF formats, use tools such as OpenXML and pdf2word to parse out their text content.

[0035] For resources such as audio and video, use speech recognition technology to convert them into text content.

[0036] (2) Remove garbled or irrelevant content from the content. If the creative content is long, use the TextRank method to obtain the content summary. Specifically:

[0037] Resource z1: "Series on Computer Technology Development and Application + This book first explains the origin of the Kunpeng architecture and the composition of the Kunpeng ecosystem, and builds a Kunpeng development environment. Then it details the reasons, methods, and the Kunpeng development kit for assisting in the migration of applications from the x86 architecture to the Kunpeng architecture. Finally, it introduces Kunpeng certification and how to pass the compatibility certification of Kunpeng cloud services. This book is targeted at beginners who hope to understand what the Kunpeng architecture is and those interested in Kunpeng. It is also of great reference value to developers with a certain technical foundation who hope to use the Kunpeng architecture in their work, architects who design Kunpeng architecture systems, and testers and developers responsible for migrating products to the Kunpeng platform."

[0038] Resource z2: Use a speech recognition tool to obtain the text content (longer) of the video. Use the TextRank algorithm to obtain the abstract of the text: "Android is an open-source, Linux-based mobile device operating system, such as smartphones and tablets. Android was developed by the Open Handset Alliance led by Google and other companies. Android provides a unified approach to application development, which means that developers only need to develop for Android, and their applications can then run on different Android-powered mobile devices. Google released the first test version of the Android Software Development Kit (SDK) in 2007, and the first commercial version, Android 1.0, was released in September 2008. On June 27, 2012, at the Google I / O conference, Google announced the release of Android version 4.1 Jelly Bean. Jelly Bean is a progressive update in terms of features and performance, mainly aimed at improving the user interface. The Android source code is under a free and open-source software license. Most of the code released by Google follows the Apache License 2.0, and changes to the Linux kernel follow the GNU General Public License version 2."

[0039] (3) For the case where the authors of two resources z1 and z2 have the same name, calculate the rarity of the name

[0040]

[0041] Where F(au_x(z1)) represents the probability that the surname of the author of z1 (e.g., 'Zhang', 'Qiu') is used as a surname. To calculate the above probability, a namebase needs to be prepared in advance. The namebase can be composed by collecting real names in various ways. Based on the namebase, the probability is calculated by the following formula:

[0042]

[0043] count(au_x(z1)) represents the number of times au_x(z1) is used as a surname, and size(namebase) represents the size of the namebase, that is, the total number of names.

[0044] F(au_m1(z1)|au_x(z1)) represents the probability that the first character of the name is au_m1(z1) when the surname is au_x(z1):

[0045]

[0046] count(au_x(z1), au_m1(z1)) represents the number of times the surname is au_x(z1) and the first character of the given name is au_m1(z1). count(au_x(z1)) represents the number of times au_x(z1) is used as a surname.

[0047] Similarly,

[0048]

[0049] count(au_m i-1 (z1), au_m i (z1)) represents the number of times the (i - 1)-th character of the given name is au_m i-1 (z1) and the i-th character is au_m i (z1). count(au_m i-1 (z1)) represents the number of times the (i - 1)-th character of the given name is au_m i-1 (z1). num(au_m(z1)) represents the number of characters in the given name.

[0050] Before calculating the above probabilities, the name has been added to the name database. Therefore, the calculated probability will not be 0.

[0051] Based on the scarcity, calculate Generally, because the scarcity value is generally very small, for example: 10 -6 , and the ratio of the scarcities of two names is very large.

[0052] In this embodiment, "Zhao Jiani" is added to the name database. Calculate respectively:

[0053] (1) The probability that the surname "Zhao" appears in the name database is 3%;

[0054] (2) The probability that the surname is "Zhao" and the first character of the given name is "Jia" is 0.5%;

[0055] (3) The probability that the first character of the given name is "Jia" and the second character is "Ni" is 0.1%;

[0056] Therefore, scarcity('Zhao Jiani') = 3% * 0.5% * 0.1% = 1.5 * 10 -6

[0057]

[0058] (4) Obtain the similarity sim(T(z1), T(z2)) of the resource summary based on the siamese algorithm; the text similarity of the resource profile in the specific embodiment of the present invention is 0.453.

[0059] (5) Calculate the relevance R(z1, z2) between the computing resource text and the author identity;

[0060]

[0061] If R(z1, z2) > the threshold L (the threshold L is usually defined as 0.3), then it is considered that the authors of resources z1 and z2 are the same person.

[0062] In a specific embodiment of the present invention, R(z1, z2) = 0.8283 * 0.453 = 0.375, which is greater than the threshold 0.3. Therefore, the author identities of these two resources are considered to be the same person.

[0063] The embodiments of the present invention are not intended to limit the present invention. Any person skilled in the art can make many possible changes and modifications to the technical solution of the present invention, or modify it into an equivalent embodiment with equivalent changes, without departing from the scope of the technical solution of the present invention. Therefore, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of protection of the technical solution of the present invention.

[0064] Reference materials:

[0065] [1] Yoon Kim: Convolutional Neural Networks for Sentence Classification. EMNLP 2014. https: / / arxiv.org / pdf / 1408.5882.pdf

[0066] [2] Jonas Mueller, Aditya Thyagarajan: Siamese recurrent architectures for learning sentence similarity. AAAI'16: Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence (2016).

Claims

1. A method for analyzing the relationship between knowledge resources and author identities, the steps of which include: 1) For resource texts z1 and z2 in different media and fields, respectively extract resource summaries T(z1) and T(z2); 2) If the authors of z1 and z2 have the same name, calculate the scarcity of this name: Among them, au(z1) represents the author of the resource text z1; au_x(z1) represents the surname of the author of z1; au_m(z1) represents the given name of the author of z1; au_m i (z1) represents the i-th character of the given name of the author of z1; F(au_x(z1)) represents the probability that the surname of the author of z1 is used as a family name; F(au_m i (z1)|au_x(z1)) represents the probability that the first character of the given name is au_m1(z1) when the family name is au_x(z1): the probability of the family name count(au_x(z1)) represents the number of times au_x(z1) is used as a family name; namebase represents the name database, and size(namebase) represents the size of namebase, that is, the number of names contained in namebase; count(au_x(z1), au_m1(z1)) represents the number of times the family name is au_x(z1) and the first character of the given name is au_m1(z1), and count(au_x(z1)) represents the number of times au_x(z1) is used as a family name; num(au_m(z1)) is the number of characters of au_m(z1); count(au_m i-1 (z1), au_m i (z1)) represents the number of times the (i - 1)-th character of the given name is au_m i-1 (z1) and the i-th character of the given name is au_m i (z1); count(au_m i-1 (z1)) represents the number of times the (i - 1)-th character of the given name is au_m i-1 (z1); 3) Obtain the similarity sim(T(z1), T(z2)) of the content summaries of the resource texts based on the siamese algorithm; 4) Calculate the name-weighted content similarity R(z1, z2); Among them, represents the logarithmic operation of scarcity(au(z1)) with base 10; if R(z1, z2) > threshold L, then the authors of resource text z1 and resource text z2 are considered to be the same person.

2. The method for analyzing the relationship between knowledge resources and author identity according to claim 1, wherein Preprocess the resource text to remove the garbled or irrelevant content in the text.

3. The method for analyzing the relationship between knowledge resources and author identity according to claim 1, characterized in that, For knowledge resources such as books, papers, and periodicals with abstracts, directly obtain the abstract content and the title to jointly form the resource summary.

4. The method for analyzing the relationship between knowledge resources and author identity according to claim 1, characterized in that, For text-type knowledge resources without abstracts, use the TextRank method to obtain the abstract content and the title to jointly form the resource summary.

5. The method for analyzing the relationship between knowledge resources and author identity according to claim 1, characterized in that For knowledge resources of PPT and PDF types, use the OpenXML and pdf2word tools to parse out the text content, and then use the TextRank method to obtain the abstract content and the title to jointly form the resource summary.

6. The method for analyzing the relationship between knowledge resources and author identities according to claim 1, wherein For knowledge resources of audio and video types, use speech recognition tools to obtain the text content, and use the TextRank method to obtain the abstract content and the title to jointly form the resource summary.

Citation Information

Patent Citations

  • Literature author name duplication disambiguation method and literature author name duplication disambiguation construction system

    CN112131872A

  • Author name disambiguation method and author name disambiguation device

    CN114510568A