Name identification device and program

The name identification device addresses the challenge of identifying unique names in social networking services by using a system that extracts and compares alternative names from user conversations, effectively identifying certified names without prior database storage of characteristic words, enhancing name recognition accuracy.

JP2025116582APending Publication Date: 2025-08-08NIPPON HOSO KYOKAI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024011085
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-29
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Existing name identification technologies, such as those described in Patent Document 1, require pre-storing characteristic words related to formal names in a database, making them ineffective for unique names used in social networking services where users express themselves freely, and struggle with extremely short sentences lacking context.

Method used

A name identification device that includes a name database storage unit, an alternative name list storage unit, a behavioral history acquisition unit, a used name extraction unit, and a name identification unit, which identifies certified names by comparing alternative names extracted from user conversations without prior database storage of characteristic words.

Benefits of technology

Enables accurate identification of certified names from user-generated content on social networking services, even when using unique or contextually limited names, by leveraging conversation history to establish associations between alternative and certified names.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025116582000001_ABST
    Figure 2025116582000001_ABST
Patent Text Reader

Abstract

To provide a name identification device which can identify a certified name corresponding to an identification object name on the basis of conversations without previously storing information such as feature words in a database or the like.SOLUTION: The name identification device comprises a name database storage unit, an alternative name list storage unit, a behavioral history acquisition unit, a use name extraction unit, and a name identification unit. The name database storage unit stores therein certified names. The alternative name list storage unit stores therein a list of alternative names related to an identification target name. The behavioral history acquisition unit acquires a text of a history of conversations between users. The use name extraction unit extracts alternative names related to the identification target name from the text acquired by the behavioral history acquisition unit and stores the extracted alternative names in the alternative name list storage unit. The name identification unit identifies a certified name corresponding to the identification target name by collating the list of alternative names stored in the alternative name list storage unit with the certified names stored in the name database storage unit.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a name identification device and a program. [Background technology]

[0002] Social networking services (SNS) are widely used. In social networking services, users can post text and other information by operating their terminal devices. Technology is used to analyze the text posted by many users. By analyzing the text posted by users on social networking services, it is possible to estimate the trends in the interests of many users and the interests of individual users. It is also conceivable to implement a service that recommends various content to users based on the results of this estimation.

[0003] For example, if a particular user frequently posts about a particular person on a social networking service, a service could be provided that recommends content (such as television programs) in which that person appears.

[0004] In addition, if a user frequently posts text that includes specific expressions (proper nouns), such as specific place names or specific organization names (e.g., company names), it is possible to recommend content related to the place represented by the place name or the organization represented by the organization name.

[0005] Incidentally, the names used to refer to people, organizations, places, etc. in posts on social networking services are not necessarily unique among all users who post. On social networking services, users often freely express their feelings, and each user tends to refer to objects (people, organizations, places, etc.) by a variety of names. These various names are often different from the official names (people's names, organization names, place names, etc.). In other words, there is a need for technology that can identify the objects that users are referring to in posts on social networking services.

[0006] One possible method is to register in advance what various names refer to.

[0007] For example, one possible method is to register in advance the nicknames and abbreviations of well-known people as a group, associated with their official name. In other words, nicknames and abbreviations such as A1, A2, A3, etc. are registered as a group associated with an object (person, etc.) with the official name A. In this way, when a posted text contains a nickname such as A1 or A2, the text can be interpreted as being written about the object with the official name A.

[0008] A similar method can also be used to deal with variations in the spelling of names. For example, names such as "David," "David," and "David" in katakana could be registered in advance as a group in association with the English name "David." This allows a posted text containing any of the names "David," "David," or "David" to be interpreted as being about the original "David."

[0009] Patent Document 1 describes a technology for identifying a person's formal name from an incomplete name. The technology described in Patent Document 1 stores a formal person's name (called a complete name), multiple characteristic words related to the person's name, and the weight of each characteristic word in a database in advance. [Prior art documents] [Patent documents]

[0010] [Patent Document 1] Japanese Patent Application Laid-Open No. 2009-181183 Summary of the Invention [Problem to be solved by the invention]

[0011] However, in order to use the technology described in Patent Document 1, it is necessary to store in advance in a database a plurality of characteristic words related to formal person names (full names) and the weights of each characteristic word.

[0012] Furthermore, the technology described in Patent Document 1 is a method that is effective only for names that are commonly used in society.

[0013] Furthermore, because social networking services allow users to freely express their feelings, they often use names that are unique to each user and are not generally used in society. In such cases, methods such as registering commonly known nicknames in advance are of no use.

[0014] Furthermore, among social networking services, those where short sentences tend to be posted in particular often have extremely short sentences posted, with almost no other words before or after the user's name. In such cases, it is difficult to extract characteristic words that characterize the name of the object (such as a person's name) referred to by the name from within the posted sentence. For this reason, the method of Patent Document 1 is useless. In other words, there is a problem in that it is difficult to identify the object referred to by the name in the sentence.

[0015] The present invention was made based on the recognition of the above problem, and aims to provide a name identification device and program that can identify a certified name (full name) corresponding to a name to be identified based on text of conversation between users, without having to previously store information on multiple characteristic words related to full names in a database or the like. [Means for solving the problem]

[0016] [1] In order to solve the above problem, a name identification device according to one aspect of the present invention includes a name database storage unit that stores in advance certified names, which are names that have been certified for an object; an alternative name list storage unit that stores a list of alternative names related to the identification target name that is the object of identification; a history acquisition unit that acquires text of a conversation history between users; a used name extraction unit that extracts alternative names related to the identification target name from the text acquired by the history acquisition unit and stores the extracted alternative names in the alternative name list storage unit; and a name identification unit that identifies the certified name that corresponds to the identification target name by comparing the list of alternative names stored in the alternative name list storage unit with the certified name stored in the name database storage unit.

[0017] [2] Furthermore, one aspect of the present invention is that in the name identification device of [1] above, the used name extraction unit extracts the alternative name from the text of the conversation history between a specific user who posted text including the name to be identified, and the specific user and other users.

[0018] [3] Furthermore, one aspect of the present invention is that in the name identification device of [2] above, the used name extraction unit further extracts the alternative names from the text of the conversation history between the other user and yet another user, and from the text of the conversation history between the other user and yet another user recursively, with the other user as the other user.

[0019] [4] Furthermore, one aspect of the present invention is that in any of the name identification devices described above in [1] to [3], the alternative name list storage unit stores frequency information regarding the frequency of appearance of each of the alternative names, and the name identification unit compares only the alternative names whose frequency information satisfies a predetermined condition with the recognized name, thereby identifying the recognized name corresponding to the name to be identified.

[0020] [5] Furthermore, one aspect of the present invention is that in any of the name identification devices described above in [1] to [4], when the name identification unit is able to identify the certified name corresponding to the name to be identified, it stores the name to be identified in the name database storage unit as a certified alternative name, which is an alternative name certified in relation to the certified name.

[0021] [6] Furthermore, one aspect of the present invention is that in any of the name identification devices [1] to [5] above, the history acquisition unit acquires text of a history of conversations between users based on a chain of posts in a social networking service.

[0022] [7] Furthermore, one aspect of the present invention is a program for causing a computer to function as a name identification device, including: a name database storage unit that stores in advance certified names, which are names that have been certified for an object; an alternative name list storage unit that stores a list of alternative names related to the identification target name that is the object of identification; a history acquisition unit that acquires text of a conversation history between users; a used name extraction unit that extracts alternative names related to the identification target name from the text acquired by the history acquisition unit and stores the extracted alternative names in the alternative name list storage unit; and a name identification unit that identifies the certified name that corresponds to the identification target name by comparing the list of alternative names stored in the alternative name list storage unit with the certified name stored in the name database storage unit. [Effects of the Invention]

[0023] According to the present invention, the name identification device compares a list of alternative names related to the name to be identified with the certified names stored in the name database storage unit, thereby enabling the name identification device to identify the certified name corresponding to the name to be identified. [Brief explanation of the drawings]

[0024] [Figure 1] 1 is a block diagram showing a schematic configuration of a system including a name identification device according to an embodiment of the present invention. [Figure 2] FIG. 2 is a schematic diagram showing an example of text posted by a particular user on a social networking service, which is used in the embodiment. [Figure 3] 1 is a schematic diagram showing an example of a conversation formed by a chain of posts on a social networking service, which is a target of processing in the embodiment. [Figure 4] FIG. 10 is a schematic diagram (1 / 2) showing an example of text posted in multiple conversations that start with one post by a user and are subject to processing in the embodiment. [Figure 5]FIG. 2 is a schematic diagram (2 / 2) showing an example of text posted in multiple conversations that start with one post by a user and that is the subject of processing in this embodiment. [Figure 6] FIG. 10 is a schematic diagram showing an example of text posted in another conversation that is processed in the embodiment. [Figure 7] 7 is a schematic diagram showing names used by multiple users for the same object when the name identification device according to the embodiment acquires the data shown in FIGS. 4 and 5 and the data shown in FIG. 6. FIG. [Figure 8] 2 is a schematic diagram showing a list of alternative names registered in an alternative name list storage unit and the frequency of appearance of each alternative name in the embodiment. FIG. [Figure 9] FIG. 7 is a schematic diagram showing an example of a post text that is a processing target in the embodiment and is different from those in FIGS. 4 to 6. [Figure 10] 10 is a schematic diagram showing an example of a list of alternative names obtained based on a large number of conversations in the embodiment. FIG. [Figure 11] FIG. 11 is a schematic diagram in which the ratio of the number of times each alternative name appears (number of times ratio) is added to the alternative name list shown in FIG. [Figure 12] 10 is a schematic diagram showing the range of different names related to the name to be identified that the name identification device according to the embodiment acquires. FIG. [Figure 13] FIG. 2 is a block diagram showing an example of the internal configuration of a name identification device (including modifications) according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0025] Next, an embodiment of the present invention will be described with reference to the drawings. This embodiment is intended to solve the above-mentioned problems. That is, in this embodiment, when a user uses a special name for an object (person, organization, place, etc.) in a sentence posted on a social networking service, the name identification device 1 identifies which object the name refers to. Furthermore, the name identification device 1 identifies which object the name refers to even if the name is not necessarily a name that is widely used in society. Furthermore, the name identification device 1 can identify which object the name refers to even if the posted sentence is extremely short and almost no other words are used before or after the name.

[0026] FIG. 1 is a block diagram showing a schematic configuration of a system according to this embodiment. As shown in the figure, the system 10 includes a name identification device 1, a PDS device 2, a server device 3, and a communication network 8. The name identification device 1, the PDS device 2, and the server device 3 can each be realized, for example, by a computer and a program. Each device also has a storage means as needed. The storage means is, for example, a program variable or memory allocated by program execution. Non-volatile storage means such as a magnetic hard disk drive or a solid-state drive (SSD) may also be used as needed. At least some of the functions of each functional unit may be realized as a dedicated electronic circuit rather than a program. The functions of each component of the system 10 are as follows:

[0027] The name identification device 1 identifies a name to be identified (referred to as an "identification target name"). Specifically, the name identification device 1 identifies a certified name corresponding to the identification target name based on a conversation between users that includes the identification target name. A more detailed functional configuration of the name identification device 1 will be described later.

[0028] The PDS device 2 is a device that stores and manages personal data. PDS stands for "Personal Data Store." A PDS device allows users to manage their own personal data. The PDS itself is realized using existing technology. The PDS device 2 can store text (including conversations) posted by users on social networking services. The PDS device may also store and manage log data of user behavior on services other than social networking services. Although only one PDS device 2 is shown in the figure, the system 10 may include multiple PDS devices 2. Multiple users may each use a different PDS device 2.

[0029] The server device 3 is a device that provides a social networking service. The system 10 may be configured to include multiple server devices 3. Here, to distinguish between the individual server devices 3, different reference symbols such as server device 3-1, server device 3-2, and server device 3-3 may be used.

[0030] The server device 3 accepts sentences (text, etc.; hereinafter, sometimes referred to as "posts") posted by users, sent from the users' terminal devices (not shown), and stores and manages the information on the posts. The server device 3 can arrange the text, etc. posted by users into a format corresponding to the services it provides and present it to the terminal devices of the poster and other users. The server device 3 can manage the relationship between a post and other posts. The server device 3 manages replies from other users to a post. The server device 3 can manage a chain of replies to a post as a "conversation" and present it on the user's terminal device. Conversations will be explained further below.

[0031] The social networking service itself provided by the server device 3 is realized by using existing technology.

[0032] The communication network 8 is a network that realizes data communication between the name identification device 1, the PDS device 2, and the server device 3. On the communication network 8, communication is carried out using a standard protocol such as the Internet Protocol (IP), which enables communication between different computers.

[0033] As shown in FIG. 1, the name identification device 1 includes a behavioral history acquisition unit 101, a name extraction unit 102, a used name extraction unit 103, a name identification unit 104, a name database storage unit 111, and an alternative name list storage unit 112.

[0034] The behavioral history acquisition unit 101 acquires and stores a history (text, etc.) of posts that users have made in the past on the social networking service from a server device 3 (for example, any of the server devices 3-1 to 3-3) via the communication network 8. Each post acquired by the behavioral history acquisition unit 101 is designed to be able to identify which user posted it. For example, the behavioral history acquisition unit 101 can acquire a history of past posts with a user ID attached to identify the user. As will be described later, a chain of posts acquired by the behavioral history acquisition unit 101 may constitute a conversation between users.

[0035] That is, the behavior history acquisition unit 101 acquires the text of the history of conversations between users based on the chain of posts in the social networking service. In other words, the behavior history acquisition unit 101 regards the chain of posts in the social networking service as a conversation between users and acquires the text of the conversation.

[0036] The behavior history acquisition unit 101 may acquire the user's posting history from the PDS device 2 instead of acquiring the posting history from the server device 3. In other words, the PDS device 2 may acquire, for example, the posting history of a specific user from the server device 3 in advance and store it.

[0037] The behavioral history acquisition unit 101 is also simply referred to as a "history acquisition unit." That is, the behavioral history acquisition unit 101 acquires text of a conversation history between users. The behavioral history acquisition unit 101 may also acquire text of a conversation history between users from a service other than a social networking service.

[0038] The name extraction unit 102 extracts names from the history of posts (text) acquired by the behavior history acquisition unit 101. A name is a nickname that indicates an object. A name is often a proper noun (or named entity). The name extraction unit 102 can extract names from text by using, for example, an existing technology. Specifically, the existing technology is, for example, named entity recognition (NER) technology.

[0039] To extract names, the name extraction unit 102 can use, for example, a method of extracting named entities using a span-based model. Alternatively, the name extraction unit 102 may extract names using a sequential labeling method. Reference 1 below describes a method of extracting named entities using a span-based model. Reference 2 also describes a technology for labeling words using a sequential labeling method.

[0040] [Reference 1] Markus Eberts, Adrian Ulges, Span-based Joint Entity and Relation Extraction with Transformer Pre-training, the Proceedings of ECAI 2020 (DOI: 10.3233 / FAIA200321), 2020.

[0041] [Reference 2] Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, Jaewoo Kang, BioBERT: a pre-trained biomedical language representation model for biomedical text mining,Bioinformatics, 2019, 1-7,doi: 10.1093 / bioinformatics / btz682

[0042] Alternatively, the name extraction unit 102 may extract names from text based on the judgment of a person (for example, an administrator of the name identification device 1). In this case, the person refers to the text and instructs the name extraction unit 102 which part of the text corresponds to the name. Then, the name extraction unit 102 extracts the part of the text instructed by the person as the name.

[0043] The name extraction unit 102 can refer to the name database storage unit 111. In other words, the name extraction unit 102 can make an inquiry about the certified names stored in the name database storage unit 111. In addition, the name extraction unit 102 can write (register) the extracted names in the different name list storage unit 112.

[0044] The used name extraction unit 103 extracts alternative names related to the name to be identified from the text of the post or the like acquired by the behavior history acquisition unit 101, and stores the extracted alternative names in the alternative name list storage unit 112.

[0045] The used name extraction unit 103 extracts alternative names from the text of the conversation history between a specific user (user A in the example described below) who posted the text including the name to be identified, and the specific user and other users (users B, C, D, E, and F in the example described below).

[0046] The name identification unit 104 identifies a certified name corresponding to the name to be identified by comparing the list of alternative names stored in the alternative name list storage unit 112 with the certified names stored in the name database storage unit 111.

[0047] The name database storage unit 111 stores in advance certified names, which are names that have been certified for an object. A certified name is a name that is widely recognized and certified for a specific object. A certified name may also be called an "official name." A certified name may be the name of a well-known object.

[0048] The name database storage unit 111 may store one or more authorized alternative names in association with an authorized name. An authorized alternative name is an alternative name authorized through a predetermined procedure, etc. An alternative name may also be called a "nickname" or an "alias," etc.

[0049] The name database storage unit 111 may have a management function for managing the stored information. The management function may include a function for receiving queries regarding certified names from the name extraction unit 102, the name identification unit 104, etc., and returning a response thereto. The queries are used to perform operations such as referencing, updating, and inserting information. The database management function itself is realized by existing technology.

[0050] The alternative name list storage unit 112 stores a list of alternative names related to the identification target name that is the target of identification. The alternative name list storage unit 112 may also store information (frequency information) relating to the frequency with which the alternative names appear. A specific example of the alternative name list storage unit 112 storing frequency information will be described later as "Modification 1."

[0051] FIG. 2 is a schematic diagram showing an example of text posted by a particular user on a social networking service. This diagram includes six pieces of text, each numbered 1 to 6 for convenience. The poster of each of the six pieces of text shown in this diagram is "User A." Each of these six pieces of text includes a name for a person. Each of these six posts may also be a reply to another post.

[0052] The behavioral history acquisition unit 101 acquires the text of a post (including when the post is a reply to another post) from a server device 3 (for example, any of the server devices 3-1 to 3-3) or a PDS device 2 via the communication network 8. The name extraction unit 102 can extract a name related to a specific target (such as a person) from the text of the post acquired by the behavioral history acquisition unit 101.

[0053] In the example shown in FIG. 2, the name extraction unit 102 may extract names such as "A-chan," "Tom XXX," "Bukunin Tetsuya," and "Miya XXX." These names were used by user A. In the example shown in FIG. 2, posts 1, 4, and 6 include the name "A-chan." Post 2 includes the name "Tom XXX." Post 3 includes the name "Bukunin Tetsuya." Post 5 includes the name "Miya XXX."

[0054] Note that the name extraction unit 102 may extract only posts that include a name beforehand, prior to the process of extracting the name from the text of the post.

[0055] The name extraction unit 102 can inquire whether the name extracted from the text of the post is stored in the name database storage unit 111 or not.

[0056] In other words, by referring to the name database storage unit 111, the name extraction unit 102 can know which object (person, etc.) the name extracted from the text of the post refers to when used. In the example shown in FIG. 2, for example, three names, "Tom XXX," "Bukunin Tetsuo," and "Miya XXX," are registered in the name database storage unit 111. In this case, the name database storage unit 111 stores information indicating which object each of the three names, "Tom XXX," "Bukunin Tetsuo," and "Miya XXX," refers to. On the other hand, for example, the name "A-chan" is a name that has not yet been registered in the name database storage unit 111. The name extraction unit 102 can register names that are not registered in the name database storage unit 111 in the alternative name list storage unit 112. The names registered in the alternative name list storage unit 112 are names that can be processed by the name identification unit 104. In other words, the names registered in the alternative name list storage unit 112 are names that are targets or candidates for matching with the certified names registered in the name database storage unit 111.

[0057] In the following, the term "A-chan" will be used as an example for explanation. That is, the following describes a process for identifying what the term "A-chan" refers to.

[0058] Furthermore, the used name extraction unit 103 receives information on posts including a name to be identified (for example, "A-chan") from the behavior history acquisition unit 101. In other words, if "A-chan" is a name to be identified, the used name extraction unit 103 receives posts related to posts in which user A uses the name "A-chan" and in which conversations are held between user A and other users from the behavior history acquisition unit 101. In other words, the used name extraction unit 103 can receive from the behavior history acquisition unit 101 the text of conversations related to a name to be identified (for example, "A-chan").

[0059] FIG. 3 is a schematic diagram showing an example of a conversation formed by a chain of posts on a social networking service. The diagram shows the relationship between Post 1, Reply 3, Reply 5, Reply 7, and Reply 9 by User A, Reply 2 and Reply 4 by User B, and Reply 6 and Reply 8 by User C. A reply is a type of post that responds to a previous post. A reply may have subsequent replies. A reply clearly indicates which previous post it responds to. A post may be associated with two or more replies. In the illustrated example, there are two parallel replies, Reply 2 and Reply 6, for Post 1. A chain of posts can form a conversation. In the illustrated example, the chain of Post 1 (User A) → Reply 2 (User B) → Reply 3 (User A) → Reply 4 (User B) → Reply 5 (User A) is referred to as "Conversation 1." Furthermore, the chain of Post 1 (User A) → Reply 6 (User C) → Reply 7 (User A) → Reply 8 (User C) → Reply 9 (User A) is referred to as "Conversation 2." There is no particular limit to the length of a chain within a single conversation. In the illustrated example, Conversation 1 above is a dialogue exchanged only between User A and User B. Conversation 2 above is a dialogue exchanged only between User A and User C. However, a conversation may be formed through posts by three or more users.

[0060] 4 and 5 are schematic diagrams showing examples of texts of posts (including replies) in multiple conversations starting with one post (Post 1) by User A. In FIGS. 4 and 5, for convenience, each post is assigned a number (numbers 1 to 13). In the illustrated example, the chain starting with number 1 (Post 1, User A) and continuing with numbers 2 (User B) → 3 (User A) → 4 (User B) → 5 (User A) is referred to as "Conversation 1." Furthermore, the chain starting with number 1 (Post 1, User A) and continuing with numbers 6 (User C) → 7 (User A) → 8 (User C) → 9 (User A) is referred to as "Conversation 2." Furthermore, the chain starting with number 1 (Post 1, User A) and continuing with numbers 10 (User D) → 11 (User A) → 12 (User D) → 13 (User A) is referred to as "Conversation 3."

[0061] 4 and 5, names that can be extracted from the text of each post are noted. For example, the text of the first post, number 1 (post 1), is "A-chan is so cute. This photo really captures A-chan's personality and is so A-chan-like. The best photo." (Note: line breaks are omitted.) The name that the name extraction unit 102 can extract based on the text of number 1 is "A-chan." The text of post (reply) number 2 is "Aya-chan is so cute, isn't she?" The name that the name extraction unit 102 can extract based on the text of number 2 is "Aya-chan." The same applies to other texts. In the relationship between posts number 1 and number 2, the names "A-chan" and "Aya-chan" refer to the same object (person) as intended by each user. The name identification unit 104 identifies the name to be identified (e.g., "A-chan") based on the different names (but referring to the same object) among the multiple posts.

[0062] In the examples of Figures 4 and 5, the names "A-chan" used by user A, "Aya-chan" used by user B, "Yasuda Aya〇-san" used by user C, and "Ayachi" used by user D refer to the same object (person).

[0063] The used name extraction unit 103 can receive the conversation examples shown in Figures 4 and 5 from the behavior history acquisition unit 101. When the used name extraction unit 103 receives the conversation data shown in Figures 4 and 5 (the names included in the text can be extracted by the name extraction unit 102), it registers the names "Aya-chan" used by user B, "Yasuda Aya〇-san" used by user C, and "Ayachi" used by user D as names related to the name "A-chan" used by user A in the different name list storage unit 112. In other words, by the used name extraction unit 103 registering the names "Aya-chan", "Yasuda Aya〇-san", and "Ayachi" related to the name "A-chan" in the different name list storage unit 112, the name identification device 1 can treat these names as names for the same object (here, the same person).

[0064] FIG. 6 is a schematic diagram showing an example of text of a post different from those shown in FIGS. 4 and 5. The data shown in FIG. 6 also includes the nickname "A-chan" used by user A and nicknames used by other users. Based on the data shown in FIG. 6, other nicknames related to the nickname "A-chan" may be extracted and registered in the alternative nickname list storage unit 112. Specifically, in the data shown in FIG. 6, the text of posts numbered 14, 16, and 19 (including replies) includes the nickname "A-chan" used by user A. That is, the nickname extraction unit 102 may extract the nickname "A-chan" from each of posts numbered 14, 16, and 19. Furthermore, the text of posts numbered 15 and 17 includes the nickname "Yasuda Aya-san" used by user E. That is, the nickname extraction unit 102 may extract the nickname "Yasuda Aya-san" from each of posts numbered 15 and 17. Furthermore, the text of the posts numbered 18 and 20 includes the nickname "Yasu" used by user F. That is, the nickname extraction unit 102 can extract the nickname "Yasu" from each of the posts numbered 18 and 20.

[0065] In other words, from the data shown in Figures 4 and 5 and the data shown in Figure 6, the used name extraction unit 103 obtains multiple names, such as "Aya-chan" (used by user B), "Yasuda Aya〇-san" (used by users C and E), "Ayachi" (used by user D), and "Yasu" (used by user F), as names related to the name "A-chan" used by user A, and registers these names in the alternative name list storage unit 112.

[0066] In this way, based on the data of conversations in which user A participates, acquired by the behavioral history acquisition unit 101, information on names (alternative names) related to the name "A-chan" is accumulated in the alternative name list storage unit 112.

[0067] Figure 7 is a schematic diagram showing how each user, including user A, calls the subject (person) that user A calls "A-chan" when the data shown in Figures 4 and 5 and the data shown in Figure 6 are acquired. In other words, in the text posted by user A, the subject is called "A-chan." In the text posted by user B, the subject is called "Aya-chan." In the text posted by user C, the subject is called "Yasuda Aya〇-san." In the text posted by user D, the subject is called "Ayachi." In the text posted by user E, the subject is called "Yasuda Aya〇-san." In the text posted by user F, the subject is called "Yasu."

[0068] FIG. 8 is a schematic diagram showing a list of alternative names registered in the alternative name list storage unit 112 and the frequency of appearance of each name, based on the data shown in FIGS. 4 and 5 and the data shown in FIG. 6. The table shown in FIG. 8 includes fields for the name to be identified, the alternative name, and the number of times the alternative name appeared in conversations with user A. In the example shown, the name to be identified is "A-chan." Alternative names corresponding to the name "A-chan" include "Aya-chan," "Yasuda Aya-san," "Ayachi," and "Yasu." The number of times each of these alternative names appeared in conversations with user A is as follows: the alternative name "Aya-chan" appeared once in conversation 1, the alternative name "Yasuda Aya-san" appeared twice in conversations 2 and 4, the alternative name "Ayachi" appeared once in conversation 3, and the alternative name "Yasu" appeared once in conversation 5. The alternative name list storage unit 112 also stores the number of times each alternative name appears. Alternatively, the number of times each alternative name appears can be found based on the information stored in the alternative name list storage unit 112.

[0069] The name identification unit 104 reads out the name of the target to be identified and a list of alternative names related to that name from the alternative name list storage unit 112. The name of the target to be identified is a name for which it is unclear which object it refers to. The alternative names related to that name are names used by other users (users B, C, D, E, and F in the examples of FIGS. 4 to 6) in conversation with the user who uses the name of the target to be identified (user A in the examples of FIGS. 4 to 6). It can be assumed that the name of the target to be identified and the alternative names refer to the same object (such as a person).

[0070] As mentioned above, the alternative name list storage unit 112 stores information on the number of times each alternative name appears. In other words, the name identification unit 104 can read multiple alternative names and information on the number of times each alternative name appears from the alternative name list storage unit 112.

[0071] The name identification unit 104 can make an inquiry to the name database storage unit 111 about each alternative name read out from the alternative name list storage unit 112. In other words, the name identification unit 104 can inquire whether each alternative name is registered in the name database storage unit 111. When the information shown in FIG. 8 is registered in the alternative name list storage unit 112 as a list of alternative names, the name identification unit 104 inquires whether each alternative name is registered in the name database storage unit 111. Specifically, an inquiry is made about four alternative names: "Aya-chan," "Yasuda Aya〇-san," "Ayachi," and "Yasu."

[0072] In addition, in the case of a person's name, honorifics such as "chan," "san," and "kun" may be included in the name. When the alternative name includes such honorifics, the name identification unit 104 may query the name database storage unit 111 for both names that include these honorifics and names that exclude these honorifics. With regard to the alternative names included in FIG. 8, the name identification unit 104 may query the name database storage unit 111 for both "Aya-chan" and "Aya." Similarly, the name identification unit 104 may query the name database storage unit 111 for both "Yasuda Aya〇-san" and "Yasuda Aya〇."

[0073] If the alternative name inquired by the name identification unit 104 is registered in the name database storage unit 111, the name identification unit 104 identifies the alternative name as the object indicated by the name of the original identification target. For example, among the examples of alternative names shown in FIG. 8, the name obtained by excluding honorifics from "Yasuda Aya〇-san" is "Yasuda Aya〇." If this "Yasuda Aya〇" is registered in the name database storage unit 111, the name identification unit 104 identifies (or presumes) that the official name of the original identification target name "A-chan" (the name used by user A) is "Yasuda Aya〇."

[0074] The name identification unit 104 may make an estimation based on the likelihood value. For example, the name database storage unit 111 may store nicknames such as "Aya-chan," "Ayachi," and "Yasu" in association with the formal name "Yasuda Aya." In such a case, the list of alternative names shown in FIG. 8 includes "Aya-chan," "Ayachi," and "Yasu," increasing the likelihood that the original name to be identified, "A-chan," refers to the subject (person) with the formal name "Yasuda Aya." In other words, the name identification unit 104 can identify (estimate) the name to be identified while changing the likelihood of the formal name depending on whether the alternative name queried in the name database storage unit 111 is registered as a nickname.

[0075] Furthermore, the name identification unit 104 may register alternative names that have not yet been registered in the name database storage unit 111 in the name database storage unit 111. For example, after identifying the name of the person to be identified, the name identification unit 104 registers alternative names associated with the name of the person to be identified that have not yet been registered in the name database storage unit 111 as nicknames related to the identified formal name. Specifically, when the name of the person to be identified is "A-chan," the name identification unit 104 identifies that the formal name of "A-chan" is "Yasuda Ayako." At this time, among the alternative names of the name "A-chan" of the person to be identified, such as "A-chan," "Ayachi," or "Yasu," the name identification unit 104 registers an alternative name that has not yet been registered in the name database storage unit 111 as a nickname associated with the formal name "Yasuda Ayako" as a certified alternative name. In this way, by the name identification unit 104 registering an unregistered nickname in the name database storage unit 111, when identifying other names later, the likelihood can be calculated based on those recognized names. In other words, it becomes easier to identify names to be identified later.

[0076] That is, when the name identification unit 104 is able to identify a certified name corresponding to the name to be identified, it stores the name to be identified as a certified alternative name, which is an alternative name certified in relation to the certified name, in the name database storage unit 111. As described above, a certified alternative name is an alternative name certified by the name identification device 1 through a predetermined procedure or the like.

[0077] This embodiment can also be implemented in the following modified examples. Note that multiple modified examples may be implemented in combination as long as they can be combined.

[0078] [Variation 1] Next, let us consider the case where, when identifying a target name, a name that is different from the target referred to by the target name appears in the conversation.

[0079] FIG. 9 is a schematic diagram showing an example of post text that is different from those in FIGS. 4 to 6. The posts shown in FIG. 9 are numbered 22 to 25. The post numbered 22 is by user A (post 3). Following this post numbered 22 by user A, there is a chain of posts numbered 23 (user E), numbered 24 (user A), and numbered 25 (user E). This series of posts numbered 22 to 25 constitutes conversation 6.

[0080] In Conversation 6 shown in Figure 9, the text of post number 22 (by user A) contains the name "A-chan." Additionally, the text of posts numbered 23 and 25 (by user E) contains the name "Kobayashi △ko-chan." In other words, in a conversation about an object (person) that user A calls "A-chan," an object (person) referred to by the name "Kobayashi △ko-chan" appears. Here, the object referred to by the name "A-chan" and the object referred to by the name "Kobayashi △ko-chan" are different.

[0081] When identifying the name "A-chan" to be identified based on conversation 6 shown in Figure 9, even if the name "Kobayashi △ko-chan" is registered in the name database storage unit 111 (the official name is "Kobayashi △ko"), it is desirable for the name identification device 1 to operate in such a way that it does not include "Kobayashi △ko" as a name for the object referred to by the name "A-chan."

[0082] To achieve this, the name identification device 1 may be configured and operated as follows. That is, the name identification device 1 (name identification unit 104) acquires many conversations containing the name to be identified, "A-chan" (user A), and collects alternative names that are thought to be related to the name "A-chan." Then, the name identification device 1 (name identification unit 104) excludes alternative names that appear extremely few times compared to the total number of conversations (or the total number of times alternative names appear). Then, the name identification unit 104 queries the name database storage unit 111 based on the list of alternative names remaining after the exclusion.

[0083] FIG. 10 is a schematic diagram showing an example of an alternative name list obtained based on a large number of (e.g., 106) conversations. The table shown in FIG. 10 lists a large number of alternative names corresponding to the name of the target to be identified, "A-chan" (the name used by user A), and also contains information on the number of times each alternative name appeared during the conversation with user A. The list in FIG. 10 includes the aforementioned name "Kobayashi △ko-chan," but the names "A-chan" and "Kobayashi △ko-chan" refer to different objects (people).

[0084] FIG. 11 is a schematic diagram illustrating the frequency ratio (frequency ratio) of each alternative name in the alternative name list shown in FIG. 10 . The frequency ratio is expressed as a percentage and is rounded to one decimal point. If the name identification unit 104 were to query the name database storage unit 111 only for alternative names with a frequency ratio of 5.0% or more, the only alternative names that could be queried would be those marked with a "〇." In other words, the alternative names that could be queried in the name database storage unit 111 are the five alternative names "Aya-chan," "Yasuda Aya〇-san," "Yasu," "Aya-chan," and "Yasu," as well as names excluding honorifics such as "chan" and "san." By restricting the alternative names to be matched with the certified name based on the frequency of appearance of the alternative names, it is possible to prevent "Kobayashi △ko-chan" (or "Kobayashi △ko"), which refers to a different subject from "A-chan," from being identified as the official name of "A-chan." In other words, it is possible to avoid erroneous identification.

[0085] In the first modification, the information on the number of times an alternative name appears and the information on the ratio of the number of times is frequency information related to the frequency of appearance of each alternative name. That is, in the first modification, the alternative name list storage unit 112 stores frequency information related to the frequency of appearance of each alternative name. Furthermore, the name identification unit 104 identifies the certified name corresponding to the name to be identified by comparing only alternative names whose frequency information satisfies a predetermined condition with the certified name. Here, an example of the "predetermined condition" is the condition that "the ratio of the number of times of appearance is 5.0% or more," as described with reference to FIG. 11. However, the threshold value for the ratio of the number of times of appearance (5.0%) is merely an example and can be set arbitrarily. Furthermore, the name identification unit 104 may compare only alternative names that satisfy the condition with the certified name according to other conditions related to the frequency information, not limited to the conditions exemplified here.

[0086] [Variation 2] In the example posts shown in the data of Figures 4 to 6, the name "Yasuda Aya〇-san" was used by users C and E. Therefore, the name "Yasuda Aya〇" became the subject of a query to the name database storage unit 111, and it was possible to identify "Yasuda Aya〇" as the formal name of the name to be identified, "A-chan."

[0087] However, if the amount of conversation acquired is too small, it is possible that the list of alternative names that the name identification unit 104 reads from the alternative name list storage unit 112 does not include any of the names registered in the name database storage unit 111.

[0088] Taking the above-mentioned cases into consideration, it is conceivable to expand the range of conversations that the name identification device 1 acquires. In other words, the name identification device 1 may use the fact that user A calls a person "A-chan" and user B calls him "Aya-chan" to extract further alternative names from a conversation between user B and yet another user on a social networking service.

[0089] 4 (user B), the behavioral history acquisition unit 101 acquires and stores the text of the conversation including the reply to the post in which user B used the name "Aya-chan." As a result, using the method already explained, the used name extraction unit 103 extracts alternative names used by other users for the object (person, etc.) that user B calls "Aya-chan," and stores them in the alternative name list storage unit 112.

[0090] As a result, the name identification unit 104 reads out a list of alternative names other than "A-chan" and "Aya-chan" for the target that user A calls "A-chan" and user B calls "Aya-chan" from the alternative name list storage unit 112. In other words, the name identification unit 104 can acquire a larger list of alternative names than when not using conversations between user B and other users. The name identification unit 104 can inquire whether each of the alternative names obtained in this way (including names excluding honorifics) is registered in the name database storage unit 111. In other words, by using alternative names obtained from conversations between user B on a social networking service, the possibility of identifying user A's "A-chan" (the name to be identified) increases.

[0091] If the formal name registered in the name database storage unit 111 cannot be found even when using alternative names obtained from conversations on user B's social networking service, the same process may be performed by acquiring conversations of further users. This makes it possible to extract a larger number of alternative names related to the original name to be identified. In other words, the possibility of identifying the formal name of the original name to be identified is further increased.

[0092] FIG. 12 is a schematic diagram showing the expansion of the range in which the name identification device 1 acquires alternative names related to the name to be identified. In FIG. 12, the name identification device 1 extracts alternative names from within the range of conversations involving user A. That is, the name identification device 1 extracts alternative names from conversations between user A and user B, between user A and user C, between user A and user D, between user A and user E, and between user A and user F. In Modification 2, the name identification device 1 further expands the range to extract alternative names from conversations between user B and other users. That is, the name identification device 1 may extract alternative names from conversations between user B and user G, between user B and user H, and between user B and user I, in addition to the above. The name identification device 1 may further expand the range to extract alternative names from conversations between user H and other users. That is, the name identification device 1 may extract other names from the conversation between user H and user J, the conversation between user H and user K, and the conversation between user H and user L, in addition to the above.

[0093] That is, in Modification 2, the used name extraction unit 103 not only extracts alternative names from the text of the conversation history between a specific user (user A in the example of FIG. 12) who posted text including the name to be identified and the specific user and other users (users B, C, D, E, and F in the example of FIG. 12), but also expands the range of alternative names to be extracted. That is, in Modification 2, the used name extraction unit 103 further extracts alternative names from the text of the conversation history between the other user (e.g., user B in the example of FIG. 12) and yet another user (e.g., users G, H, and I in the example of FIG. 12). In the second modification, the used name extraction unit 103 further extracts alternative names from the text of the conversation history between the other user (e.g., user H in the example of FIG. 12) and yet another user (e.g., users J, K, and L in the example of FIG. 12) recursively, regarding the other user (e.g., user H in the example of FIG. 12) as the other user. Note that there is no limit to the range of recursion, and alternative names may be extracted from the text of the conversation history with users in a range beyond users J, K, and L in the example of FIG. 12.

[0094] In this way, the name identification device 1 can widen the range of conversations from which alternative names are extracted, thereby increasing the possibility that the name identification device 1 can identify the correct name from among the alternative names of the name to be identified.

[0095] Note that if the target name cannot be identified (does not match the official name registered in the name database storage unit 111) even after expanding the range of conversations as described above and acquiring conversations for all users whose conversation content can be acquired and extracting alternative names, the identification process is suspended. After that, after the passage of time and the amount of available conversation text increases, it is possible to arrive at the official name by performing the above identification process again.

[0096] [Variation 3] As a third modification, the alternative name list storage unit 112 may store, in association with each alternative name, the identification information (user ID) of the user who uses the alternative name. In this case, the name identification unit 104 can identify the name to be identified based on the identification information of the user associated with the alternative name by referring to the alternative name list storage unit 112. In this case, even when the same alternative name exists for multiple different objects (people, organizations, places, etc.), it becomes possible to easily identify the name to be identified based on the user's identification information.

[0097] [Variation 4] When the name identification device 1 acquires text of a post on a social networking service from the PDS device 2, the name identification device 1 may also acquire log data from services other than the social networking service from the PDS device 2. The name identification device 1 can use such log data from services other than the social networking service as a hint for name identification. As an example, a record of a user A who uses the name "A-chan" in a post on the social networking service searching for the term "Yasuda Aya" on another service (e.g., a search site) may remain in the log data. Alternatively, a record of user A purchasing a product containing the name "Yasuda Aya" on a shopping site may remain in the log data. In the case of Variation 4, the name identification device 1 can use these log data as reference information when identifying the name "A-chan" used by user A.

[0098] [Variation 5] The above description has been given using an example in which the name is that of a person (famous person). This embodiment is not limited to people, but can be similarly applied to any object (especially a famous object). The name that the name identification device 1 identifies may include, but is not limited to, a person's name, an organization name (group name), a place name, and a store name.

[0099] As an example of a store name, a supermarket in a certain region called "XX Mart" may be referred to by multiple users as "XX Mart," the old store name "XX store," or simply as "XX." As an example of a place name, the name "South Africa" may refer to the official name "Minami Alps City" (city name) or the official name "Republic of South Africa" (country name). In these cases, the name can be identified by applying this embodiment.

[0100] FIG. 13 is a block diagram showing an example of the internal configuration of each device constituting the information provision system in the name identification device 1 (including modified examples). The name identification device 1 can be realized using a computer. As shown in the figure, the computer includes a central processing unit 901, a RAM 902, an input / output port 903, input / output devices 904 and 905, etc., and a bus 906. The computer itself can be realized using existing technology. The central processing unit 901 executes instructions contained in a program read from the RAM 902, etc. In accordance with each instruction, the central processing unit 901 writes data to the RAM 902, reads data from the RAM 902, and performs arithmetic and logical operations. The RAM 902 stores data and programs. Each element contained in the RAM 902 has an address and can be accessed using the address. Note that RAM is an abbreviation for "random access memory." The input / output port 903 is a port through which the central processing unit 901 exchanges data with external input / output devices, etc. Input / output devices 904 and 905 exchange data with the central processing unit 901 via an input / output port 903. A bus 906 is a common communication path used within the computer. For example, the central processing unit 901 reads and writes data from and to RAM 902 via the bus 906. Also, for example, the central processing unit 901 accesses the input / output port 903 via the bus 906.

[0101] At least some of the functions of the name identification device 1 in the above-described embodiment can be realized by a computer and a program. In this case, the functions can be realized by recording a program for realizing the functions on a computer-readable recording medium and loading and executing the program recorded on the recording medium into a computer system. Note that the term "computer system" as used herein includes hardware such as an OS and peripheral devices. Furthermore, the term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, CD-ROMs, DVD-ROMs, and USB memory, as well as storage devices such as hard disks built into computer systems. In other words, a "computer-readable recording medium" may be a non-transitory computer-readable recording medium. Furthermore, the term "computer-readable recording medium" may also include media that temporarily and dynamically store programs, such as communication lines used when transmitting programs via networks such as the Internet or telephone lines, or media that store programs for a certain period of time, such as volatile memory within a computer system that serves as a server or client in such cases. The program may also be designed to realize some of the functions described above, or may be capable of realizing the functions described above in combination with a program already stored in the computer system.

[0102] According to the above-described embodiment (including the modified examples), when a user uses a specific name for a specific object in a post on a social networking service or the like, the name identification device 1 extracts alternative names from posts by other users related to the post. The name identification device 1 identifies the object of the original name by comparing this list of alternative names with a name database storage unit in which names including the official name are registered in advance. In other words, according to this embodiment, even when a name unique to the user that is not commonly used is used, or when the text of the post is extremely short and there are almost no other words before or after the name that can serve as a clue, it is possible to identify which object the user is referring to in the post.

[0103] Although an embodiment of the present invention has been described above in detail with reference to the drawings, the specific configuration is not limited to this embodiment, and includes designs within the scope of the gist of the present invention. [Industrial Applicability]

[0104] The present invention can be used, for example, to analyze and utilize text posted on social networking services, etc. However, the scope of use of the present invention is not limited to the examples given here. [Explanation of symbols]

[0105] 1. Name Identification Device 2 PDS device 3 Server equipment 8. Communication Networks 10 Systems 101 Behavioral history acquisition unit (history acquisition unit) 102 Name Extraction Unit 103 Usage name extraction unit 104 Name Identification Section 111 Name database storage unit 112 Alternative name list storage unit 901 Central Processing Unit 902 RAM 903 Input / Output Ports 904,905 Input / Output Devices 906 Bus

Claims

1. a name database storage unit that stores in advance authorized names that are names authorized for the object; an alternative name list storage unit that stores a list of alternative names related to an identification target name that is an object of identification; a history acquisition unit that acquires text of a conversation history between users; a used name extraction unit that extracts alternative names related to the name to be identified from the text acquired by the history acquisition unit and stores the extracted alternative names in the alternative name list storage unit; a name identification unit that identifies the certified name corresponding to the name to be identified by comparing a list of alternative names stored in the alternative name list storage unit with the certified name stored in the name database storage unit; A name identification device comprising:

2. The used name extraction unit extracts the alternative name from the text of the conversation history between a specific user who posted a text including the identification target name, and the specific user and another user. The name identification device according to claim 1 .

3. the used name extraction unit further extracts the alternative names from the text of the conversation history between the other user and another user, and from the text of the conversation history between the other user and another user recursively, with the other user being the other user; The name identification device according to claim 2.

4. the alternative name list storage unit stores frequency information regarding the frequency of appearance of each alternative name; the name identification unit identifies the certified name corresponding to the name to be identified by comparing only the alternative names whose frequency information satisfies a predetermined condition with the certified name; The name identification device according to claim 1 .

5. When the name identification unit is able to identify the certified name corresponding to the identification target name, it stores the identification target name in the name database storage unit as a certified alternative name which is an alternative name certified in relation to the certified name. The name identification device according to claim 1 .

6. the history acquisition unit acquires text of a conversation history between users based on a chain of posts on the social networking service; The name identification device according to claim 1 .

7. a name database storage unit that stores in advance authorized names that are names authorized for the object; an alternative name list storage unit that stores a list of alternative names related to an identification target name that is an object of identification; a history acquisition unit that acquires text of a conversation history between users; a used name extraction unit that extracts alternative names related to the name to be identified from the text acquired by the history acquisition unit and stores the extracted alternative names in the alternative name list storage unit; a name identification unit that identifies the certified name corresponding to the name to be identified by comparing a list of alternative names stored in the alternative name list storage unit with the certified name stored in the name database storage unit; A program for causing a computer to function as a name identification device comprising:

Citation Information

Patent Citations

  • Anthroponym expression identification device, its method, program, and recording medium

    JP2009181183A