Information processing device, information processing method, and program

The information processing device enhances the estimation of user activity areas and relationships by utilizing social media data, providing accurate insights into user activities and area significance.

JP7740364B2Active Publication Date: 2025-09-17NEC CORP
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2023563438
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-11-26
Publication Date
2025-09-17
Estimated Expiration
2041-11-26

AI Technical Summary

Technical Problem

Existing techniques for estimating user activity ranges and relationships based on social media information are not fully effective in utilizing the available data for generating useful insights.

Method used

An information processing device and method that estimates a user's activity area and relationship with that area using public information from social media accounts, including profile data, posts, and relationships with other users, employing various methods to enhance accuracy and relevance.

Benefits of technology

Generates useful information about user activity areas and relationships, including estimating time periods of activity and significance of areas, even when hometown information is absent or incorrect, by leveraging social media data effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007740364000007
    Figure 0007740364000007
  • Figure 0007740364000008
    Figure 0007740364000008
  • Figure 0007740364000009
    Figure 0007740364000009
Patent Text Reader

Abstract

The present invention provides an information processing device (1000) comprising: an activity area estimation unit (1001) that estimates, on the basis of public information which is published on the Internet and which is linked to a social media account, the activity area of a user of the account; and a relationship estimation unit (1002) that estimates, on the basis of public information, the relationship between the user of the account and the activity area.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing method, and a program. [Background technology]

[0002] Techniques related to the present invention are disclosed in Patent Documents 1 to 4 and Non-Patent Documents 1 to 7.

[0003] Patent Document 1, Non-Patent Documents 1 to 4, and Non-Patent Document 7 disclose techniques for estimating the range of activities of a user who has an account on a social media such as an SNS (social networking service) based on friendship relationships.

[0004] Patent Documents 2 to 4, Non-Patent Document 5 and Non-Patent Document 6 disclose techniques for identifying social media accounts such as SNSs owned by the same person.

[0005] Patent Document 5 discloses a technique for identifying the location of a user when text data or the like is posted on a social media such as an SNS. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] International Publication No. 2021 / 028988 [Patent Document 2] International Publication No. 2019 / 187107 [Patent Document 3] International Publication No. 2019 / 234827 [Patent Document 4] Japanese Patent Application Laid-Open No. 2013-122630 [Patent Document 5] Japanese Patent Application Publication No. 2018-010378 [Non-patent literature]

[0007] [Non-Patent Document 1] Keisuke Ikeda, Kazufumi Kojima, Masahiro Tani, "A Study on Residential Area Estimation Methods Focusing on Geographical Proximity of Friends," IEICE Technical Report, Vol. 119, No. 317, pp. 37-42, AI2019-36, November 2019. [Non-patent document 2] Dan Xu, Peng Cui, Wenwu Zhu, Shiqiang Yang, "Graph-based residence location inference for social media users", IEEE Computer Society, IEEE MultiMedia, Volume 21, Issue 4, pp 76-83, October 2014 [Non-patent document 3] Backstrom Lars, Eric Sun, Cameron Marlow, "Find me if you can: Improving geographical prediction with social and spatial proximity" Proceedings of the 19th international conference on World Wide Web, 2010, pp.61-70 [Non-patent document 4] Liu Zhi, Yan Huang, "Closeness and structure of friends help to estimate user locations", International Conference on Database Systems for Advanced Applications, Springer, pp. 33-48 [Non-Patent Document 5] Y. Li, Y. Peng, W. Ji, Z. Zhang, and Q. Xu, "User Identification Based on Display Names Across Online Social Networks," IEEE Access, vol. 5, pp. 17342-17353, August 25, 2017. [Non-patent document 6] X. Han, X Liang and et al. "Linking social network accounts by modeling user spatiotemporal habits", Intelligence and Security Informatics (ISI), IEEE International Conference on, 2017 [Non-Patent Document 7] Keisuke Ikeda, Kazufumi Kojima, Masahiro Tani, "A Method for Estimating Social Media User Activity Areas Using Kernel Density Estimation," IEICE Technical Report, Vol. 120, No. 379, pp. 18-23, AI2020-42, February 2021 Summary of the Invention [Problem to be solved by the invention]

[0008] As described above, the range of a user's activities can be estimated based on information published on social media such as SNS. There is a demand for more effective use of information published on social media such as SNS.

[0009] An object of the present invention is to generate useful information based on information published on social media such as SNS. [Means for solving the problem]

[0010] According to the present invention, an activity area estimation means for estimating an activity area of ​​a user of a social media account based on public information linked to the account and published on the Internet; A relationship estimation means for estimating a relationship between a user of the account and the activity area based on the public information; An information processing device having the above configuration is provided.

[0011] Further, according to the present invention, The computer an activity area estimation step of estimating an activity area of ​​a user of a social media account based on public information linked to the account and published on the Internet; a relationship estimation step of estimating a relationship between the user of the account and the activity area based on the public information; An information processing method is provided that performs the above.

[0012] Further, according to the present invention, Computer, an activity area estimation means for estimating the activity area of ​​a user of a social media account based on public information linked to the account and published on the Internet; a relationship estimation means for estimating a relationship between the user of the account and the activity area based on the public information; A program is provided to function as a [Effects of the Invention]

[0013] According to the present invention, useful information can be generated based on information published on social media such as SNS. [Brief explanation of the drawings]

[0014] [Figure 1] FIG. 1 is a diagram illustrating an example of a hardware configuration of an information processing apparatus according to an embodiment of the present invention. [Figure 2] FIG. 2 is a functional block diagram illustrating an example of an information processing apparatus according to the present embodiment. [Figure 3]10 is a flowchart illustrating an example of processing operation of the information processing apparatus according to the present embodiment. [Figure 4] 1 is a configuration diagram illustrating an outline of an estimation device according to an embodiment of the present invention. [Figure 5] 1 is a configuration diagram showing an example of the configuration of an activity area estimation system according to an embodiment of the present invention; [Figure 6] 10 is a flowchart illustrating an example of the operation of the activity area estimation device of the present embodiment. [Figure 7] FIG. 10 is a diagram illustrating an example of generating a posting distribution according to the present embodiment. [Figure 8] FIG. 10 is a diagram illustrating an example of generating a post distribution according to the present embodiment. [Figure 9] FIG. 10 is a diagram illustrating an example of generating a friend distribution according to the present embodiment. [Figure 10] FIG. 10 is a diagram illustrating an example of generating an activity area distribution according to the present embodiment. [Figure 11] FIG. 10 is a diagram showing an example of an output of an active area distribution according to the present embodiment. [Figure 12] 1 is a configuration diagram showing an example of the configuration of an activity area estimation device according to an embodiment of the present invention; [Figure 13] 10 is a flowchart illustrating an example of the operation of the activity area estimation device of the present embodiment. [Figure 14] 1 is a configuration diagram showing an example of the configuration of an activity area estimation device according to an embodiment of the present invention; [Figure 15] 10 is a flowchart illustrating an example of the operation of the activity area estimation device of the present embodiment. [Figure 16] 1 is a configuration diagram showing an example of the configuration of an activity area estimation device according to an embodiment of the present invention; [Figure 17] 10 is a flowchart illustrating an example of the operation of the activity area estimation device of the present embodiment. [Figure 18] 10 is a flowchart illustrating an example of processing operation of the information processing apparatus according to the present embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0015] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings, like components are designated by like reference numerals, and the description thereof will be omitted as appropriate.

[0016] First Embodiment "overview" The information processing device of this embodiment estimates the real-world activity area of ​​the account user based on public information linked to the account on social media such as an SNS and published on the Internet. Furthermore, the information processing device estimates the relationship between the account user and the estimated activity area based on the public information. Thus, the information processing device of this embodiment can estimate not only the activity area of ​​the account user but also the relationship between the activity area and the account user based on the public information on social media.

[0017] "Hardware Configuration" Next, an example of the hardware configuration of an information processing device will be described. Each functional unit of the information processing device is realized by any combination of hardware and software, centered on a CPU (Central Processing Unit) of any computer, memory, programs loaded into the memory, a storage unit such as a hard disk that stores the programs (this can store programs that are pre-loaded when the device is shipped, as well as programs downloaded from storage media such as CDs (Compact Discs) or servers on the Internet), and a network connection interface. Those skilled in the art will understand that there are many variations in the realization methods and devices.

[0018] FIG. 1 is a block diagram illustrating an example of the hardware configuration of an information processing device. As shown in FIG. 1, the information processing device has a processor 1A, a memory 2A, an input / output interface 3A, a peripheral circuit 4A, and a bus 5A. The peripheral circuit 4A includes various modules. The information processing device does not necessarily have to have the peripheral circuit 4A. Note that the information processing device may be composed of multiple devices that are physically and / or logically separated. In this case, each of the multiple devices can have the above hardware configuration.

[0019] The bus 5A is a data transmission path for the processor 1A, memory 2A, peripheral circuit 4A, and input / output interface 3A to transmit and receive data among them. The processor 1A is an arithmetic processing device such as a CPU or a GPU (Graphics Processing Unit). The memory 2A is a memory such as a RAM (Random Access Memory) or a ROM (Read Only Memory). The input / output interface 3A includes an interface for acquiring information from an input device, an external device, an external server, an external sensor, a camera, etc., and an interface for outputting information to an output device, an external device, an external server, etc. Examples of input devices include a keyboard, a mouse, a microphone, physical buttons, a touch panel, etc. Examples of output devices include a display, a speaker, a printer, a mailer, etc. The processor 1A can issue commands to each module and perform calculations based on the results of those calculations.

[0020] "Function Configuration" Next, a functional configuration of the information processing device will be described. An example of a functional block diagram of the information processing device 1000 is shown in Fig. 2. As shown in the figure, the information processing device 1000 has an activity area estimation unit 1001 and a relationship estimation unit 1002.

[0021] The activity area estimation unit 1001 estimates the activity area of ​​the user of the account based on public information that is linked to the account of a social media such as an SNS and that is made public on the Internet.

[0022] "Public information" can include any information that is linked to the user of each account and that is made public on social media. For example, public information can include at least one of the following: the profile of the user of each account, posts posted by the user of each account, relationship information indicating relationships with users of other accounts on social media, profiles of users of other accounts that have a predetermined relationship with the user of each account on social media, and posts posted by users of the other accounts.

[0023] The items included in a "profile" may differ depending on the social media, but may include, for example, username, nickname, gender, date of birth, nationality, age (or generation), birthplace, current place of residence, affiliation (company name, school name), alma mater, etc.

[0024] "Posts" include messages, still images, moving images, and audio.

[0025] The "relationship information" is information indicating connections with users of other accounts on social media. For example, the relationship information may indicate at least one of users of other accounts who have a mutual follow relationship with the user of each account, users of other accounts who are followed by the user of each account, users of other accounts who follow the user of each account, users of other accounts who have exchanged messages with the user of each account, and users of other accounts who have been in the same place at the same time as the user of each account.

[0026] "Having a history of exchanging messages" may mean that at least one user has sent text data, emoticons, photos, videos, audio, icons, etc. to the other user, or has taken an action by pressing the like button. Alternatively, "having a history of exchanging messages" may mean that both users have sent each other text data, emoticons, photos, videos, audio, icons, etc., or have taken an action by pressing the like button.

[0027] "Users of other accounts who have been in the same place at the same time as a user of each account" may be identified, for example, based on the posting location and posting date and time. If the difference in posting date and time between posts by a user of one account and a user of another account is within a reference value and the posting locations are the same or the difference is within a reference value, the two users may be determined to have been in the same place at the same time. Additionally, if the locations of users of each account are tracked using a Global Positioning System (GPS), two users who have been within a threshold distance from each other, or two users who have been within a threshold distance from each other for a predetermined period of time or longer, may be determined to have been in the same place at the same time. Additionally, if the facilities (stores, etc.) used by each user and the dates and times of use can be acquired, two users who used the same facility and whose difference in dates and times of use is within a reference value may be determined to have been in the same place at the same time. The facilities (stores, etc.) used by each user and the dates and times of use may be identified based on each user's posts or by other methods.

[0028] "Users of other accounts who have a specified relationship with the user of each account" refers to, for example, at least one of users of other accounts who have a mutual follow relationship with the user of each account, users of other accounts who are followed by the user of each account, users of other accounts who follow the user of each account, users of other accounts who have exchanged messages with the user of each account, and users of other accounts who have been in the same place at the same time as the user of each account.

[0029] An "activity area" is the area in the real world where the account user is active, and may be a city, ward, town, village, or a larger or smaller area.

[0030] The activity area estimation unit 1001 acquires the public information from a server that provides social media services. Then, the activity area estimation unit 1001 estimates the activity area of ​​the user of each account based on the acquired public information. There are no particular limitations on the method for estimating the activity area of ​​the user of each account, and any technique can be adopted, such as the techniques disclosed in Patent Document 1, Non-Patent Documents 1 to 4, and Non-Patent Document 7. In addition, the following embodiment will describe another example of the method for estimating the activity area of ​​the user of each account.

[0031] The activity area estimation unit 1001 may identify multiple accounts owned by the same user. When estimating the activity area of ​​a user of a certain account, not only public information linked to the account but also public information linked to other accounts of the user may be used. Using more public information improves the accuracy of estimating the activity area. There are no particular limitations on how multiple accounts owned by the same user may be identified, and any technique, such as those disclosed in Patent Documents 2 to 4, Non-Patent Document 5, and Non-Patent Document 6, may be employed.

[0032] The relationship estimation unit 1002 estimates the relationship between the user of the account and the estimated activity area of ​​the user of the account (hereinafter, sometimes referred to as "user-activity area relationship") based on public information.

[0033] The "user-activity area relationship" can be estimated from public information and indicates the relationship between the user of the account and the activity area. The user-activity area relationship may indicate, for example, a temporal relationship (i.e., what period of time the relationship exists), or may indicate the significance of the activity area to the user of the account (e.g., hometown, current residence, etc.), or may indicate other content. In the following embodiment, a specific example of the user-activity area relationship will be described.

[0034] In addition, the activity area of ​​an account user may be divided into multiple child areas based on geographical relationships. For example, if the activity area includes multiple areas that can be distinguished from each other, such as City A and City B, the activity area may be divided into multiple child areas based on the relationships. Also, if the activity area includes multiple areas that are separated from each other, the combined areas may be treated as a single child area.

[0035] In this way, when the activity area of ​​the user of the account can be divided into multiple child areas, the relationship estimation unit 1002 can estimate the relationship with the user of the account (user-activity area relationship) for each child area.

[0036] Next, an example of the processing flow of the information processing device 1000 will be described with reference to the flowchart of FIG.

[0037] First, the information processing device 1000 estimates the activity area of ​​the user of the account based on public information linked to the social media account and published on the Internet (S10). Then, the information processing device 1000 estimates the relationship between the user of the account and the activity area estimated in S10 (user-activity area relationship) based on the public information (S11).

[0038] The information processing device 1000 may associate the estimated activity area and the user-activity area relationship with the user of the account and register them in a storage device. The information processing device 1000 may also output information in which the estimated activity area and the user-activity area relationship are associated with the user of the account via an output device. Examples of the output device include, but are not limited to, a display, a projection device, a printer, a mailer, etc.

[0039] "Action and effect" According to the information processing device 1000 of this embodiment, not only the activity area of ​​the account user but also the relationship between the activity area and the account user (user-activity area relationship) can be estimated based on public information on social media. According to the information processing device 1000, such useful information can be generated based on public information.

[0040] <Second embodiment> The information processing device 1000 of this embodiment estimates the period when the user of the account was active in the activity area (user-activity area relationship), which will be described in detail below.

[0041] The relationship estimation unit 1002 estimates the posting location of a post, i.e., the location of the account user at the time the post was posted. Then, based on the posting date and time of a post whose posting location is included in the activity area, the relationship estimation unit 1002 estimates the period when the account user was active in the activity area.

[0042] The posting location can be estimated using any technology. For example, if a post is accompanied by metadata (such as a geotag) indicating the posting location, the relationship estimation unit 1002 can estimate the location indicated by the metadata as the posting location. Alternatively, the relationship estimation unit 1002 may identify the shooting location based on landmarks or the like appearing in the posted image (including still images and moving images) and estimate the identified shooting location as the posting location. Alternatively, the relationship estimation unit 1002 may identify the location where the audio was recorded based on background sounds in the posted audio data and estimate the identified location as the posting location. Alternatively, the relationship estimation unit 1002 may estimate the posting location based on the location indicated by keywords included in the posted text data, audio data, or image. Examples of keywords include, but are not limited to, place names and landmark names. Alternatively, the relationship estimation unit 1002 may use other technologies such as those disclosed in Patent Document 5.

[0043] Next, a process for estimating the period when the account user was active in an activity area based on the posting date and time of a post whose posting location is included in the activity area will be described.

[0044] First, the relationship estimation unit 1002 identifies posts whose posting locations are included in the activity area. Then, based on the posting dates and times of the identified posts, the relationship estimation unit 1002 estimates the period when the account user was active in the activity area.

[0045] For example, the relationship estimation unit 1002 may estimate the period from the posting date and time of the first posted post among the identified posts to the posting date and time of the last posted post as the period during which the account user was active in that activity area.

[0046] Alternatively, the relationship estimation unit 1002 may first remove outlier data (posts) whose posting dates and times deviate significantly from those of other posts from among posts whose posting locations are included in the activity area. After removing the outlier data from posts whose posting locations are included in the activity area, the relationship estimation unit 1002 may estimate the period from the posting date and time of the first posted post to the posting date and time of the last posted post among the remaining posts as the period during which the account user was active in that activity area. Detection of outlier data can be achieved using any conventional technology.

[0047] Alternatively, the relationship estimation unit 1002 may estimate a period obtained by extending the period estimated as described above by a predetermined length before and after the period as the period when the account user was active in the activity area. Examples of the predetermined length include, but are not limited to, "X days," "Y months," "Z% of the length of the estimated period," etc.

[0048] Other configurations of the information processing device 1000 of this embodiment are the same as those of the first embodiment.

[0049] The information processing device 1000 of this embodiment achieves the same effects as those of the first embodiment. Furthermore, the information processing device 1000 of this embodiment can estimate not only the activity area of ​​the account user but also the time period when the account user was active in the activity area based on public information on social media. The information processing device 1000 can generate such useful information based on public information.

[0050] <Third embodiment> The information processing device 1000 of this embodiment estimates what meaning the activity area has for the user of the account (user-activity area relationship), which will be described in detail below.

[0051] Incidentally, the language used by the account user can be classified according to language type (Japanese, English, etc.). Furthermore, one language type can be classified into multiple types according to dialect. The information processing device 1000 estimates what meaning the activity area has for the account user based on the relationship between the language used by the account user and the language commonly used in that activity area.

[0052] Specifically, the relationship estimation unit 1002 compares the language characteristics of the account user with the language characteristics of the language commonly used in the activity area. Based on the comparison result, the relationship estimation unit 1002 estimates whether the activity area is the hometown of the account user. If the language characteristics of the account user match the language characteristics of the language commonly used in the activity area, the relationship estimation unit 1002 estimates that the activity area is the hometown of the account user. On the other hand, if the language characteristics of the account user do not match the language characteristics of the language commonly used in the activity area, the relationship estimation unit 1002 estimates that the activity area is not the hometown of the account user. The information processing device 1000 stores information indicating the language characteristics commonly used in each region in advance and can perform the above processing using the information. The language used by the account user is the language used in their profile and posts.

[0053] In addition, if the user of an account uses multiple languages ​​(Japanese, English, etc.), the relationship estimation unit 1002 may determine one of them as the language used by the user of that account based on the frequency of use, language proficiency, etc., and make the above estimation based on the results of comparing the characteristics of the determined language with the characteristics of the language used in the activity area.

[0054] For example, the relationship estimation unit 1002 may determine the language with the highest usage frequency as the language used by the user of the account. Alternatively, the relationship estimation unit 1002 may determine the language with the highest language proficiency as the language used by the user of the account. Language proficiency can be evaluated using any technology. For example, it can be evaluated based on various items such as the degree of grammatical errors, the degree of spelling errors and omissions, and the difficulty of the words used. The fewer grammatical errors, the fewer spelling errors and omissions, and the more difficult words used, the higher the proficiency.

[0055] Other configurations of the information processing device 1000 of this embodiment are the same as those of the first and second embodiments.

[0056] The information processing device 1000 of this embodiment achieves the same effects as those of the first and second embodiments. Furthermore, the information processing device 1000 of this embodiment can estimate not only the activity areas of the account users but also whether each activity area is the user's hometown based on public information on social media. The information processing device 1000 can generate such useful information based on public information.

[0057] The information processing device 1000 of this embodiment is useful when the hometown is not included in the profile (public information) of the user of the account. Even if the hometown is included in the profile (public information) of the user of the account, the user of the account may have registered a false hometown. Considering this, the information processing device 1000 of this embodiment is useful even when the hometown is included in the profile (public information) of the user of the account.

[0058] <Fourth embodiment> The information processing device 1000 of this embodiment estimates what meaning the activity area has for the user of the account (user-activity area relationship) based on public information of users of other accounts that have a predetermined relationship with the user of the account. This will be described in detail below.

[0059] The relationship estimation unit 1002 estimates the relationship between the user of the account and the activity area (user-activity area relationship) based on public information published on the Internet linked to users of other accounts that have a predetermined relationship with the user of the account.

[0060] Specifically, the relationship estimation unit 1002 estimates that an activity area that matches the hometown of a user of another account that has a predetermined relationship with the user of the account is the hometown of the user of the account. In this case, the relationship estimation unit 1002 may estimate that an activity area that matches the hometown of a user who satisfies a predetermined condition among users of other accounts that have a predetermined relationship with the user of the account is the hometown of the user of the account.

[0061] The predetermined condition is “friends since childhood, elementary school, junior high school, or high school.” Whether a user of another account that has a predetermined relationship with the user of the account satisfies the predetermined condition can be determined based on publicly available information.

[0062] For example, the determination may be based on the timing of establishing a predetermined relationship on social media (mutual following, following, history of exchanging messages, etc.). A user of another account who established the predetermined relationship with the user of the account when the user of the account was a child, elementary school student, junior high school student, or high school student is determined to be a "friend from childhood, elementary school, junior high school, or high school."

[0063] Additionally, if a user of an account refers to a user of another account as a "childhood friend," "a friend since elementary school," "a friend since middle school," or "a friend since high school" in public information, or conversely, if a user of another account refers to a user of the account in such ways in public information, the user of the other account may be determined to be a "friend since childhood, elementary school, middle school, or high school." Note that the names exemplified here are merely examples. By predefining all possible names that indicate "friends since childhood, elementary school, middle school, or high school" and detecting when a user is referred to in such a way, it is possible to reduce missed detections.

[0064] Here, a method for identifying the hometown of a user of another account will be described. For example, the hometown included in the profile (public information) of the user of another account may be identified as the hometown of the user of the other account. Alternatively, the hometown of the user of the other account may be estimated using the method described in the third embodiment.

[0065] Other configurations of the information processing apparatus 1000 of this embodiment are the same as those of the first to third embodiments.

[0066] The information processing device 1000 of this embodiment achieves the same effects as those of the first to third embodiments. Furthermore, the information processing device 1000 of this embodiment can estimate not only the activity area of ​​the account user but also whether each activity area is the user's hometown based on public information on social media. The information processing device 1000 can generate such useful information based on public information.

[0067] <Fifth embodiment> The information processing device 1000 of this embodiment estimates what the activity area means to the user of the account (user-activity area relationship) based on public information of users of other accounts that have a predetermined relationship with the user of the account, using a method different from that of the fourth embodiment. This will be described in detail below.

[0068] Generally, during childhood, elementary school, junior high school, and high school, people's circle of friends and areas of activity are smaller than those of university students and working adults, and they tend to become friends with people they meet relatively close to them, such as at school or in their neighborhood. For this reason, the hobbies and preferences of the multiple friends they have during childhood, elementary school, junior high school, and high school tend to be quite diverse.

[0069] On the other hand, when people enter university or work, their circle of friends and areas of activity become wider than when they were children, elementary school students, junior high school students, or high school students, and they tend to become friends with people who share common interests, such as hobbies and tastes. For this reason, the hobbies and tastes of multiple friends during university or work life tend to be similar to each other.

[0070] Therefore, the information processing device 1000 of this embodiment estimates the relationship between each activity area and the user of the account (user-activity area relationship) based on the degree of variation in the hobbies and preferences of friends in each activity area.

[0071] The relationship estimation unit 1002 executes steps 1 to 3 shown in the flowchart of FIG.

[0072] In step 1, the relationship estimation unit 1002 identifies users of other accounts related to the activity area of ​​the user of the account from users of other accounts that have a predetermined relationship with the user of the account (S20). If the activity area of ​​the user of the account can be divided into multiple child areas, the relationship estimation unit 1002 identifies users of other accounts related to each child area for each child area.

[0073] "Users of other accounts related to the activity area of ​​the account user" may include, for example, at least one of "users whose own activity area is included in the activity area of ​​the account user," "users whose posting locations in any of their posts are included in the activity area of ​​the account user," "users whose posting locations in a certain percentage or more of their posts are included in the activity area of ​​the account user," "users whose hometown or current place of residence included in their profile is included in the activity area of ​​the account user," and "users whose affiliation or alma mater location included in their profile is included in the activity area of ​​the account user."

[0074] Step 2 is performed after identifying users of other accounts related to the activity area of ​​the user of the account. In step 2, the relationship estimation unit 1002 calculates the degree of variation in the hobbies and preferences of the identified users of the other multiple accounts (S21).

[0075] The hobbies and preferences of users of other accounts can be estimated based on public information (profiles, posts, etc.) linked to and published by users of other accounts. For example, the hobbies and preferences may be estimated based on the frequency of appearance of words related to each of a plurality of hobbies and preferences (baseball, soccer, music, piano, ice cream, donuts, bread, etc.) in the public information, or may be estimated using other methods.

[0076] The degree of variation in tastes and preferences can be expressed by, for example, information entropy, but is not limited to this.

[0077] Step 3 is performed after calculating the degree of variation in the hobbies and preferences of the users of the identified multiple other accounts. In step 3, the relationship estimation unit 1002 estimates the activity area of ​​the user of the account and the relationship with the user of the account (user-activity area relationship) based on the calculated degree of variation in the hobbies and preferences (S22).

[0078] If the hobbies and preferences of users of multiple other accounts related to an activity area vary by more than a reference level, the relationship estimation unit 1002 estimates that the activity area is an area where the user of that account was active as a child, elementary school student, junior high school student, or high school student, that is, the hometown of the user of that account.

[0079] Furthermore, if the hobbies and preferences of users of multiple other accounts related to the activity area do not vary by more than a reference level, the relationship estimation unit 1002 estimates that the activity area is not an area in which the user of that account was active when they were a university student or working adult, i.e., it is not the hometown of the user of that account.

[0080] Other configurations of the information processing device 1000 of this embodiment are the same as those of the first to fourth embodiments.

[0081] The information processing device 1000 of this embodiment achieves the same effects as those of the first to fourth embodiments. Furthermore, the information processing device 1000 of this embodiment can estimate not only the activity areas of the account users but also whether each activity area is the user's hometown based on public information on social media. The information processing device 1000 can generate such useful information based on public information.

[0082] Sixth Embodiment In this embodiment, a method for estimating an activity area of ​​a user of an account is embodied. In this embodiment, an activity area estimation unit 1001 is realized by an estimation device 10 described below. Other configurations of the information processing device 1000 of this embodiment are the same as those of the first to fifth embodiments.

[0083] FIG. 4 shows an overview of the estimation device 10. The estimation device 10 is a device that estimates activity locations of a target user in a physical space (sometimes called the "real world" or "actual space") using information from social media. As shown in FIG. 4, the estimation device 10 includes a first location distribution generation unit 11, a second location distribution generation unit 12, and an estimation unit 13. The first location distribution generation unit 11 generates a first location distribution of the target user based on account information of the target user on social media. For example, the first location distribution generation unit 11 may generate a posting distribution based on posted information (posting locations) of the target user. "Posted information" is synonymous with "posted material" in the first to fifth embodiments.

[0084] The second location distribution generation unit 12 generates a second location distribution of friends based on account information of friends who are related to the target user on social media. For example, the second location distribution generation unit 12 may generate a friend distribution based on activity base information (residential information) of friends. "Friends who are related to the target user on social media" is synonymous with "users of other accounts who have a predetermined relationship with users of each account" in the first to fifth embodiments.

[0085] The estimation unit 13 estimates the activity location of the target user based on the generated first location distribution and the generated second location distribution. For example, the estimation unit 13 may estimate the activity location of the target user based on the overlap between the first location distribution and the second location distribution. Alternatively, the first location distribution and the second location distribution may be generated using a nonparametric method such as a kernel density estimation function, and the activity location may be estimated. Either the first location distribution or the second location distribution may be generated using a nonparametric method. The estimated activity location may be an activity area, or may be an everyday activity location that the target user visits in their daily life (such as a residence or workplace, a store visited for shopping or eating, or a travel route between these locations), or may be an unusual activity location that the target user does not visit in their daily life (such as a tourist spot or hotel during a trip or business trip, or a travel route).

[0086] As described above, in this embodiment, by using the location distribution based on the target user's account information and the location distribution based on the friend's account information, the target user's activity location (activity area) can be estimated with less information. For example, the activity area may be estimated when only the target user's posted information or the friend's friend information is available. When two types of information are available, the activity area can be estimated more accurately by combining them. Furthermore, by using a nonparametric method that does not require large-scale data collection, the cost of collecting social data, which has limitations on data collection, can be reduced. Furthermore, the information processing device 1000 of this embodiment achieves the same effects as those of the first to fifth embodiments.

[0087] Hereinafter, embodiments 6-1 to 6-4 that are more specific versions of the sixth embodiment will be described.

[0088] (Embodiment 6-1) Hereinafter, a 6-1 embodiment will be described with reference to the drawings. Fig. 5 shows an example of the configuration of an activity area estimation system 1 according to this embodiment. As shown in Fig. 5, the activity area estimation system 1 according to this embodiment includes an activity area estimation device 100 (one embodiment of the estimation device 10) and a social media system 200.

[0089] The social media system 200 is a system that provides social media services such as SNS. The social media system 200 may include multiple social media services. A social media service is an online service that enables multiple accounts (users) to send (publish) information and communicate with each other over the Internet (online). Social media services are not limited to SNS, but also include messaging services such as chat, blogs, electronic bulletin boards (forum sites), video sharing sites, information sharing sites, social games, social bookmarks, etc.

[0090] For example, the social media system 200 includes a server on the cloud and a user terminal. The server may be a social media server or a web server. The user terminal logs in with a user account via an API (Application Programming Interface) provided by the server, inputs and views posts, and registers account connections such as friend relationships and following relationships. The social media system 200 and the activity area estimation device 100 are connected to each other so as to be able to communicate with each other via the Internet or the like.

[0091] The activity area estimation device 100 includes a post information acquisition unit 101, a post distribution generation unit 102, a friend information acquisition unit 103, a friend distribution generation unit 104, an activity area estimation unit 105, and an activity area output unit 106. Note that the configuration of each unit (block) is an example, and the device may be configured with other units as long as the operation (method) described below is possible. Furthermore, each unit may be provided in one device or in multiple devices. For example, the post information acquisition unit 101 and the post distribution generation unit 102 may be a first location distribution generation unit, and the friend information acquisition unit 103 and the friend distribution generation unit 104 may be a second location distribution generation unit.

[0092] The posted information acquisition unit (target account information acquisition unit) 101 acquires posted information of a target account from the social media system 200. The posted information acquisition unit 101 also functions as a target account identification unit that identifies a target account of a target user whose activity area is to be estimated. The posted information acquisition unit 101 acquires account information (social media information) of the identified target account from the social media system 200. The account information is synonymous with the "public information" in the first to fifth embodiments and includes account profile information, posted information, and the like. The posted information acquisition unit 101 may acquire account information for multiple social media platforms. The posted information acquisition unit 101 may acquire the posted information from a server that provides a social media service via an API or a crawler (acquisition tool), or may acquire the information from a database in which social media account information is stored in advance.

[0093] The posted information acquisition unit 101 acquires all posted information (synonymous with posts) from the account information of the target account. The posted information includes images, text, etc. posted by the account (user) to a timeline, etc. The posted information acquisition unit 101 extracts the posting location and posting date and time from the acquired images and text of the posted information. The posting location is the location where the target user posted the posted information, and the posting date and time is the date and time when the posted information was posted. The posting date and time are registered in association with the posted image or text at the time of posting. The posted location is location information that can be extracted from the posted information, and may be a geotag such as GPS (Global Positioning System) information assigned to the posted image, or a location identified from the appearance of a landmark or the like in the posted image. Furthermore, the posted location is not limited to an image, and may also be a location mentioned in the posted text (text). The location mentioned in the posted text is extracted, for example, by natural language processing of the posted text. The posting location is an example of location information for estimating the target user's activity location (place of affiliation) from the target user's account information, and may not be limited to the posting location but may also be an activity base such as a residence included in the profile information.

[0094] The post distribution generation unit 102 generates a post distribution (first location distribution) of the target account based on the post information of the target account. The post distribution generation unit 102 generates a post distribution of the post locations of the extracted target account. The post distribution is a distribution of post locations (post locations) in physical space (a spatial distribution specific to the post locations), and is, for example, a two-dimensional geographical spatial distribution consisting of latitude and longitude coordinates. For example, the post distribution is a distribution of post locations in distribution area units of a predetermined size. The granularity level of the distribution area may be an administrative district unit such as a country, prefecture, or city, town, or village, or may be a mesh unit of a predetermined size such as 1 km x 1 km, 100 m x 100 m, or 10 m x 10 m.

[0095] The post distribution generation unit 102 calculates the post distribution using a predetermined distribution function. It is preferable to use a density estimation function that estimates the distribution using a non-parametric method. In this embodiment, a kernel density estimation function is used as an example of a density estimation function using a non-parametric method. When generating (calculating) the post distribution, weighting may be applied to each piece of post information based on the post information. For example, weighting may be applied to the post information based on the posting date and time. Note that the post distribution may be calculated using other statistical processing, not just a distribution function. For example, the post distribution (histogram) may be generated by counting the number of posting locations included in each distribution area.

[0096] The friend information acquisition unit 103 acquires friend information of the friend account from the social media system 200. The friend information acquisition unit 103 also functions as a friend account identification unit that identifies the friend account of the target user. A friend account is an account that has a connection, such as a friendship relationship, with the target account in social media. The friend account may be an account on the same social media as the target user, or on a different social media. For example, a friend account is an account that has a friendship relationship registered with the target account, but may also be an account (associated account) that has another connection (relationship) with the target account. The associated account is synonymous with "other accounts that have a predetermined relationship with the user of each account" in the first to fifth embodiments. The associated account may be, for example, an account that has a follow relationship (following or following), a connection through posts (comments on posts, quotes such as retweets, reactions such as "likes," mentions through mentions, etc.), or a history of message exchange with the target account. Note that a retweet is a comment or the like that quotes a post from another account or one's own account. A mention is a comment or the like that includes a specific account name.

[0097] The friend information acquisition unit 103 acquires account information of the identified friend account from the social media system 200. The method of acquiring information from the social media system 200 is the same as that of the posted information acquisition unit 101, and the account information is acquired using an API of a server or the like. The friend information acquisition unit 103 extracts friend information from the account information of all acquired friend accounts. The friend information is location information related to the friend account, such as a residence (residential area) extracted from the account information. The friend information acquisition unit 103 extracts residence information from profile information included in the account information. It is not limited to a residence, but other bases of activity such as a hometown, workplace, or school may also be extracted. Note that the friend information is an example of location information for estimating a place of activity (a place of affiliation) of a friend from the friend's account information, and is not limited to a base of activity such as a residence, but may also be a place where posted information is posted.

[0098] The friend distribution generation unit 104 generates a friend distribution (second location distribution) of the friend accounts based on the friend information (activity bases) of the friend accounts. The friend distribution generation unit 104 generates a friend distribution of the residences of the extracted friend accounts. Like the post distribution, the friend distribution is a distribution of friends' residences (friend locations) in physical space (a spatial distribution specific to the residences of friends). The granularity level of the distribution area of ​​the friend distribution is the same as that of the post distribution, but may be different. Like the post distribution generation unit 102, the friend distribution generation unit 104 calculates the friend distribution using a distribution function of a nonparametric method such as a kernel density estimation function, but may also calculate the friend distribution using other statistical processing. When generating (calculating) the friend distribution, each piece of residence information may be weighted based on the residence information.

[0099] The activity area estimation unit 105 estimates the activity area of ​​the target user based on the generated post distribution and the generated friend distribution. The activity area estimation unit 105 generates the activity area distribution of the target user by overlapping the post distribution and the friend distribution. The granularity level of the generated activity area distribution is the same as the granularity of the post distribution and / or the friend distribution, but may be different. The activity area estimation unit 105 estimates the activity area based on the overlap (amount of overlap) between the post distribution and the friend distribution. The overlap of the distributions is represented by the scores of the post distribution and the friend distribution, respectively, calculated using a kernel density estimation function. That is, the activity area is estimated based on the score of the post distribution obtained using the kernel density estimation function and the score of the friend distribution obtained using the kernel density estimation function. The activity area estimation unit 105 estimates the activity area based on a predetermined calculation result of the score of the post distribution and the score of the friend distribution, respectively, calculated using the kernel density estimation function. For example, the product of the score of the post distribution and the score of the friend distribution is calculated, and the area with the highest score is determined as the activity area. Note that addition, subtraction, etc. may also be used instead of multiplication. By multiplying or adding the score of the post distribution and the score of the friend distribution, the target user's everyday activity area can be estimated. By subtracting the score of the friend distribution from the score of the post distribution, the unusual activity area can be estimated. The activity area estimation unit 105 may determine an area whose calculated score is equal to or greater than a predetermined value as the activity area, or may determine an area with the top N scores (top 5, etc.) as the activity area.

[0100] The activity area output unit 106 outputs the estimated activity area. The activity area output unit 106 may be used as a display device to display the activity area in a predetermined format using a GUI (Graphical User Interface). The post distribution and friend distribution may be displayed, and areas where the distributions overlap may be highlighted. For example, the score for each activity area may be displayed in a heat map format. Alternatively, the score may be output to the outside as a file in a predetermined format. For example, the score for each activity area may be output in a list format, with only a predetermined number of items being output.

[0101] 6 shows an example of the operation of the activity area estimation device (activity area estimation method) according to this embodiment. As shown in FIG. 6, first, the activity area estimation device 100 identifies the target account of the target user (S101). The posted information acquisition unit 101 accepts input of information related to the target account and identifies the target account based on the input information. The account may be identified by inputting the account ID (identification information) of the target account, or the account may be identified by searching social media or the Internet using the input name, keywords, etc.

[0102] Next, the activity area estimation device 100 acquires posted information of the target account (S102). The posted information acquisition unit 101 accesses a server or database of the social media system 200 and acquires publicly available account information of the target account that is available. For example, the account information of the target account is acquired to the extent possible using an API of the social media service, etc. The posted information acquisition unit 101 acquires all posted information included in the account information of the target account.

[0103] Next, the activity area estimation device 100 extracts the posting location and posting date and time of the posted information (S103). The posted information acquisition unit 101 extracts the posting location and posting date and time from all posted information of the target account. Note that the extraction is not limited to all posted information, and the posting location and posting date and time may be extracted from some of the posted information. For example, posted information older than a predetermined date and time may be excluded from the extraction, or if there are two posted information with the same content, one of the posted information may be excluded from the extraction. If a geotag is assigned to the posted image, the posted information acquisition unit 101 acquires the posting location (location information) from the geotag. If a geotag is not assigned to the posted image, the posted location may be acquired from a building, landscape, or the like that can identify the location by image analysis of the posted image. If location information cannot be acquired from the posted image, the posted location may be acquired from words that can identify the location by natural language processing the text of the posted message. If the posted information acquisition unit 101 cannot acquire the posting location from the posted information, the posted information acquisition unit 101 may exclude the posted information from the information for generating the post distribution. Furthermore, the posted information acquisition unit 101 acquires the date and time attached to the posted information as the posted date and time.

[0104] Next, the activity area estimation device 100 generates a post distribution of the target account (S104). The post distribution generation unit 102 generates the post distribution based on the posting locations and posting dates and times of the extracted plurality of pieces of post information. In this example, the post distribution generation unit 102 uses a kernel density estimation function to calculate the post distribution p(L p ) is calculated. p ) is a set of kernel density estimates (scores) of posted information in each distribution area.

[0105]

number

[0106] In equation (1), l p is the set of posting locations, h p is the bandwidth for posting, w p is the weight for posting, K pis the kernel function for posts. The bandwidth is a parameter that indicates the range of influence of each sample in kernel density estimation. The bandwidth for posts is a predetermined value for the distribution of posts, and may be set in advance or may be a value obtained by learning from multiple posting locations in advance. The bandwidth for posts may be changed depending on the output activity area (estimation result).

[0107] Figure 7 shows an image of the distribution of posts obtained by kernel density estimation. As shown in Figure 7, the posting location of each piece of posted information is plotted on a two-dimensional coordinate system of latitude and longitude, and the distribution shows the range of influence of the posting bandwidth (for example, a circle of normal distribution) with the posting location at the center. Within the range of influence of each posting location (sample), the center (posting location) has the highest score, and the score decreases as you move away from the center. In the example shown, higher scores are indicated by darker colors.

[0108] The posting weight in formula (1) is the weight of posted information in the posting distribution based on each piece of posted information. The posting weight indicates the degree of importance of each piece of posted information and sets the magnitude of the score. As an example, the posting weight is a weight based on the posting date and time of the posted information. For example, as shown in Figure 8, the importance of posted information and the elapsed time are inversely proportional to each other, with importance decreasing as time passes. For this reason, the newer the posted information, the higher the weight (higher the importance), and the older the posted information, the lower the weight (lower the importance). By changing the weight in formula (1) according to the posting date and time, the scope of influence remains unchanged, but the newer the information, the higher the score and the older the information, the lower the score.

[0109] Meanwhile, after identifying the target account (S101), the activity area estimation device 100 identifies friend accounts (S105). The friend information acquisition unit 103 identifies friend accounts that have a friend relationship with the target account, etc., from the account information of the target account. For example, an account that is registered as a friend relationship in the account information of the target account is considered to be a friend account. In addition, friend accounts may include accounts that follow or follow posts of the target account, accounts that have posted information that quotes posted information of the target account, accounts that have given "likes" to posted information of the target account, accounts with which messages have been exchanged, and users of other accounts that have been in the same place at the same time as the user of the target account.

[0110] Next, the activity area estimation device 100 acquires friend information of the friend accounts (S106). Similar to acquiring the account information of the target account, the friend information acquisition unit 103 acquires account information of all friend accounts to the extent possible using the API of the social media service from the server of the social media system 200 or the like.

[0111] Next, the activity area estimation device 100 extracts residence information of the friend accounts (S107). The friend information acquisition unit 103 extracts residence information from the account information of all acquired friend accounts. The friend information acquisition unit 103 acquires profile information from the friend's account information and acquires residence information registered in the profile information. If the residence cannot be acquired from the profile information, the residence information may be the hometown, workplace, school, or other activity base registered in the profile information. Posting locations may be extracted from posted information, and locations with a high posting frequency may be used as residence information. Furthermore, if residence information cannot be acquired from the account information of the friend account, the residence of the friend may be estimated from the account information of the friend's friend (other friend) who is further friends of the friend. For example, the residence of the friend may be estimated based on the distribution of residences obtained from the account information of the friend's further friends. In other words, the friend distribution may be generated based on the residences of the friends identified from the residences of the friend's further friends. If the friend information acquisition unit 103 cannot acquire the residence information of a friend account, the friend information acquisition unit 103 may exclude the information of that friend account from the information for generating the friend distribution.

[0112] Next, the activity area estimation device 100 generates a friend distribution of the friend accounts (S108). The friend distribution generation unit 104 generates the friend distribution based on the residential address information of the extracted multiple friend accounts. In this example, the friend distribution generation unit 104 uses a kernel density estimation function to generate the friend distribution p(L f ) is calculated. f ) is a set of kernel density estimates (scores) of friend information for each distribution area.

[0113]

number

[0114] In equation (2), l f is the set of friends' residences, h f is the friend bandwidth, w f is the weight for friends, K fis the friend kernel function. The friend bandwidth is a predetermined value for friend distribution, and like the posting bandwidth, it may be set in advance or may be a value obtained by learning from the residences of multiple friends. The friend bandwidth may be different from the posting bandwidth, or may be the same. The friend bandwidth may be changed depending on the output activity area (estimation result).

[0115] The friend weight in Equation (2) is the weight of the friend information (place of residence) in the friend distribution based on each friend information (account information). The friend weight indicates the degree of importance of each friend information and sets the magnitude of the score. As an example, the friend weight may be based on the time when the user became friends with the target user (became friends or connected). For example, if the date and time when the user became friends with the target user can be obtained, the weight of old friend information may be lowered (less important), and the weight of new friends may be higher (more important). This is because if the target user moves, old friends may reside near their original address. Conversely, new friends may be weighted less. For example, if there is a city that the user aspires to live in, it is estimated that the user has been friends with people living in that city before moving in order to gather information about the city. In such cases, older friends may be given more importance. As a specific calculation method, the weight value may be set to an initial value (100), and this weight value may be decreased based on the time elapsed since the user became friends with the target user. In a simple example, weight can be calculated using a linear function such as weight = ax + b (where a is a negative value, x is the number of days elapsed, and b is the initial value of 100). Alternatively, a certain reference date can be set, and if a friend became a friend within x days, a certain weight can be assigned, and if a friend became a friend more than x days ago, no weight can be assigned.

[0116] The weight for friends may also be a weight based on conversation frequency, such as the number of mentions or retweets of the target user's account. For example, a friend who has conversations with the target user more frequently than other friends may be weighted (emphasized). As a specific calculation method, the total number of conversations with the target user may be used as the denominator and the number of conversations with each friend may be used as the numerator to assign a weight to the friend, or friends with whom the target user has conversations more than a certain number of times may be weighted and friends with conversations less than the certain number may not be weighted.

[0117] Furthermore, the friend weight may be a weight based on the reliability of the friend account. Since fake accounts that misrepresent information exist among social media users, if such fake accounts are included in the friend list, estimation may be performed without prioritizing the friend's information. The reliability indicates the degree of reliability of an account, with a higher reliability indicating a higher reliability. The reliability may also be a numerical index calculated based on distance. The activity area estimation device 100 may further include a reliability calculation unit (not shown), which calculates the reliability based on the account's personal attribute information. For example, the reliability calculation unit acquires personal attribute information (such as profile information) of the target account whose reliability is to be calculated and personal attribute information of the target account's friend accounts, and estimates the personal attributes of the target account from the personal attribute information of the friend accounts. If the personal attribute information of the friend account includes a residence, the residence of the user of the target account is estimated based on the physical distance from the residence. Furthermore, the reliability is calculated based on the distance between the acquired personal attribute information (residence) of the target account and the estimated personal attribute information (residence) of the target account. For example, the reliability calculated by the reliability calculation unit (or a value based on the reliability) is set as the friend weight.

[0118] The friend weight may also be a weight based on the offline friend degree of a friend. An offline friend is a friend account that has a friend relationship with the target user on social media and that also has a friend relationship (connection) with the target user in physical space (the real world). The estimation may be performed by prioritizing information about offline friends over information about online friends. The offline friend degree indicates whether an offline friend relationship has also been formed in physical space. The activity area estimation device 100 may further include an offline friend determination unit, and the offline friend determination unit may calculate a score indicating the degree of offline friend for each friend account of the target user. Specific examples of the offline friend determination unit and a method for calculating the offline friend degree will be described in the embodiments described later. For example, the offline friend degree calculated by the offline friend determination unit (or a value based on the offline friend degree) is set as the friend weight.

[0119] Figure 9 shows an image of the friend distribution obtained by kernel density estimation. As shown in Figure 9, similar to the post distribution, the residence of each friend is plotted on a two-dimensional coordinate system of latitude and longitude, and the distribution shows the range of influence of the friend bandwidth (for example, a circle of normal distribution) centered on the friend's residence.

[0120] Following the generation of the posting distribution and the friend distribution, the activity area estimation device 100 generates the activity area distribution of the target user (S109). The activity area estimation unit 105 generates the activity area distribution of the target user by superimposing the posting distribution and the friend distribution of the same area (space). For example, the activity area estimation unit 105 calculates the activity area l of the target user by taking the product of the posting distribution and the friend distribution obtained from the above equations (1) and (2), as shown in the following equations (3) and (4). t Estimate (estimated activity area).

[0121]

number

[0122]

number

[0123] In equation (3), L is l f and p As shown in equation (4), the score p(L) of each distribution area is proportional to the score of the post distribution and the score of the friend distribution, and as shown in equation (3), the area with the highest score p(L) is estimated to be the activity area.

[0124] Figure 10 shows an image of the posting distribution and friend distribution overlaid on the same space (coordinates). As shown in Figure 10, the range of influence of each location in the posting distribution is overlaid on the range of influence of each location in the friend distribution. The area where the distribution of friends' residences and posting locations overlap is the activity area, and the greater the amount of overlap (the denser the area), the greater the activity area.

[0125] Next, the activity area estimation device 100 outputs the generated activity area distribution (S110). The activity area output unit 106 displays the generated activity area distribution in a predetermined format. FIG. 11 shows an example of how the activity area distribution is displayed. As shown in FIG. 11, for example, the activity area distribution is displayed as a heat map. In the heat map, a distribution of colors and intensities according to the score of each area is displayed on a map (such as a world map, a map of Japan, or a regional map).

[0126] As described above, in this embodiment, a place with a strong trace of activity, such as a place with a personal connection, is considered to be an activity area. Specifically, a distribution based on friend information (place of residence) and a distribution based on posted information (place of posting) are generated simultaneously and in parallel, and the activity area distribution of the target user is generated by overlaying them.

[0127] In this embodiment, an estimation method that does not require prior model preparation is used, eliminating the need to prepare large amounts of data. Specifically, kernel density estimation is used, which does not require parameter learning using large amounts of data. In addition, data collection costs can be reduced by limiting the information used for estimation to the residential locations of the target user's friends and the target user's own posting locations. Furthermore, collection costs can be reduced both during learning and estimation.

[0128] In addition, in this embodiment, the activity area of ​​a target user can be estimated using two types of information. Specifically, the information used for estimation is the target user's friend's residence location and the target user's own posting location. This makes it possible to estimate the activity area even for a target user for whom only one of the two types of information can be obtained. Furthermore, by limiting the information to the above two types, it is possible to reduce collection costs.

[0129] (Embodiment 6-2) Hereinafter, a 6-2 embodiment will be described with reference to the drawings. In this embodiment, an example in which posted information and friend information are filtered in the activity area estimation device 100 of the 6-1 embodiment will be described.

[0130] Fig. 12 shows an example of the configuration of the activity area estimation device 100 according to this embodiment. As shown in Fig. 12, the activity area estimation device 100 according to this embodiment includes a posted information filter unit 107 and a friend information filter unit 108 in addition to the configuration of the embodiment 6-1.

[0131] The posted information filter unit 107 filters the posted information of the target account acquired by the posted information acquisition unit 101 according to a predetermined condition. The posted information filter unit 107 is a selection unit (first selection unit) that selects posted information to be used to generate a post distribution from multiple pieces of posted information included in the account information of the target user. The posted information filter unit 107 selects posted information based on the granularity of the posting location, and, for example, excludes posted information whose granularity of the posting location is greater than a predetermined granularity level. As a specific example, posted information with a granularity of country or prefecture, which is greater than the city, ward, town, or village unit, may be excluded, or posted information with a granularity of 1 km x 1 km or 100 m x 100 m, which is greater than 10 m x 10 m, may be excluded.

[0132] The friend information filter unit 108 filters the friend information of the friend accounts acquired by the friend information acquisition unit 103 according to predetermined conditions. The friend information filter unit 108 is a selection unit (second selection unit) that selects residence information to be used to generate a friend distribution from multiple pieces of residence information (activity base information) included in the friend's account information. As with posted information, the friend information filter unit 108 selects residence information based on the granularity of the residence information, and, for example, excludes friend information whose granularity of residence information is greater than a predetermined granularity level.

[0133] FIG. 13 shows an example of the operation of the activity area estimation device according to this embodiment. As shown in FIG. 13, after extracting the posting location and posting date and time (S103), the posted information filter unit 107 filters the posted information (S111). The posted information filter unit 107 determines the granularity of the posting location of each piece of extracted posted information, and if the granularity of the posting location is greater than a predetermined granularity level, excludes the posted information from the information for generating a posted distribution. For example, the predetermined granularity level is the granularity level of the posted distribution to be generated (or the activity area distribution to be output). Subsequently, the posted distribution generation unit 102 generates a posted distribution using the filtered posted information (S104), as in the 6-1 embodiment.

[0134] In this example, posted information is filtered according to the granularity of the posting location, but filtering may be performed based on other criteria. Posted information may also be filtered based on the posting date and time used in the posting weight in the 6-1 embodiment. For example, posted information whose posting date and time is older than a predetermined date and time may be excluded.

[0135] Furthermore, in this example, the granularity of the posting location is used as the filtering criterion, but the granularity of the posting location may also be used as the posting weight in the 6-1 embodiment. That is, in the above formula (1), the posting weight (wp) may be a weight based on the granularity level of the posting location, and a posting distribution may be generated. For example, the smaller the granularity of the posting location, the more detailed the distribution can be generated. Therefore, the smaller the granularity of the posting location, the larger the weight may be, and the larger the granularity of the posting location, the smaller the weight may be.

[0136] Meanwhile, after extracting the residence information of friends (S107), the friend information filter unit 108 filters the friend information (S112). As with the posted information, the friend information filter unit 108 determines the granularity of the extracted residence information of each friend, and if the granularity of the friend's residence information is greater than a predetermined granularity level, it excludes the friend information from the information used to generate the friend distribution. For example, the predetermined granularity level is the granularity level of the friend distribution to be generated (or the activity area distribution to be output). Next, the friend distribution generation unit 104 generates a friend distribution using the filtered friend information (S108), as in the 6-1 embodiment.

[0137] As with posted information, filtering may be performed based on other criteria, not just the granularity of residence information. Friend information may be filtered based on the time when the user became friends, the frequency of conversations, the reliability of the friend account, the offline friend degree of the friend, and the like, which were used in the friend weights of embodiment 6-1. For example, friend information in which the user became friends with the target user older (or newer) than a predetermined date and time, friend information in which the number of conversations with the target user is a predetermined number or less, friend information in which the reliability of the friend account is a predetermined value or less, friend information in which the offline friend degree is a predetermined value or less, and the like may be excluded.

[0138] Furthermore, similar to the posted information, the granularity of the residence information is not limited to the criterion for filtering, and may be used as the friend weight in the 6-1 embodiment. That is, in the above formula (2) in the 6-1 embodiment, the friend weight (wf) may be set as a weight based on the granularity level of the friend's residence information (activity base), and a friend distribution may be generated. For example, similar to the posted information, the smaller the granularity of the residence information, the larger the weight may be, and the larger the granularity of the residence information, the smaller the weight may be.

[0139] As described above, in this embodiment, the post information that generates the post distribution and the friend information that generates the friend distribution are filtered based on their respective information. This allows the distribution to be generated using information at a predetermined granularity level, thereby obtaining a distribution with desired accuracy.

[0140] (Embodiment 6-3) Hereinafter, a sixth embodiment will be described with reference to the drawings. In this embodiment, an example will be described in which weighting is applied to the post distribution and friend distribution to be superimposed in the activity area estimation device 100 of the sixth embodiment.

[0141] FIG. 14 shows an example of the configuration of the activity area estimation device 100 according to this embodiment. As shown in FIG. 14, the activity area estimation device 100 according to this embodiment includes a weighting unit 109 in addition to the configuration of the 6-1 embodiment. The weighting unit 109 weights the post distribution and friend distribution to be superimposed (weighting of superimposition). For example, the friend distribution and the post distribution may be weighted according to the number of friend information (sample number) in the friend distribution and the number of posted information (sample number) in the post distribution, or weighted according to the difference between the number of friend information and the number of posted information. Alternatively, either the friend distribution or the post distribution may be weighted. The activity area estimation unit 105 estimates the activity area of ​​the target user based on the weighting of the post distribution and / or the friend distribution.

[0142] FIG. 15 shows an example of the operation of the activity area estimation device according to this embodiment. As shown in FIG. 15, after generating the post distribution (S104) and the friend distribution (S108), the weighting unit 109 weights the friend distribution and the post distribution to overlap each other (S113). The weighting unit 109 counts the number of posted information (posting locations) in the generated post distribution and the number of friend information (residential locations) in the generated friend distribution, calculates the difference between the number of posted information and the number of friend information, and weights the post distribution and the friend distribution according to the calculated difference. For example, if there is a large difference between the number of posted information and the number of friend information, there is a risk that one of the pieces of information will be overemphasized. Therefore, it is possible to balance the number of posted information and the number of friend information. For example, if the number of friends is 100 and the number of posts is 200, the friend distribution and the post distribution may be overlapped at a ratio of 2:1.

[0143] Next, the activity area estimation unit 105 generates an activity area distribution by overlaying the weighted friend distribution and post distribution (S109). For example, as shown in the following formula (5), the score p(L) is calculated by multiplying each distribution by the weight WF of the friend distribution and the weight WP of the post distribution.

[0144]

number

[0145] As described above, in this embodiment, when the friend distribution and the post distribution are superimposed, weighting is performed on each distribution. This allows the activity area of ​​the target user to be estimated by placing emphasis on either the friend distribution or the post distribution. For example, by performing weighting based on the number of friends and the number of posts, the activity area can be estimated in a balanced manner.

[0146] (Embodiment 6-4) Hereinafter, a sixth embodiment will be described with reference to the drawings. In this embodiment, as another example of weighting the overlap of the sixth embodiment, an example of weighting the distribution of online friends and the distribution of offline friends will be described.

[0147] FIG. 16 shows an example of the configuration of an activity area estimation device 100 according to this embodiment. As shown in FIG. 16, the activity area estimation device 100 according to this embodiment includes an offline friend determination unit 110 in addition to the configuration of the sixth embodiment. The offline friend determination unit 110 determines offline friends who are friends (connected) with the target user in physical space (the real world) from among friend accounts that are friends with the target user on social media. That is, it determines offline friends and online friends other than offline friends from the friends of the target user. The activity area estimation unit 105 estimates the activity area of ​​the target user based on the posting distribution, the friend distribution of offline friends, and the friend distribution of online friends. The activity area is also estimated based on weighting the friend distribution of offline friends and the friend distribution of online friends.

[0148] FIG. 17 illustrates an example of the operation of the activity area estimation device 100 according to this embodiment. As illustrated in FIG. 17, after extracting the residences of friends (S107), the offline friend determination unit 110 determines offline friends (S114). Based on the acquired account information of the friend accounts, the offline friend determination unit 110 determines whether each friend who has a friend account is also a friend of the target user in the physical space or not. The offline friend determination unit 110 calculates the offline friend degree of the friend account and determines whether the friend is an offline friend or an online friend based on the offline friend degree. For each friend account of the target user, the offline friend determination unit 110 calculates a score indicating the degree of offline friend. For example, if the score exceeds a certain threshold, the offline friend degree is set to a value indicating offline friend (e.g., "1"), and if the score is equal to or less than the threshold, the offline friend degree is set to a value indicating not offline friend (e.g., "0"). The threshold is, for example, set arbitrarily by the user of the activity area estimation device 100.

[0149] The offline friend determination unit 110 may determine whether a friend account is a local account related to a specific region. For example, a local account is a social media account that is operated for a specific location or region. Examples of local accounts include accounts operated by local newspapers, local governments, and community-based businesses such as privately owned restaurants. The offline friend determination unit 110 may calculate the offline friend degree of a friend based on the determination result of whether the friend account is a local account. For example, the offline friend determination unit 110 may refer to the friend information (profile information and posted information) of the friend account, calculate a score depending on whether or not there is information indicating whether the account is operated for a specific location or region, and the amount of such information, and determine whether or not the friend account is a local account.

[0150] Furthermore, if the offline friend determination unit 110 determines that it is unclear whether a friend account is a local account, it may refer to the further friend information of the friend account to determine whether the friend account is a local account. For example, the offline friend degree of the friend account of the target user may be calculated based on whether the account of the further friend of the friend account is a local account. Alternatively, the method described in Non-Patent Document 1 may be used to determine whether an offline friend is an online friend.

[0151] The friend distribution generation unit 104 generates a friend distribution of the determined offline friends and a friend distribution of the online friends (S108). As in the 6-1 embodiment, the friend distribution generation unit 104 generates a friend distribution of the offline friends based on the residence information of the offline friends, and generates a friend distribution of the online friends based on the residence information of the online friends.

[0152] Next, the weighting unit 109 weights the generated friend distribution of offline friends and the generated friend distribution of online friends (S113). For example, offline friends are more important than online friends in terms of the target user's activity area. For this reason, the weighting unit 109 weights the friend distribution of offline friends more than the friend distribution of online friends.

[0153] Next, the activity area estimation unit 105 generates an activity area distribution by superimposing the weighted friend distribution of offline friends and the friend distribution of online friends on the post distribution (S109). Note that the activity area distribution may be generated by superimposing only the friend distribution of offline friends on the post distribution. For example, as shown in the following equation (6), the weight WF of the friend distribution of offline friends off , the weight of the friend distribution of online friends WF on The score p(L) is calculated by multiplying each distribution by the post distribution and taking the product of the multiplied values. Note that it is preferable that the friend weights in this case do not include weights based on the offline friend degree.

[0154]

number

[0155] In addition, in equation (6), h f1 , w f1 is the value in the friend distribution of online friends, and h f2 , w f2 is a value in the friend distribution of offline friends. In other words, when generating a friend distribution for offline friends and a friend distribution for online friends, the bandwidth and friend weights may be set to different values. This allows the generated friend distributions to be different from each other.

[0156] As described above, in this embodiment, the friend distribution is divided into a distribution of only offline friends and a distribution of only online friends, and the offline friend distribution is weighted when overlaying the post distribution. This makes it possible to estimate the target user's activity area by placing emphasis on the friend distribution of offline friends.

[0157] Although the embodiments of the present invention have been described above with reference to the drawings, these are merely examples of the present invention, and various other configurations may be adopted. The configurations of the above-described embodiments may be combined with each other, or some of the configurations may be replaced with other configurations. Furthermore, various modifications may be made to the configurations of the above-described embodiments without departing from the spirit of the invention. Furthermore, the configurations and processes disclosed in the above-described embodiments and modified examples may be combined with each other.

[0158] In this specification, "acquisition" includes at least one of the following: "the device retrieves data stored in another device or storage medium (active acquisition)" based on user input or program instructions, such as receiving data by making a request or inquiry to another device, or accessing and reading out another device or storage medium; "the device inputs data output from another device (passive acquisition)" based on user input or program instructions, such as receiving data that is distributed (or transmitted, push notification, etc.), and selecting and acquiring data from received data or information; and "the device generates new data by editing data (converting it to text, rearranging data, extracting some data, changing the file format, etc.), and then acquires the new data."

[0159] Some or all of the above-described embodiments can be described as, but are not limited to, the following supplementary notes. 1. An activity area estimation means for estimating the activity area of ​​a user of a social media account based on public information linked to the account and published on the Internet; A relationship estimation means for estimating a relationship between a user of the account and the activity area based on the public information; An information processing device having the above. 2. The public information includes any postings made online by the user of the account; The relationship estimation means Inferring the posting location of the posting; An information processing device according to claim 1, which estimates the period when the user of the account was active in the activity area based on the posting period of the post whose posting location is included in the activity area. 3. The relationship estimation means An information processing device as described in 1 or 2, which estimates whether the activity area is the hometown of the user of the account based on the results of comparing the language characteristics used by the user of the account with the language characteristics used in the activity area. 4. The public information includes information indicating connections between accounts on the social media; The relationship estimation means An information processing device described in any one of 1 to 3, which estimates the relationship between the user of the account and the activity area based on the public information that is linked to users of other accounts that have a predetermined relationship with the user of the account and is published on the Internet. 5. The relationship estimation means 5. The information processing device according to 4, wherein the activity area that matches the hometown of the user of the other account is estimated to be the hometown of the user of the account. 6. The relationship estimation means Identifying users of the other accounts associated with the activity area; 6. An information processing device according to claim 4 or 5, which estimates the relationship between the user of the account and the activity area based on the public information linked to the identified other account and published on the Internet. 7. The relationship estimation means Inferring the hobbies and preferences of the user of the identified other account based on the public information linked to the identified other account and published on the Internet; 7. The information processing device according to claim 6, wherein the relationship between the user of the account and the activity area is estimated based on the degree of variation in hobbies and preferences of the users of the identified other accounts. 8. The relationship estimation means If the hobbies and preferences of the identified users of the other accounts vary by more than a reference level, it is estimated that the activity area is the hometown of the user of the account; 8. An information processing device according to claim 7, wherein if the hobbies and preferences of the users of the identified multiple other accounts do not vary by more than a reference level, it is inferred that the activity area is not the hometown of the user of the account. 9. The computer an activity area estimation step of estimating an activity area of ​​a user of a social media account based on public information linked to the account and published on the Internet; a relationship estimation step of estimating a relationship between the user of the account and the activity area based on the public information; An information processing method that performs the above. 10. Computer an activity area estimation means for estimating the activity area of ​​a user of a social media account based on public information linked to the account and published on the Internet; a relationship estimation means for estimating a relationship between the user of the account and the activity area based on the public information; A program that functions as a [Explanation of symbols]

[0160] 10 Estimation device 11 First position distribution generation unit 12 Second position distribution generation unit 13 Estimation part 100 Activity area estimation device 101 Posted Information Acquisition Department 102 Post Distribution Generation Unit 103 Friend Information Acquisition Department 104 Friend distribution generator 105 Activity Area Estimation Department 106 Activity Area Output Unit 107 Posting Information Filter 108 Friend Information Filter 109 Weighting section 110 Offline Friend Identification Unit 200 Social Media Systems 1000 Information Processing Device 1001 Activity Area Estimation Department 1002 Relationship Estimation Unit 1A processor 2A Memory 3A input / output I / F 4A peripheral circuit 5A Bus

Claims

1. an activity area estimation means for estimating an activity area of ​​a user of a social media account based on public information linked to the account and published on the Internet; A relationship estimation means for estimating a relationship between a user of the account and the activity area based on the public information; and The public information includes posts posted on the Internet by the user of the account; The relationship estimation means Inferring the posting location of the posting; An information processing device that estimates a period when the user of the account was active in the activity area based on the posting period of the post whose posting location is included in the activity area.

2. an activity area estimation means for estimating an activity area of ​​a user of a social media account based on public information linked to the account and published on the Internet; A relationship estimation means for estimating a relationship between a user of the account and the activity area based on the public information; and The relationship estimation means An information processing device that estimates whether the activity area is the hometown of the user of the account based on the results of comparing the language characteristics used by the user of the account with the language characteristics used in the activity area.

3. an activity area estimation means for estimating an activity area of ​​a user of a social media account based on public information linked to the account and published on the Internet; A relationship estimation means for estimating a relationship between a user of the account and the activity area based on the public information; and the public information includes information indicating connections between accounts on the social media; The relationship estimation means Inferring whether the activity area is the hometown of the user of the account based on the public information that is linked to users of other accounts that have a predetermined relationship with the user of the account and that is published on the Internet; An information processing device that estimates that the activity area that matches the hometown of the user of the other account is the hometown of the user of the account.

4. The relationship estimation means Identifying users of the other accounts associated with the activity area; The information processing device according to claim 3 , wherein the information processing device estimates whether the activity area is the hometown of the user of the account based on the public information linked to the identified other account and published on the Internet.

5. The relationship estimation means Inferring the hobbies and preferences of the user of the identified other account based on the public information linked to the identified other account and published on the Internet; The information processing device according to claim 4 , wherein the information processing device estimates whether the activity area is the hometown of the user of the account based on the degree of variation in hobbies and preferences of the users of the identified multiple other accounts.

6. The relationship estimation means If the hobbies and preferences of the identified users of the other accounts vary by more than a reference level, it is estimated that the activity area is the hometown of the user of the account; The information processing device according to claim 5 , wherein, when the hobbies and preferences of the identified users of the other accounts do not vary by more than a reference level, it is estimated that the activity area is not the hometown of the user of the account.

7. The computer an activity area estimation step of estimating an activity area of ​​a user of a social media account based on public information linked to the account and published on the Internet; a relationship estimation step of estimating a relationship between the user of the account and the activity area based on the public information; Run The public information includes posts posted on the Internet by the user of the account; In the relationship estimation step, Inferring the posting location of the posting; An information processing method for estimating the period when the user of the account was active in the activity area based on the posting period of the post whose posting location is included in the activity area.

8. The computer an activity area estimation step of estimating an activity area of ​​a user of a social media account based on public information linked to the account and published on the Internet; a relationship estimation step of estimating a relationship between the user of the account and the activity area based on the public information; Run In the relationship estimation step, An information processing method that estimates whether the activity area is the hometown of the user of the account based on the results of comparing the language characteristics used by the user of the account with the language characteristics used in the activity area.

9. Computer, an activity area estimation means for estimating the activity area of ​​a user of a social media account based on public information linked to the account and published on the Internet; a relationship estimation means for estimating a relationship between the user of the account and the activity area based on the public information; It functions as The public information includes posts posted on the Internet by the user of the account; The relationship estimation means Inferring the posting location of the posting; A program that estimates the period when the user of the account was active in the activity area based on the posting time of the post whose posting location is included in the activity area.

10. Computer, an activity area estimation means for estimating the activity area of ​​a user of a social media account based on public information linked to the account and published on the Internet; a relationship estimation means for estimating a relationship between the user of the account and the activity area based on the public information; It functions as The relationship estimation means A program that estimates whether the activity area is the hometown of the user of the account based on the results of comparing the language characteristics used by the user of the account with the language characteristics used in the activity area.

Citation Information

Patent Citations

  • Plate type heat exchanger

    JP1986076889A

  • Identification information management support system, identification information management support method, and program

    JP2013122630A

  • Behavior analysis device, behavior analysis method, and program

    JP2018010378A

  • Information processing device, information processing method, and information processing program

    JP2018045532A

  • Proposal device, proposal method and proposal program

    JP2020004211A