Excursion estimation system and excursion estimation method
The migration estimation system uses nonnegative tensor factorization to analyze user terminal data, addressing privacy concerns by estimating group behaviors and interests without personal information, thereby providing accurate insights into user group activities and preferences.
Patent Information
- Application Number
- JP2024106497
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-01
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2044-07-01
AI Technical Summary
Existing technologies struggle to accurately estimate daily behaviors and interest characteristics of user groups without relying on personal information and considering privacy, as they often rely on individual user identification which may be restricted by laws and regulations.
A migration estimation system and method that collects terminal IDs, category information, and contact dates for user terminals, performs nonnegative tensor decomposition, and extracts migration patterns to estimate linguistic meanings of group behaviors and interests using nonnegative tensor factorization.
Accurately estimates daily activities and interest characteristics of specific groups without personal information, enabling cohort analysis that respects privacy.
Smart Images

Figure 2026007041000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a migration estimation system and a migration estimation method. [Background technology]
[0002] Conventionally, there are technologies for analyzing and estimating user attributes and characteristics based on the movement paths of multiple users. For example, Japanese Patent Application Laid-Open Publication No. 2019-23851 (Patent Document 1) discloses a data analysis system in which a user terminal operated by a user, a service platform provider device, and a service receiver device are interconnected via the Internet. The service platform provider device in this data analysis system includes a web server, a log server, and an analysis server. The web server provides web content to the user terminal, and the log server stores the user terminal's location information and user attribute information input from the user terminal as logs. The analysis server includes a database that collects and saves the logs stored in the log server, a statistical processing unit that statistically processes the logs stored in the database, and a rendering unit that creates attribute analysis data and people flow analysis data based on the statistical data processed by the statistical processing unit. In response to a request from the service receiver device, the statistical processing unit of the service platform provider device performs statistical processing on the collected data in accordance with the request, and creates attribute analysis data or people flow analysis data based on the obtained statistical data and transmits it to the service receiver device. Here, the people flow analysis data created by the rendering unit is people flow information in which nodes representing areas where users have moved are arranged in a first direction in the order of their movements based on movement information between spots, and the nodes are connected by links, and for spots where multiple users have moved along different routes, the nodes are arranged in a second direction that intersects with the first direction. This makes it possible to statistically process the vast amount of attribute information and behavior information of users using user devices, making visualization easier and supporting analysis.
[0003] Furthermore, Non-Patent Document 1 (Kubo Motoi et al., "Information Extraction from Tourist Behavior Data Using Nonnegative Tensor Factorization," Multimedia, Distributed Collaboration and Mobile Symposium 2019, Proceedings 2019, pp. 1259-1263, June 26, 2019) discloses that the behavioral patterns of foreign tourists throughout Japan were analyzed using multidimensional tourist behavior log data obtained from a smartphone app, constructing them as a third-order tensor of time of stay, place of stay, and estimated country of residence, and applying nonnegative tensor factorization (NTF). By decomposing the third-order tensor into five clusters, it was possible to extract two clusters for July and August: one centered on American tourists and the other centered on Italian tourists.
[0004] Furthermore, Non-Patent Document 1 (Kumagaya Yusuke et al., "Analysis of the Migratory Behavior of Foreign Tourists Visiting Japan Using Nonnegative Complex Tensor Factorization," IEICE Technical Report (Institute of Electronics, Information and Communication Engineers), Vol. 115, No. 112, pp. 15-19, June 16, 2015) discloses that the migratory behavior of foreign tourists centered around Fukuoka City was analyzed using nonnegative complex tensor factorization (NMTF) based on various information obtained from smartphone apps. It was confirmed that this method was able to extract migratory patterns that are strongly dependent on spatiotemporal information and user attributes. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Publication No. 2019-23851 [Non-patent literature]
[0006] [Patent Document 1] Motoi Kubo et al., "Information Extraction from Tourist Behavior Data Using Nonnegative Tensor Factorization," Proceedings of the 2019 Multimedia, Distributed Collaboration and Mobile Symposium, pp. 1259-1263, 2019-06-26 [Patent Document 2] Yusuke Kumagai et al., "Analysis of the Travel Behavior of Foreign Tourists Visiting Japan Using Nonnegative Complex Tensor Factorization," IEICE Technical Report (Institute of Electronics, Information and Communication Engineers), Vol. 115, No. 112, pp. 15-19, June 16, 2015. Summary of the Invention [Problem to be solved by the invention]
[0007] Conventionally, in attribute analysis of user devices, attribute analysis is performed taking into account the movement and tracking of specific user devices using identification information such as the MAC address and GPS information of the user device. Such identification information is called a traceable ID. Here, a traceable ID refers to identification information that allows you to know what the identification information is and to know that the identification information is the same as something else.
[0008] On the other hand, while currently the acquisition of user device identification information does not violate laws and regulations or privacy, there is a possibility that the acquisition of user device identification information may be restricted from an ESG perspective. Furthermore, in the future, the acquisition of user device identification information may be included in the definition of personal information and its handling as personally identifiable information under the Personal Information Protection Act and other laws, making it difficult to acquire user device identification information at all.
[0009] Therefore, in the future, rather than analyzing the attributes of a specific user device based on its movements and tracking, the mainstream approach will likely be to classify multiple user devices into groups based on certain conditions, analyze changes in each group's behavior over time, and analyze anonymous groups (cohort analysis) that abstract the characteristics of the groups.In other words, even if tracking IDs are collected, it will be important to perform cohort analysis that takes privacy into consideration and does not rely on personal information by performing certain processing.
[0010] Here, the technology described in Patent Document 1 mentioned above is based on people flow information of the identification information of each individual user, does not perform pattern extraction or pattern clustering, and simply calculates the causal relationships between time-series events of multiple user terminals, and does not assume cohort analysis.
[0011] Furthermore, the technology described in Non-Patent Document 1 extracts clusters of a specific group of foreign tourists visiting Japan. Furthermore, the technology described in Non-Patent Document 2 extracts migration patterns that are highly dependent on spatiotemporal information and user attributes. However, there is a problem in that it is not possible to clarify what characteristics of daily activities and interests are present in these clusters and migration patterns.
[0012] Therefore, the present invention has been made to solve the above-mentioned problems, and aims to provide a migration estimation system and migration estimation method that can accurately estimate what kind of daily behaviors and interest characteristics a specific group has, without relying on personal information and taking privacy into consideration. [Means for solving the problem]
[0013] The migration estimation system according to the present invention includes a collection control unit, a calculation control unit, an extraction control unit, and an estimation control unit. The collection control unit collects, for each user terminal, the terminal IDs of each of a plurality of user terminals migrating at each location, category information for each location where the user terminal stayed, and the contact date and time when the user terminal contacted the location. The calculation control unit performs nonnegative tensor decomposition on a tensor with the collected plurality of device IDs, the category information for the device IDs, and the contact date and time for the category information as three axes, and calculates, for each device ID, a migration pattern indicating a distribution of stay times calculated from the contact date and time for each category information of the device IDs. The extraction control unit extracts, as a migration cluster, migration patterns of a predetermined number of similar device IDs based on the similarity of the arrangement of stay times for each category information in the calculated migration pattern for each device ID and the similarity of the stay times for each category information. The estimation control unit estimates linguistic meanings that indicate characteristics of the extracted migration cluster based on the proportion of stay time for each category information of the migration cluster of a predetermined number of terminal IDs that belong to the cluster.
[0014] The migration estimation method according to the present invention includes a collection control step, a calculation control step, an extraction control step, and an estimation control step. Each control step of the migration estimation method corresponds to each control unit of the migration estimation system. [Effects of the Invention]
[0015] According to the present invention, it is possible to accurately estimate what kind of daily activities and interest characteristics a specific group has, without relying on personal information and taking privacy into consideration. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a functional block diagram of a wandering estimation system according to an embodiment of the present invention. [Figure 2] 1 is a flowchart showing an execution procedure of a migration estimation method according to an embodiment of the present invention. [Figure 3]FIG. 3A shows an example of obtaining terminal ID, device ID, contact start date and time, and contact end date and time from short-range wireless communication between four wireless devices and a user terminal, and FIG. 3B shows an example of a log table. [Figure 4] FIG. 4A shows an example of an equipment category table and a navigation table, and FIG. 4B shows an example of a GPS table when the terminal ID, GPS information, and communication date and time are obtained using the GPS function of the user terminal. [Figure 5] Figure 5A shows an example of a GPS category table and a migration table, and Figure 5B shows an example of a tensor-format graph in which the rows (X-axis) represent device IDs, the columns (Y-axis) represent category information, and the depth (Z-axis) represents stay time. [Figure 6] Figure 6A shows an example of the arrangement of stay times for each category of information in the navigation pattern of terminal ID "a01" and linear user behavior information, and Figure 6B shows an example of extracting navigation clusters. [Figure 7] FIG. 7A shows an example of calculating the time proportion for each category information in the first migration cluster ("cluster A"), and FIG. 7B shows an example of the time proportion for each category information, with the vertical axis (X-axis) set as the time proportion and the horizontal axis (Y-axis) set as the category information. [Figure 8] Figure 8A shows an example of matching the first semantic hierarchy of a migratory cluster ("cluster A") with a time ratio table, and Figure 8B shows an example of converting from the first semantic hierarchy to the second semantic hierarchy. [Figure 9] FIG. 9A shows an example of a case where the category information in each of the equipment category table and the GPS category table is used as main category information and subcategory information is added, and FIG. 9B shows an example of a case where the first semantic hierarchy is converted to the second semantic hierarchy using the subcategory information. [Figure 10] FIG. 10 is a diagram showing an example of a tensor graph and floor description in which rows (X-axis) are set as terminal IDs, columns (Y-axis) are set as category information, and depth (Z-axis) is set as stay time in an embodiment. [Figure 11] FIG. 10 is a diagram showing an example of an arrangement of stay times for each category of information in the migration pattern of the terminal ID "USER ID 000001" and user behavior linear information in the embodiment. [Figure 12] FIG. 10 is a diagram illustrating an example of extracting a predetermined number of migratory clusters in an embodiment. [Figure 13] FIG. 10 is a diagram showing an example of a time ratio for each piece of category information, in which the vertical axis (X axis) is set as a time ratio and the horizontal axis (Y axis) is set as category information, in an embodiment. [Figure 14] In the embodiment, this figure shows an example of the first and second semantic hierarchies of a first wandering cluster ("top-level cluster") composed of main category information, and the first and second semantic hierarchies of a first wandering cluster utilizing subcategory information. [Figure 15] FIG. 10 is a diagram illustrating an example of generating a user's context from linguistic semantics. [Figure 16] FIG. 10 is a diagram illustrating an example of generating language or images from a user's context. DETAILED DESCRIPTION OF THE INVENTION
[0017] Hereinafter, an embodiment of the present invention will be described with reference to the accompanying drawings to help understand the present invention. Note that the following embodiment is an example of the present invention and is not intended to limit the technical scope of the present invention.
[0018] As shown in FIG. 1, the migration estimation system 1 of the present invention includes a plurality of user terminals 10, a plurality of wireless devices 11 installed at each location, and a server 13 capable of wireless communication with the user terminals 10 and the wireless devices 11 via a network 12.
[0019] Here, the user terminal 10 includes a display unit (output unit) that displays a screen, a reception unit (input unit) that receives input of predetermined instructions through user operation, a communication unit for wireless communication, a storage unit that stores data, and a control unit that controls each unit. The communication unit of the user terminal 10 is capable of wireless communication with both the wireless device 11 and the server 13. The user terminal 10 is also provided with a GPS function, and GPS information is acquired as location information of the user terminal 10 based on the GPS function of the user terminal 10. Examples of the GPS information include the longitude and latitude of the user terminal 10. Furthermore, the user terminal 10 is, for example, a mobile terminal device (smartphone) with a touch panel, a tablet terminal device, a portable notebook computer, etc.
[0020] Further, the wireless device 11 includes a communication unit capable of short-distance wireless communication with the user terminal 10. Short-distance wireless communication means transmitting or receiving data within a range of several tens of centimeters to several hundred meters using radio waves of a wireless signal. Here, when the user terminal 10 enters the radio wave reception area (reception range) of the wireless device 11, it performs short-distance wireless communication with the wireless device 11. The wireless device 11 includes a Wi-Fi sensor, a Beacon terminal, or a combination thereof.
[0021] Furthermore, the location where the wireless device 11 is installed is a place where the user of the user terminal 10 can enter and exit and move around inside and outside, and examples of such places include stores that sell products, supermarkets, mass retailers, stations, roads, universities, schools, parks, hospitals, gambling halls, etc. Furthermore, location information based on GPS information is set in advance for each location.
[0022] The network 12 is communicatively connected to each of the user terminal 10, the wireless device 11, and the server 13. The network 12 includes a LAN (Local Area Network) via a Wifi (registered trademark) access point, a WAN (Wide Area Network) via a wireless base station, a third generation (3G) communication method, a fourth generation (4G) communication method such as LTE, a fifth generation (5G) or later communication method, and a wireless communication network such as Bluetooth (registered trademark).
[0023] The server 13 is a commonly used computer or the like, and includes a communication unit for wireless and wired communication, a storage unit for storing data, and a processing unit for various processes. The server 13 receives data from the user terminal 10 and the wireless device 11 via the network 12, and transmits data to the user terminal 10. The server 13 may be a single server or a combination of multiple servers.
[0024] Furthermore, the user terminal 10, the wireless device 11, and the server 13 each incorporate a CPU, ROM, RAM, HDD, SSD, etc. (not shown), and the CPU uses, for example, the RAM as a work area to execute programs stored in the ROM, HDD, SSD, etc. Furthermore, each control unit (described later) is realized by the CPU executing a program.
[0025] Next, the configuration and execution procedure according to an embodiment of the present invention will be described with reference to Figures 1 to 9. First, when multiple users who own user terminals 10 roam (commute) various places, the collection control unit 101 of the server 13 collects, for each user terminal 10, the terminal ID of each of the multiple user terminals 10 roaming around at each place, category information for each place where the user terminal 10 stayed, and the contact date and time when the user terminal 10 came into contact with the place (Figure 2: S101).
[0026] Here, the method by which the collection control unit 101 collects the information is not particularly limited, but for example, when the location of the user terminal 10 is identified by a wireless device 11 that has been installed in advance, the following occurs: That is, as shown in Fig. 3A, when a user terminal 10 with a terminal ID of "a01" enters the radio wave reception range of a wireless device 11 with a device ID of "abc," it establishes short-range wireless communication with the wireless device 11. Then, the collection control unit 101 acquires the device ID "abc" of the wireless device 11, the terminal ID "a01" of the user terminal 10, and the date and time when contact with the user terminal 10 via short-range wireless communication began (for example, "2023 / 4 / 11 7:30"). Furthermore, when a user terminal 10 with a terminal ID of "a01" leaves the radio wave reception range of a wireless device 11 with a device ID of "abc," the collection control unit 101 acquires the device ID "abc" of the wireless device 11, the terminal ID "a01" of the user terminal 10, and the contact end date and time of the user terminal 10 via short-range wireless communication (for example, "2023 / 4 / 11 7:40").Then, the collection control unit 101 creates a log table that associates the device ID "abc" of the wireless device 11, the terminal ID "a01" of the user terminal 10, the contact start date and time "2023 / 4 / 11 7:30," and the contact end date and time "2023 / 4 / 11 7:40."
[0027] 3B, the log table 300 stores a device ID 301, a terminal ID 302, a contact start date and time 303, and a contact end date and time 304 in association with each other. The device ID is identification information for identifying the wireless device 11, and may be, for example, a MAC address. The terminal ID is identification information for identifying the user terminal 10, and may similarly be, for example, a MAC address.
[0028] There are no particular limitations on the method by which the collection control unit 101 acquires the device ID, the terminal ID, and the contact start date and time or the contact end date and time. For example, when the wireless device 11 is a Wi-Fi sensor, the user terminal 10 normally periodically transmits radio waves (beacons, omnidirectional radio waves) to search for access points for wireless LAN (Local Area Network) communication. These radio waves contain the terminal ID ("a01") of the user terminal 10. Therefore, when the user terminal 10 enters the radio wave reception range of the Wi-Fi sensor 11, the Wi-Fi sensor 11 receives the radio waves from the user terminal 10 and transmits the device ID ("abc") for identifying the Wi-Fi sensor 11, the terminal ID ("a01") in the radio waves, and the contact start date and time "2023 / 4 / 11 7:30" to the server 13 via the network 12. At this time, radio wave intensity may also be transmitted. Furthermore, when the user terminal 10 leaves the radio wave reception range of the Wi-Fi sensor 11, the Wi-Fi sensor 11 stops receiving radio waves from the user terminal 10, and transmits the device ID ("abc"), the terminal ID ("a01") whose radio wave reception has stopped, and the contact end date and time ("2023 / 4 / 11 7:40") to the server 13 via the network 12. Note that the Wi-Fi sensor 11 may also identify the contact end date and time by detecting that the user terminal 10 has stopped receiving radio waves.
[0029] On the other hand, when the wireless device 11 is a Beacon terminal, the Beacon terminal 11 periodically transmits radio waves for short-range wireless communication. These radio waves include a device ID ("abc") for identifying the Beacon terminal 11. Therefore, when the user terminal 10 enters the radio wave reception range of the Beacon terminal 11, the user terminal 10 receives radio waves from the Beacon terminal 11 and transmits the device ID ("abc") in the radio waves, the terminal ID ("a01") of the user terminal 10, and the contact start date and time ("2023 / 4 / 11 7:30") to the server 13 via the network 12. In this case, a specific program (application, etc.) that (automatically) transmits information to the Beacon terminal 11 to the server 13 can be installed in advance in the user terminal 10. Furthermore, when the user terminal 10 leaves the radio wave reception range of the Beacon terminal 11, the user terminal 10 stops receiving radio waves from the Beacon terminal 11, and transmits the device ID ("abc"), terminal ID ("a01"), and contact end date and time ("2023 / 4 / 11 7:40") for which radio wave reception has stopped to the server 13 via the network 12. The user terminal 10 may also determine the contact end date and time by detecting that reception of radio waves from the Beacon terminal 11 has stopped.
[0030] In this way, regardless of whether the wireless device 11 is a Wi-Fi sensor or a Beacon terminal, the collection control unit 101 of the server 13 can receive and acquire the device ID, terminal ID, and contact start date and time or contact end date and time from one wireless device 11 via the network 12 by short-range wireless communication with a user terminal 10 located in the vicinity of the wireless device 11.
[0031] For example, if four wireless devices 11 with device IDs "abc," "def," "ghi," and "jkl" are installed in four locations as shown in FIG. 3A, when a user possessing a user terminal 10 with terminal ID "a01" wanders around a location of interest to the user and approaches and stays near a wireless device 11 with device ID "abc," the collection control unit 101 associates and stores the device ID 301 ("abc"), terminal ID 302 ("a01"), contact start date and time 303 ("2023 / 4 / 11 7:30"), and contact end date and time 304 ("2023 / 4 / 11 7:40") in the log table 300, as shown in FIG. 3B. Furthermore, when a user of a user terminal 10 with terminal ID "a01" approaches and stays near a wireless device 11 with device ID "jkl", the collection control unit 101 associates and stores in the log table 300 the device ID 301 ("jkl"), terminal ID 302 ("a01"), contact start date and time 303 ("2023 / 4 / 11 8:00"), and contact end date and time 304 ("2023 / 4 / 11 8:20").
[0032] Once the contact start date and time 303 and contact end date and time 304 for one terminal ID 302 are stored in association with one device ID 301, the collection control unit 101 references a predetermined device category table. As shown in Fig. 4A, the device category table 400 stores the device ID 401 (e.g., "abc") in association with category information 402 (e.g., "PC01") of the location where the wireless device 11 with that device ID is installed. The category information is preset as information indicating the characteristics of the location, and examples of such category information include office, education, temples and shrines, leisure, amusement, sports, learning, culture, retail stores, eating out, clinics, and transportation.
[0033] Therefore, the collection control unit 101 references the device ID 401 ("abc") in the device category table 400 that corresponds to the device ID 301 ("abc") in the log table 300, and acquires the category information 402 ("PC01") associated with the referenced device ID 401 ("abc"). The collection control unit 101 then creates a migration table that associates the terminal ID ("a01"), the acquired category information ("PC01"), the contact start date and time ("2023 / 4 / 11 8:00"), and the contact end date and time ("2023 / 4 / 11 8:20").
[0034] 4A, the migration table 403 stores a terminal ID 404 ("a01"), category information 405 ("PC01"), and contact date and time 406 {contact start date and time ("2023 / 4 / 11 7:30") and contact end date and time ("2023 / 4 / 11 7:40")} in association with each other. Also, as described above, when a user of a user terminal 10 with a terminal ID ("a01") migrates around the installation location of a wireless device 11 with a device ID ("jkl"), the collection control unit 101 stores the terminal ID 404 ("a01"), category information 405 ("PC04"), and contact date and time 406 {contact start date and time ("2023 / 4 / 11 8:00") and contact end date and time ("2023 / 4 / 11 8:20")} in association with each other. In this way, the migration table 403 can store the terminal ID 404 of the user terminal 10 of the user who roamed around each wireless device 11, category information 405 of the installation location of the wireless device 11, and contact date and time 406 in a database.
[0035] In the above description, the wireless device 11 is used to create the migration table for the user terminal 10. However, if the wireless device 11 is not present in the migration location of the user terminal 10 and the migration location of the user terminal 10 is identified by the GPS function of the user terminal 10, the situation will be as follows. That is, as shown in FIG. 4B , a user of the user terminal 10 with a terminal ID of "a02" migrates to a location of interest to the user and transmits the terminal ID "a02," GPS information obtained by the GPS function (e.g., "X21, Y21"), and the communication date and time (e.g., "2023 / 4 / 11 7:30") to the server 13 via the network 12. The collection control unit 101 then creates a GPS table that associates the terminal ID "a02," the GPS information ("X21, Y21"), and the communication date and time ("2023 / 4 / 11 7:30").
[0036] 4B, the GPS table 407 stores a terminal ID 408, GPS information 409, and communication date and time 410 in association with each other. For example, as shown in FIG. 4B, when a user of a user terminal 10 with a terminal ID of "a02" visits and stays at a place of interest that is identified by GPS information ("X20, Y20"), the collection control unit 101 stores the terminal ID 408 ("a02"), GPS information 409 ("X21, Y21"), and communication date and time 410 ("2023 / 4 / 11 7:30") in association with each other in the GPS table 407, and also stores the terminal ID 408 ("a02"), GPS information 409 ("X22, Y22"), and communication date and time 410 ("2023 / 4 / 11 7:45") in association with each other in the GPS table 407. Furthermore, when this user travels around and stays at a location identified by other GPS information ("X30, Y30"), the collection control unit 101 associates and stores in the GPS table 407 the terminal ID 408 ("a02"), GPS information 409 ("X31, Y31"), and communication date and time 410 ("2023 / 4 / 11 8:00"), and also associates and stores the terminal ID 408 ("a02"), GPS information 409 ("X32, Y32"), and communication date and time 410 ("2023 / 4 / 11 8:20").
[0037] Then, every time the collection control unit 101 receives information such as the terminal ID from the user terminal 10, the collection control unit 101 stores the received information such as the terminal ID in the GPS table 407. As a result, even in a location where no wireless device 11 is installed, the location where the user terminal 10 has roamed can be identified by using the GPS function of the user terminal 10.
[0038] Now, when the communication date and time 410 of one terminal ID 408 is associated with a predetermined number of adjacent GPS information 409, the collection control unit 101 refers to a predetermined GPS category table. Here, as shown in Fig. 5A, GPS information 501 indicating a specific location (e.g., "X20, Y20") and category information 502 of the location (e.g., "PC02") are stored in association with each other in the GPS category table 500.
[0039] Therefore, the collection control unit 101 references GPS information 501 ("X20, Y20") in the GPS category table 500, which is included within a predetermined range (for example, a circular range with a predetermined radius, a square range with predetermined sides, etc.) centered on GPS information 409 ("X21, Y21", "X22, Y22") in the GPS table 407, and acquires category information 402 ("PC02") associated with the referenced GPS information 501 ("X20, Y20"). The collection control unit 101 then acquires, as the contact start date and time, the communication date and time 410 ("2023 / 4 / 11 7:30") indicating the first date and time from the GPS information 409 ("X21, Y21", "X22, Y22") in the GPS table 407, which includes GPS information 501 ("X20, Y20") in the GPS category table 500. Furthermore, the collection control unit 101 selects the GPS information 409 ("X21, Y21" and "X22, Y22") of the GPS table 407 that includes the GPS information 501 ("X20, Y20") of the GPS category table 500. The communication date and time 410 indicating the last date and time ("2023 / 4 / 11 7:45") is acquired as the contact end date and time. Furthermore, the collection control unit 101 creates a migration table that associates the terminal ID ("a02"), the acquired category information ("PC02"), the acquired contact start date and time ("2023 / 4 / 11 7:30"), and the contact end date and time ("2023 / 4 / 11 7:45").
[0040] 5A, the migration table 503 stores a terminal ID 504 ("a02"), category information 505 ("PC02"), and contact date and time 506 {contact start date and time ("2023 / 4 / 11 7:30"), contact end date and time ("2023 / 4 / 11 7:45")} in association with each other. Furthermore, as described above, when the user terminal 10 with the terminal ID ("a02") roams a location with other GPS information ("X30, Y30"), the collection control unit 101 stores the terminal ID 504 ("a02"), category information 505 ("PC03"), and contact date and time 506 {contact start date and time ("2023 / 4 / 11 8:00"), contact end date and time ("2023 / 4 / 11 8:20")} in association with each other. In this way, even in a location where wireless device 11 is not installed, by utilizing the GPS function of user terminal 10, it is possible to create a database of terminal ID 504 of user terminal 10 that has roamed each location, category information 505 of the roaming location, and contact date and time 506.
[0041] Now, when the collection control unit 101 completes collection, the calculation control unit 102 of the server 13 performs non-negative tensor decomposition on a tensor whose three axes are the collected multiple terminal IDs, the category information of the terminal IDs, and the contact date and time of the category information, and calculates, for each terminal ID, a migration pattern indicating the distribution of stay times calculated from the contact date and time for each category information of the terminal ID (Figure 2: S102).
[0042] Here, the calculation method performed by the calculation control unit 102 is not particularly limited. For example, data on category information and contact date and time associated with one device ID can be expressed in tensor format. A tensor is an extension of a matrix and refers to data having three axes: row (X axis), column (Y axis), and depth (Z axis). When data on category information and stay time for one device ID is converted into tensor format, the row axis represents the device ID, the column axis represents the category information, and the depth axis represents the contact date and time. Non-negative tensor factorization (NTF) is effective for discovering clusters present in tensor format data. A cluster is, for example, a group classified based on a certain degree of similarity and refers to a group having similar characteristics. When NTF is applied to tensor format data, the tensor format data is decomposed into a product of matrices of lower rank (hereinafter referred to as "low-order matrices"). Each low-order matrix indicates the contribution of each row, column, and depth in the tensor-format data to the cluster, making it possible to discover clusters in the tensor-format data.
[0043] In the present invention, as shown in FIG. 5B, the calculation control unit 102 uses a navigation table to set rows (X-axis) as terminal IDs and columns (Y-axis) as category information. Next, for a contact date and time specified by the terminal ID and category information, the calculation control unit 102 calculates the user's stay time at the location of the category information by subtracting the contact start date and time from the contact end date and time. For example, for a contact date and time specified by the terminal ID "a01" and category information "PC01," the stay time ("10") is calculated by subtracting the contact start date and time ("2023 / 4 / 11 7:30") from the contact end date and time ("2023 / 4 / 11 7:40"). Here, the calculation control unit 102 calculates the stay time for each category information from the contact date and time for each category information belonging to the terminal ID. For category information "PC04" belonging to terminal ID "a01," the calculation control unit 102 calculates the stay time ("20") by subtracting the contact start date and time ("2023 / 4 / 11 8:00") from the contact end date and time ("2023 / 4 / 11 8:20"). On the other hand, the calculation control unit 102 treats the stay time as "null" ("0") for category information in which there is no contact date and time.
[0044] The calculation control unit 102 then sets the depth (Z axis) as the stay time, and treats the terminal ID, category information, and stay time as a third-order (three-dimensional) tensor. Here, the stay time for each category information of the terminal ID reflects the behavior history (contact history) of the user of the user terminal 10 of the terminal ID, and is treated as linear behavior information of the user. Furthermore, the calculation control unit 102 performs non-negative tensor factorization on the third-order tensor, setting the basis number to 3, floor 0 as the terminal ID, floor 1 as the category information, and floor 2 as the stay time, and calculates a navigation pattern indicating the arrangement of the stay time for each category information of the terminal ID for each terminal ID. Here, the navigation pattern of the terminal ID corresponds to a low-order matrix in which the stay time for each category information is arranged in accordance with the order of the multiple category information.
[0045] For example, when the migration pattern of terminal ID "a01" is calculated, as shown in FIG. 6A, the migration pattern 601 of terminal ID "a01" indicates an arrangement of stay times for each category information. The order of the multiple category information items corresponding to the Y axis is "PC01," "PC02," "PC03," "PC04," etc., and the multiple stay times calculated from the contact date and time on the Z axis are arranged as "10," "null," "null," "20," etc. for each category information item. Here, when the stay times calculated from the contact date and time on the Z axis are plotted for each category information item on the Y axis, the migration pattern indicating the arrangement of stay times for each category information item is expressed as linear behavior information (low-order matrix) of the user terminal 10 with terminal ID "a01." Then, the calculation control unit 102 calculates the migration pattern of the terminal ID for each terminal ID. This allows the migration pattern of the user terminal 10 for each terminal ID to be calculated as linear behavior information of the user.
[0046] Now, when the calculation control unit 102 completes the calculation, the extraction control unit 103 of the server 13 extracts similar migration patterns of a predetermined number of terminal IDs as migration clusters based on the similarity of the arrangement of stay times for each category information in the calculated migration patterns for each terminal ID and the approximation of the stay times for each category information (Figure 2: S103).
[0047] Here, the extraction control unit 103 may perform extraction by any method. First, the extraction control unit 103 extracts candidate terminal IDs based on the similarity of the arrangement of stay times for each category information in the visit pattern for each terminal ID. For example, the extraction control unit 103 calculates the number of category information items in which the stay time is "null" in the visit pattern 601 of a first terminal ID ("a01") that serves as a reference as the total number of zeros for the reference terminal ID. Next, as shown in FIG. 6B, the extraction control unit 103 compares the visit pattern 601 of the first terminal ID (e.g., "a01") with the visit pattern 602 of another second terminal ID (e.g., "a02"), and calculates the number of category information items in which the stay time is "null" in the category information at the same location as the number of zeros for the comparison terminal ID. Then, the extraction control unit 103 calculates the zero ratio of the comparison terminal ID by dividing the number of zeros of the comparison terminal ID by the total number of zeros of the reference terminal ID, and determines whether the zero ratio of the comparison terminal ID is greater than or equal to a predetermined zero threshold (for example, "0.8").
[0048] In the above description, it is assumed that multiple pieces of category information exist. However, here, to facilitate understanding of the present invention, it is assumed that only four types of category information exist. In this case, the arrangement of the stay times for each piece of category information in the migration pattern 601 of the first terminal ID ("a01") is "10," "null," "null," and "20," while the arrangement of the stay times for each piece of category information in the migration pattern 602 of the second terminal ID ("a02") is "null," "15," "20," and "null." Therefore, the total number of zeros of the reference terminal ID in the migration pattern 601 of the first terminal ID is "2," and the number of zeros of the comparison terminal ID calculated by comparing the migration pattern 601 of the first terminal ID ("a01") with the migration pattern 602 of the second terminal ID ("a02") is "0," so the zero ratio of the comparison terminal ID is "0." In this case, the extraction control unit 103 determines that the zero ratio of the comparison terminal ID is less than the zero threshold ("0.8"), and determines that the arrangement of stay times for each category information is not similar between the navigation pattern 601 of the first terminal ID ("a01") and the navigation pattern 602 of the second terminal ID ("a02").
[0049] Then, the extraction control unit 103 refers to the migration pattern 603 of the third terminal ID (for example, "a03") that is next to the migration pattern 602 of the second terminal ID ("a02"), compares the migration pattern 601 of the first terminal ID ("a01") with the migration pattern 603 of the third terminal ID ("a03"), calculates the number of zeros of the comparison terminal ID (corresponding to the third terminal ID), calculates the zero ratio of the comparison terminal ID, and determines whether the zero ratio of the comparison terminal ID is greater than or equal to the zero threshold ("0.8").
[0050] Here, the sequence of stay times for each category information in the migration pattern 603 of the third terminal ID ("a03") is "5," "null," "null," and "5." The number of zeros for the comparison terminal ID calculated by comparing the migration pattern 601 of the first terminal ID ("a01") with the migration pattern 603 of the third terminal ID ("a03") is "2," and therefore the zero ratio of the comparison terminal ID is "1." In this case, the extraction control unit 103 determines that the zero ratio ("1") of the comparison terminal ID is equal to or greater than the zero threshold ("0.8"), and determines that the sequence of stay times for each category information is similar between the migration pattern 601 of the first terminal ID ("a01") and the migration pattern 603 of the third terminal ID ("a03"). This makes it possible to determine the similarity of the sequence of stay times for each category information in the migration pattern for each terminal ID and extract migration patterns of candidate terminal IDs.
[0051] Next, the extraction control unit 103 further extracts migration patterns of similar terminal IDs based on the similarity of the stay times for each category information in the migration patterns for each terminal ID. For example, the extraction control unit 103 calculates the number of category information whose stay times are not "null" in the migration pattern 601 of the reference first terminal ID ("a01") as the total number of non-zero values for the reference terminal ID. Next, as shown in FIG. 6B, the extraction control unit 103 compares the migration pattern 601 of the first terminal ID ("a01") with the migration pattern 603 of the candidate third terminal ID ("a03"), and calculates the number of category information in the same location where the difference between the stay time of the first terminal ID ("a01") and the stay time of the third terminal ID ("a03") is within a predetermined difference threshold (e.g., "5"), as the number of non-zero values for the comparison terminal ID. Then, the extraction control unit 103 calculates the non-zero ratio of the comparison terminal ID by dividing the number of non-zero values of the comparison terminal ID by the total number of non-zero values of the reference terminal ID, and determines whether the non-zero ratio of the comparison terminal ID is greater than or equal to a predetermined non-zero threshold value (for example, "0.8").
[0052] For example, the total number of non-zero values of the reference terminal ID in the migration pattern 601 of the first terminal ID ("a01") is "2," the difference value in the first category information is "5," and the difference value in the fourth category information is "15," so the number of non-zero values of the comparison terminal ID is "1," and the non-zero ratio of the comparison terminal ID is "0.5." In this case, the extraction control unit 103 determines that the non-zero ratio of the comparison terminal ID is less than the non-zero threshold value ("0.8"), and determines that the stay times for each category information are not similar between the migration pattern 601 of the first terminal ID ("a01") and the migration pattern 603 of the third terminal ID ("a03").
[0053] Then, the extraction control unit 103 refers to the navigation pattern 604 of the next fourth terminal ID (for example, "a04"), compares the navigation pattern 601 of the first terminal ID ("a01") with the navigation pattern 604 of the fourth terminal ID ("a04"), and determines whether the arrangement of stay times for each category information is similar between the navigation pattern 601 of the first terminal ID ("a01") and the navigation pattern 604 of the fourth terminal ID ("a04")
[0054] Here, the arrangement of the stay times for each category information in the navigation pattern 604 of the fourth terminal ID ("a04") is "15", "null", "null", "15", so the number of zeros in the comparison terminal ID is "2" and the zero ratio of the comparison terminal ID is "1", so the extraction control unit 103 determines that the zero ratio ("1") of the comparison terminal ID is greater than or equal to the zero threshold ("0.8") and determines that the arrangement of the stay times for each category information is similar in the navigation pattern 601 of the first terminal ID ("a01") and the navigation pattern 604 of the fourth terminal ID ("a04").
[0055] Next, the extraction control unit 103 compares the migration pattern 601 of the first terminal ID ("a01") with the migration pattern 604 of the fourth terminal ID ("a04"), and finds that the difference value in the first category information is "5" and the difference value in the fourth category information is also "5". Therefore, the number of non-zero values of the comparison terminal ID is "2", and the non-zero ratio of the comparison terminal ID is "1". In this case, the extraction control unit 103 determines that the non-zero ratio of the comparison terminal ID is equal to or greater than the non-zero threshold value ("0.8"), and determines that the stay times for each category information are similar between the migration pattern 601 of the first terminal ID ("a01") and the migration pattern 604 of the fourth terminal ID ("a04"). In this case, the extraction control unit 103 extracts the migration pattern 604 of the fourth terminal ID ("a04") as a migration pattern similar to the migration pattern 601 of the first terminal ID ("a01"), and extracts the similar migration pattern 601 of the first terminal ID ("a01") and the migration pattern 604 of the fourth terminal ID ("a04") as a first migration cluster (for example, "cluster A"). This makes it possible to extract migration patterns of a predetermined number of similar terminal IDs as a migration cluster.
[0056] The extraction control unit 103 compares the migration patterns for each terminal ID, and repeatedly determines the similarity of the arrangement of stay times for each category of information and the approximation of the stay times for each category of information, thereby extracting a predetermined number of migration clusters.
[0057] Here, as shown in Figure 6B, it is assumed that the extraction control unit 103 further extracts the migration pattern 602 of the second terminal ID ("a02") and the migration pattern 605 of the fifth terminal ID ("a05") as a second migration cluster (e.g., "cluster B").
[0058] Here, after the extraction control unit 103 extracts a predetermined number of wandering clusters, it may proceed directly to the next step, but in order to improve the accuracy of the linguistic meaning estimated from the wandering clusters, the extraction control unit 102 may determine whether the number of extracted wandering clusters is equal to or less than a predetermined cluster threshold (e.g., "2") after extracting all wandering clusters (Figure 2: S104).
[0059] As described above, the number of extracted migration clusters is "2," so the extraction control unit 102 determines that the number of migration clusters is equal to or less than the cluster threshold ("2") (FIG. 2: S104 YES). In this case, the extraction control unit 103 determines that the number of migration clusters obtained by clustering the characteristic migration patterns has converged, and ends the extraction process.
[0060] On the other hand, if the number of extracted migration clusters is "3," the extraction control unit 102 determines that the number of migration clusters exceeds the cluster threshold ("2") (FIG. 2: S104 NO). In this case, the extraction control unit 103 determines that the number of migration clusters obtained by clustering characteristic migration patterns is diffuse. Then, the extraction control unit 103 returns to S101 and collects the terminal ID, category information, and stay time (FIG. 2: S101), calculates the migration pattern for each terminal ID (FIG. 2: S102), and extracts the migration clusters (FIG. 2: S103) again to converge the number of migration clusters. Then, the extraction control unit 103 determines whether the number of extracted migration clusters is equal to or less than the predetermined cluster threshold ("2") (FIG. 2: S104). As a result, by increasing the number of migration patterns of the terminal ID to be determined and processing to converge the number of migration clusters, it is possible to extract migration clusters that collect characteristic migration patterns with high accuracy.
[0061] Now, when the extraction control unit 103 completes the extraction, the estimation control unit 104 of the server 13 estimates the linguistic meaning that indicates the characteristics of the extracted migration cluster based on the time proportion of stay time for each category information of the migration cluster of a predetermined number of terminal IDs belonging to the extracted migration cluster (Figure 2: S105).
[0062] Here, there is no particular limitation on the method of estimation performed by the estimation control unit 104. For example, as shown in Fig. 7A, the estimation control unit 104 refers to a first migration cluster (e.g., "cluster A") and calculates, for each piece of category information, the total time obtained by adding up all of the stay times of predetermined category information in the migration clusters of a predetermined number of terminal IDs belonging to the first migration cluster ("cluster A") as the total stay time of the category information. Next, the estimation control unit 104 calculates the total time obtained by adding up all of the total stay times of each piece of category information as the total time, and calculates the quotient obtained by dividing the total stay time of the predetermined category information by the total time as the time ratio of the category information.
[0063] For example, the estimation control unit 104 adds up the stay time "10" for the first category information in the migration pattern 601 of the first terminal ID ("a01") and the stay time "15" for the first category information in the migration pattern 604 of the fourth terminal ID ("a04") to calculate a total stay time of "25" for the first category information. The estimation control unit 104 also adds up the stay time "20" for the fourth category information in the migration pattern 601 of the first terminal ID ("a01") and the stay time "15" for the fourth category information in the migration pattern 604 of the fourth terminal ID ("a04") to calculate a total stay time of "35" for the fourth category information. In the first migration cluster ("cluster A"), the stay times for the second and third category information are "null," so the total stay time for the second and third category information is "0." Next, the estimation control unit 104 adds up the total stay time of "25" for the first category information and the total stay time of "35" for the fourth category information to calculate a total time of "60", divides the total stay time of "25" for the first category information by the total time of "60" to calculate a time proportion of "41.7%" (0.417) for the first category information, and divides the total stay time of "35" for the fourth category information by the total time of "60" to calculate a time proportion of "58.3%" (0.58.3) for the first category information. Note that the time proportions for the second and third category information are "0.0%".
[0064] Here, if the vertical axis (X-axis) represents the time ratio and the horizontal axis (Y-axis) represents the category information, and the calculated time ratio for each category information is plotted, it can be seen that the most time is spent on the fourth category information ("PC04"), followed by the first category information ("PC01"), as shown in Figure 7B. In other words, it can be estimated that in the first migratory cluster ("cluster A"), the fourth category information ("PC04") is of greatest interest, followed by the first category information ("PC01").
[0065] Next, the estimation control unit 104 assigns linguistic meanings that indicate the characteristics of the migration cluster using the time proportion of the stay time for each category information. Specifically, as shown in FIG. 8A, the estimation control unit 104 creates a first semantic layer in which category information and the time proportion of the category information are paired based on the time proportion of the stay time for each category information in a predetermined migration cluster ("cluster A"). Here, the pairs of category information and time proportion are arranged, for example, in ascending order of time proportion. The first semantic layer for the migration cluster ("cluster A") is {PC02: 0.0%}, {PC03: 0.0%}, {PC01: 41.7%}, {PC04: 58.3%}. Then, the estimation control unit 104 refers to the time proportion table.
[0066] In the time ratio table 800, as shown in FIG. 8A, the time ratio range 801 and the user's behavioral meaning 802 within the time ratio range 801 are stored in association with each other. The time ratio ranges 801 are arranged in ascending order of the size of the range 801. For example, when the time ratio range 801 is "0.0% < t ≤ 5.0%", the behavioral meaning 802 is set as "Stay a little", when the time ratio range 801 is "5.0% < t ≤ 25.0%", the behavioral meaning 802 is set as "View quickly", when the time ratio range 801 is "25.0% < t ≤ 50.0%", the behavioral meaning 802 is set as "Look around in a circle", and when the time ratio range 801 is "50.0% < t ≤ 75.0%", the behavioral meaning 802 is set as "Look around carefully". Thus, the longer the time ratio, the more the behavioral meaning is set to prolong the stay at that location. Note that when the time ratio is "0.0%", there is no relation at all, so the behavioral meaning 802 does not exist.
[0067] Now, as shown in FIG. 8B, the estimation control unit 104 collates with the time ratio range 801 of the time ratio table 800 that includes the time ratio of the created first-layer predetermined category information with semantic annotation, and associates the behavioral meaning 802 associated with the collated time ratio range 801 with the category information. The estimation control unit 104 changes the time ratio of the category information with semantic annotation at the first layer to the behavioral meaning 802 of the time ratio table 800, converts it to the second layer with semantic annotation that is a set of the category information and the behavioral meaning associated with the time ratio of the category information, and estimates the linguistic semantic annotation.
[0068] For example, when the inference control unit 104 refers to the time ratios of the first and second category information in the first semantic layer, since the time ratios of the first and second category information are "0.0%", even if the inference control unit 104 refers to the time ratio range 801 of the time ratio table 800, there is no matching time ratio range 801. Therefore, the inference control unit 104 refers to the time ratio of the next third category information. Since the time ratio of the third category information is "41.7%", the inference control unit 104 acquires the behavioral meaning 802 of "look around in a circle" associated with the time ratio range 801 of "25.0% < t ≤ 50.0%" that includes the time ratio "41.7%", and associates it with the third category information. Also, since the time ratio of the fourth category information is "58.3%", the inference control unit 104 acquires the behavioral meaning 802 of "look around carefully" associated with the time ratio range 801 of "50.0% < t ≤ 75.0%" that includes the time ratio "58.3%", and associates it with the fourth category information. In this way, the inference control unit 104 can convert to the second semantic layer in which a set of category information and the behavioral meaning associated with the category information is used.
[0069] Here, for example, if "PC03" of the third category information is "office" and "PC04" of the fourth category information is "sports", it is understood that the movement and stay of this roaming cluster ("cluster A") have life behaviors such as "look around in a circle" at the "office" and "look around carefully" at the "sports", or there is an interest in these. In this way, by the inference control unit 104 converting to the second semantic layer, it becomes possible to estimate the linguistic semantic meaning indicating the characteristics of the roaming cluster, and without depending on personal information, it is possible to accurately estimate what life behaviors and interest characteristics a specific group has while considering privacy.
[0070] Here, the movement and stay of a specific migratory cluster ("cluster A") makes it possible to clarify the basic purpose of the migratory cluster based on linguistic meaning, and therefore, for example, marketing and promotions targeting the migratory cluster can be carried out based on this basic purpose.
[0071] In the above description, the estimation control unit 104 uniquely calculates the time proportion for each category of information using a certain formula, but it may also be possible to weight each time proportion according to the content of the category information so that there are differences in the time proportion for each category of information. This makes it possible to more clearly estimate daily activities and interests when assigning linguistic meaning.
[0072] Incidentally, the more specific the category information, the more specific the linguistic meaning. Therefore, in order to make the category information more specific, two sets of main category information and subcategory information may be set in association with each other for the above-described device category table 400 and GPS category table 500. For example, as shown in FIG. 9A , the device category table 400 stores a device ID 401 ("abc"), main category information 402 ("PC01") of the location where the wireless device 11 with that device ID is installed, and subcategory information 900 (e.g., "SC01") of that location, in association with each other. Furthermore, the GPS category table 500 stores GPS information 501 ("X20, Y20"), main category information 502 ("PC02") of that location, and subcategory information 900 (e.g., "SC02") of that location, in association with each other. Here, as described above, the main category information 402, 502 is set to broad categories such as office, education, temples and shrines, leisure, amusement, sports, learning, culture, retail stores, eating out, clinics, transportation, etc., and the subcategory information 900 is set to specific categories such as retail store fields such as daily necessities, drug stores, convenience stores, coffee shops, etc.
[0073] Here, in S101, when collecting terminal IDs, category information, and stay times, subcategory information 900 is collected in addition to main category information 402, 502, and the calculation of the migration pattern for each terminal ID in S102 and the extraction of migration clusters in S103 are performed using the main category information 402, 502. Then, when calculating the time proportion of stay times for each category information of the migration clusters of a predetermined number of terminal IDs belonging to the migration clusters in S105, the subcategory information 900 is used as appropriate.
[0074] For example, as shown in FIG. 9B , for a first migration cluster (e.g., “cluster A”), the main category information is “PC01,” “PC02,” “PC03,” “PC04,” etc., and the time percentages of the main category information are “15.0%, 5.0%, 30.0%, 10.0%, etc. Here, if the main category information “PC03” and “PC04” are both category information related to retail stores, it is difficult to clarify the sales target that the first migration cluster is interested in even if linguistic meaning is assigned as is. Therefore, the estimation control unit 104 acquires each of the subcategory information 900 for multiple common main category information items that share a common field or purpose among the main category information items, calculates the sum of the time percentages of the common main category information items as the total time percentage, and calculates the quotient obtained by dividing the time percentage of the acquired subcategory information by the total time percentage as the time percentage of the acquired subcategory information.
[0075] For example, the subcategory information "SC03" of the main category information "PC03" and the subcategory information "SC04" of the main category information "PC04" are acquired, and the total time percentage is calculated as the sum of the time percentages of the common main category information, 40.0% (30.0% + 10.0%). The time percentage of the subcategory information "SC03" is calculated as 75.0% (30.0% / 40.0%), and the time percentage of the subcategory information "SC04" is calculated as 25.0% (10.0% / 40.0%). The estimation control unit 104 then reflects the subcategory information "SC03" with the highest time percentage in the first semantic hierarchy. In this example, the results are {PC02: 5.0%}, {PC01: 15.0%}, and {SC03: 75.0%}. Then, the inference control unit 104 converts the first semantic hierarchy into the second semantic hierarchy using the time proportion table 800, and assigns linguistic meaning to the first shopping cluster. Here, the second semantic hierarchy is {PC02: stay a little while}, {PC01: quickly appreciate}, {SC03: look around carefully}. Here, for example, if the subcategory information "SC03" is "drugstore," it can be seen that the first shopping cluster has a lifestyle behavior and interest of "looking around carefully" and shopping at the "drugstore."
[0076] Furthermore, if subcategory information is not used, for example, if the main category information "PC03" and "PC04" are "Retailer A" and "Retailer B," both of which are "retailers," then in main category information "PC03," the time percentage is "30.0%," so the behavioral meaning is "looked around," and in main category information "PC04," the time percentage is "10.0%," so the behavioral meaning is "take a quick look around." If linguistic meaning is assigned as is, "retailer A" will be converted to "take a quick look around," and "retailer B" will be converted to "take a quick look around," and it will be clear that the linguistic meaning of the first shopping cluster will be unclear.
[0077] Therefore, by adding subcategory information to main category information, it becomes possible to give clearer linguistic meaning to the wandering cluster.
[0078] An application example of the present invention will now be described. For example, wireless devices 11 are installed in various locations, and the terminal IDs (e.g., "USER ID 000001") of multiple user terminals 10 roaming around each location, category information (e.g., "PLACE CATEGORY N") for each location where the user terminal 10 has stayed, and the duration of time the user terminal 10 has stayed at each location are collected for each user terminal 10.
[0079] Then, nonnegative tensor decomposition was performed on a tensor with the collected multiple device IDs, the category information for the device IDs, and the contact dates and times for the category information as three axes. A migration pattern indicating the distribution of stay times calculated from the contact dates and times for each category information of the device ID was calculated for each device ID. As shown in Figure 10, the rows (X-axis) were set to device IDs, the columns (Y-axis) to category information, and the depth (Z-axis) to stay times, and the device IDs, category information, and stay times were treated as a three-dimensional tensor. Floor 0 was set to device IDs, the number of levels was set to n (many), and the level was set to the number of devices to be measured (number of fixed MAC addresses). Floor 1 was set to category information, the number of levels was set to 12, and the level was set to the number of installation locations. Floor 2 was set to stay times, the number of levels was set to 2,880, and the level was set to stay times. In this way, a migration pattern indicating the arrangement of stay times for each category information of the device IDs was calculated for each device ID.
[0080] For example, when the navigation pattern for terminal ID "USER ID 000001" is calculated, as shown in Figure 11, the navigation pattern for terminal ID "USER ID 000001" indicates an arrangement of stay times for each category of information, and multiple stay times for each category of information calculated from the contact date and time on the Z axis are arranged as "0.5," "10.0," "40.0," "60.0," etc., corresponding to the order of multiple category information corresponding to the Y axis: "PLACE0001," "PLACE0002," "PLACE0003," "PLACE0004," etc. Then, when the stay times calculated from the contact date and time on the Z axis are plotted for each category of information on the Y axis, the navigation pattern indicating the arrangement of stay times for each category of information is expressed as linear behavior information of the user of user terminal 10 with terminal ID "USER ID 000001."
[0081] Then, based on the similarity of the arrangement of stay times for each category information in the calculated migration patterns for each terminal ID and the approximation of stay times for each category information, migration patterns of a predetermined number of similar terminal IDs were extracted as migration clusters. For example, as shown in Figure 12, a predetermined number of migration clusters (e.g., "Cluster A" and "Cluster B") were extracted based on the similarity and approximation of the migration patterns for each terminal ID. Furthermore, the time proportion of stay times for each category information of the migration clusters of the predetermined number of terminal IDs belonging to the extracted migration clusters was calculated.
[0082] 13, the vertical axis (X-axis) represents the time proportion and the horizontal axis (Y-axis) represents the category information, and the calculated time proportions for each category information are shown. For the category information "Transportation...," "Culture...," "Retailers...," "Retailers...," "Retailers...," "Retailers...," and "Restaurants...," the time proportions for the migration cluster ("Cluster A") are "12.0," "18.2," "9.6," "20.8," "2.6," "30.9," etc. Here, the third to fifth category information are mutually related to "Retailers...," so they are treated as common main category information, and the time proportion of the common main category information "Retailers..." was used to calculate the total time proportion of the third to fifth category information, "33.0."
[0083] As shown in Figure 14, we generated a first semantic hierarchy for the first migration cluster ("top-level cluster") composed of main category information. The first semantic hierarchy was {Transportation...: 12.0%}, {Culture...: 18.2%}, {Stores...: 33.0%}, {Restaurants...: 30.9%}. Applying the time proportion table to convert this to the second semantic hierarchy resulted in the following: {Transportation...: Stay a little}, {Culture...: Quickly appreciate}, {Stores...: Look around}, {Restaurants...: Have a quick meal}. The behavioral meanings in the time proportion table were designed according to the content of the category information. For example, in the case of "Restaurants," "Look around" was changed to "Have a quick meal." This clearly clarifies the characteristics of daily activities and interests in the first migration cluster ("top-level cluster").
[0084] Furthermore, we generated a semantic first layer for the first migration cluster ("secondary cluster") using subcategory information. The semantic first layer was {Transportation...: 12.0%}, {Culture...: 18.2%}, {Drugs...: 26.0%}, {Restaurants...: 30.9%}. "Drugs" is the subcategory information of the main category information with the highest time percentage among the common main category information "Retailers...." Applying the time percentage table to this, we converted it into a semantic second layer, resulting in {Transportation...: Stay for a while}, {Culture...: Quickly appreciate}, {Drugs: Mainly shop}, {Restaurants...: Have a quick meal}. Again, for the subcategory information "Drugs," "Look around" was changed to "Mainly shop." This shows that the characteristics of daily activities and interests of the first migration cluster ("secondary cluster") have become more specific. By utilizing these linguistic meanings, it is possible to further estimate the attributes of migratory clusters (gender, age group, hobbies, tastes, etc.) and use this information for marketing and promotion.
[0085] Now, in the present invention, when the estimation control unit 104 estimates the linguistic meaning (Figure 2: S105), the linguistic meaning may be output as is, or the generation control unit 105 of the server 13 may input the estimated linguistic meaning to a predetermined context machine learning unit, thereby generating a context for the user of the user terminal 10 of the wandering cluster for which the linguistic meaning was estimated (Figure 2: S106).
[0086] Here, user context refers to the background, situation, scene, etc. of the user carrying the user terminal 10, and by generating user context, it is possible to specifically clarify the basic attributes, interests, and behavioral conditions that the user has as a person living in the world.
[0087] There are no particular limitations on the method by which the generation control unit 105 generates a context. For example, as shown in FIG. 15, when the estimation control unit 104 generates a first semantic layer ({Transportation...: 12.0%},...) and a second semantic layer ({Transportation...: Stay a little},...) for a first migration cluster ("top-level cluster") composed of main category information, the generation control unit 105 inputs the first semantic layer and the second semantic layer of the first migration cluster to the context machine learning unit 105a.
[0088] Here, the context machine learning unit 105a can be a computer, device, software, etc. that uses a machine learning technique to discover certain rules from certain data and realize inferences and predictions for unknown data based on those rules. Known techniques can be used for machine learning, and there are no particular limitations on the type of machine learning, but for example, neural networks, genetic algorithms, reinforcement learning, etc. can be used.
[0089] Here, for example, the context machine learning unit 105a is configured according to the attributes and numerical values of the estimated linguistic meaning, and the generation control unit 105 inputs the attributes and numerical values of the first semantic layer ({traffic...:12.0%}, . . .) of the first migration cluster into the context machine learning unit 105a, thereby causing the context machine learning unit 105a to generate a graph or table according to the attributes and numerical values as the user's context. Here, there is no particular limitation on the type of graph, and examples include pie charts, bar graphs, and line graphs, as shown in FIG. 15. There is also no particular limitation on the type of table, and examples include matrices and maps.
[0090] For example, if the linguistic meaning of the migratory cluster is data related to age or gender, it will become a graph of basic attributes showing age and gender; if the linguistic meaning of the migratory cluster is data related to life stage, it will become a graph of interests showing life stage; and if the linguistic meaning of the migratory cluster is data related to place of residence or place of work, it will become a graph of behavioral conditions showing place of residence or place of work.
[0091] For example, if the first semantic hierarchy of the first migration cluster is {Transportation...: 12.0%}, {Culture...: 18.2%}, {Stores...: 33.0%}, {Restaurants...: 30.9%}, etc., the user's context will be a graph of interests and behavioral conditions. Also, if the second semantic hierarchy of the first migration cluster is {Transportation...: Stay for a while}, {Culture...: Quickly appreciate}, {Stores...: Look around}, {Restaurants...: Have a quick meal}, etc., the user's context will be a graph of basic attributes, interests, and behavioral conditions. Generating user context in this way makes it possible to visualize and estimate user attributes (gender, age group, hobbies, tastes, etc.) more specifically.
[0092] Incidentally, when the generation control unit 105 generates a user context using the context machine learning unit 105a, it is preferable that the context machine learning unit 105a utilize linguistic meanings input in the past. For example, the context machine learning unit 105a generates a graph or a table using attributes and numerical values of the input linguistic meaning and attributes and numerical values of linguistic meanings input in the past related to this attribute. This allows the user context to be generated taking into account past linguistic meanings other than the input linguistic meaning, thereby making it possible to generate a more appropriate user context. The machine learning unit 105a can use past linguistic meanings to add or supplement them.
[0093] Furthermore, in S106, when the generation control unit 105 generates the user's context, the generated user's context may be input to another machine learning unit to generate language or an image that indicates the user's context.
[0094] Here, there is no particular limitation on the method by which the generation control unit 105 generates language or images, but for example, as shown in FIG. 16, the generation control unit 105 inputs the generated user context to another machine learning unit 105b.
[0095] Here, there are no particular limitations on the other machine learning unit 105b, and it may have the same configuration as the above-mentioned context machine learning unit 105a, or may be configured as an external machine learning unit different from the server 13.
[0096] For example, if the other machine learning unit 105b is an external machine learning unit specialized in language, examples of the machine learning unit include ChatGPT, GPT-3.5, GPT-4 from OpenAI (registered trademark), LLaMa from Meta AI (registered trademark), etc. In this case, if the user's context consists of five types of topics, the other machine learning unit 105b outputs a language for each topic. For example, "Topic 0: Characteristic categories: Childcare, education, academics | Travel, accommodation These are family households that often have high school-aged children, enjoy traveling, and although they are nuclear families, the number of household members is estimated to be slightly above average, and they can be said to be households that have contact with relatives, etc. Topic 1: Characteristic categories: Shopping | Leisure, sports | Hobbies, entertainment These are households that enjoy shopping in a wide range of categories, such as electronics retailers, bookstores, and household goods, and are relatively interested in facilities that can be enjoyed alone, such as karaoke boxes, game centers, and gambling. They are thought to be a group that is single and includes a relatively large number of young people. Topic 2: Characteristic categories: Public transportation, automobiles | Public facilities | Childcare, education, academics These are households that frequently use automobiles and include many households with preschool children. They also frequently use public facilities such as parks, and have distinctive ways of spending weekends with their family or children. Topic 3: Characteristic categories: Leisure, sports | Medical, health, nursing care, travel, accommodation" They have a particularly high interest in golf, and are also extremely interested in nursing homes, massage, and chiropractic. They are relatively middle-aged and older, often travel for business, and are presumed to be from families with parents who require care. Topic 4: Characteristic Categories: Food & Drink | Beauty & Fashion This demographic is highly inclined toward eating out, and has a strong interest in beauty and fashion, with a particular interest in beauty salons and hairdressing salons. On the other hand, they have little interest in hobby and entertainment-related facilities, so they are a demographic that spends a lot of time on self-improvement, which is related to external expression. Language is output for each topic, as shown above. This makes it possible to verbalize the user's context, express the user's (visitor's) interests in concrete language, and deepen understanding of the user.
[0097] Furthermore, for example, if the other machine learning unit 105b is an external machine learning unit specialized in images, an example of such a machine learning unit is OpenAI (registered trademark)'s DALL-E3. Here, if a user's context is composed of five topics, the other machine learning unit 105b outputs an image for each topic. For example, it outputs an image for each topic, such as "Topic 0: Minimal Luxury, Topic 1: Minimal Talent, Topic 2: Minimal Family, Topic 3: Luxury Executive, Topic 4: Luxury Talent." This makes it possible to visualize the user's context and visualize a representative profile of the user, thereby deepening understanding of the user.
[0098] In this way, the present invention makes it possible to accurately estimate the characteristics of daily activities and interests of a specific group without relying on personal information and while taking privacy into consideration. Furthermore, the present invention makes it possible to convert the estimated linguistic meaning into a more understandable user context. Furthermore, the present invention makes it possible to deepen understanding of the user by converting the user context into language or images.
[0099] In the embodiment of the present invention, the migration estimation system 1 is configured to include each control unit. However, the program that realizes each control unit may be stored in a storage medium and the storage medium may be provided. In this configuration, the program is read by a device, and the device realizes each control unit. In this case, the program itself read from the recording medium achieves the effects of the present invention. Furthermore, it is also possible to provide a method for storing the steps executed by each control unit on a hard disk. [Industrial Applicability]
[0100] As described above, the migration estimation system and migration estimation method of the present invention are useful for utilizing the behavioral histories of multiple user terminals, and are effective as a migration estimation system and migration estimation method that can accurately estimate what kind of lifestyle behaviors and interest characteristics a specific group has, without relying on personal information and taking privacy into consideration. [Explanation of symbols]
[0101] 1. Migration estimation system 10 sensors 11 Server 12 Network 101 Collection control unit 102 Calculation control unit 103 Extraction control section 104 Estimation control unit
Claims
1. a collection control unit that collects, for each user terminal, the terminal IDs of a plurality of user terminals roaming around at each location, category information for each location where the user terminal has stayed, and the contact date and time when the user terminal has come into contact with the location; a calculation control unit that performs non-negative tensor decomposition on a tensor having three axes of the collected multiple terminal IDs, category information of the terminal IDs, and contact dates and times of the category information, and calculates, for each terminal ID, a migration pattern that indicates a distribution of stay times calculated from contact dates and times for each category information of the terminal IDs; an extraction control unit that extracts similar migration patterns of a predetermined number of terminal IDs as a migration cluster based on the similarity of the arrangement of the stay times for each category information in the calculated migration patterns for each terminal ID and the approximation of the stay times for each category information; an estimation control unit that estimates linguistic meanings that indicate characteristics of the extracted migration cluster based on a time ratio of stay time for each category information of the migration cluster of a predetermined number of terminal IDs that belong to the extracted migration cluster; A migration estimation system comprising:
2. When all migratory clusters are extracted, the extraction control unit determines whether the number of the extracted migratory clusters is equal to or less than a predetermined cluster threshold; If the result of the determination is that the number of migratory clusters is equal to or less than the cluster threshold, the extraction control unit terminates the extraction process, If the result of the determination is that the number of the migration clusters exceeds the cluster threshold, the extraction control unit causes the collection of terminal IDs, category information, and stay times, the calculation of the migration pattern for each terminal ID, and the extraction of the migration clusters to be performed again, thereby converging the number of the migration clusters. The migration estimation system according to claim 1 .
3. The estimation control unit creates a first semantic hierarchy in which category information and the time proportion of the category information are arranged in pairs based on the time proportion of the stay time for each piece of category information, and uses a time proportion table in which ranges of time proportions are stored in association with behavioral meanings of users within the ranges of time proportions to change the time proportions of the category information in the first semantic hierarchy to behavioral meanings 802 in the time proportion table 800, converting the second semantic hierarchy into pairs of category information and the behavioral meanings associated with the time proportions of the category information, and estimating the linguistic meaning. The migration estimation system according to claim 1 .
4. a generation control unit that generates a context of a user of a user terminal of a migration cluster in which the linguistic meaning is estimated by inputting the estimated linguistic meaning to a predetermined context machine learning unit; Further comprising: The migration estimation system according to claim 1 .
5. a collection control step of collecting, for each user terminal, the terminal IDs of a plurality of user terminals roaming around each location, category information for each location where the user terminal has stayed, and the contact date and time when the user terminal has come into contact with the location; a calculation control process for performing non-negative tensor decomposition on a tensor having three axes of the collected multiple terminal IDs, category information of the terminal IDs, and contact dates and times of the category information, and calculating, for each terminal ID, a migration pattern indicating a distribution of stay times calculated from contact dates and times for each category information of the terminal IDs; an extraction control step of extracting similar migration patterns of a predetermined number of terminal IDs as a migration cluster based on the similarity of the arrangement of the stay times for each category information in the calculated migration patterns for each terminal ID and the approximation of the stay times for each category information; an estimation control step of estimating linguistic meanings that indicate characteristics of the extracted migration clusters based on the proportion of stay times for each category information of the migration clusters of a predetermined number of terminal IDs that belong to the extracted migration clusters; A migration estimation method for a migration estimation system comprising:
Citation Information
Patent Citations
Marketing system and marketing method
JP2016004336A
Information distribution device and information distribution method
JP2019175123A
Information processing apparatus, information processing method, and information processing program
JP2022144212A
Evaluation device and evaluation method
WO2018061623A1
Prediction device and prediction method
WO2018131214A1