Modeling user propensity based on location data
By integrating time-varying and static behavioral attributes into the model training, and using machine learning techniques to extract user trajectories from location data, the bias problem of self-reported data is solved, and more accurate and flexible user tendency prediction is achieved.
Patent Information
- Application Number
- CN202480043612.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-28
- Filing Date
- 2024-05-15
- Publication Date
- 2026-02-27
AI Technical Summary
Existing location-based user propensity modeling methods rely on self-reported data, which is biased and inaccurate. They also struggle to effectively integrate time-varying and static behavioral attributes and lack model flexibility, making them difficult to adapt to different applications.
By training a model to integrate time-varying and static behavioral attributes, machine learning techniques are used to extract user trajectories from location data, infer activity data, and convert it into behavioral attributes. These attributes are then combined with time-varying and static attributes to quantify user tendencies, and model parameters are adjusted to adapt to different behavioral predictions.
It achieves more accurate and robust user propensity prediction, can quickly adjust the model to adapt to predictions of different behaviors or experiences, and does not require individual identification information, thus improving the accuracy and flexibility of prediction.
Smart Images

Figure CN121586908A_ABST
Abstract
Description
Cross Reference to Related Applications
[0001] This application claims priority to and the benefit of pending U.S. Provisional Application No. 63 / 523,874, filed June 28, 2023, for all subject matter common to the two applications. The disclosure of the provisional application is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] The present invention relates to the field of data analysis and machine learning suitable for behavioral modeling. In particular, the present invention relates to modeling user predispositions based on location data. BACKGROUND
[0003] In recent years, the proliferation of mobile devices and the rapid development of location-based services have led to an exponential growth in the availability of location data. These data, which include information about users’ geographic locations and movements, open new possibilities for understanding and predicting human behavior. One area of interest is modeling user predispositions based on location data, which can provide valuable insights into users’ preferences, habits, and potential future behaviors or experiences.
[0004] Traditional methods for modeling user predispositions often rely on self-reported data, such as surveys and questionnaires. However, these methods can be subject to various biases and inaccuracies, as users may not be able to accurately recall or report their past behaviors and experiences. Moreover, self-reported data may not capture all factors that influence user predispositions, such as environmental and contextual factors that can be derived from location data.
[0005] Machine learning techniques have increasingly been used to model user predispositions based on location data. These techniques can automatically learn patterns and relationships in the data, resulting in more accurate and robust models. However, the process of training machine learning models can be complex and computationally intensive, requiring careful selection and tuning of model parameters to achieve optimal performance. Moreover, current approaches can result in models that are inflexible, making it difficult to re-tune or adapt them for other purposes.
[0006] Furthermore, integrating time-varying and static behavioral attributes in the modeling process can be challenging. Time-varying attributes, such as the frequency and duration of visits to specific locations, change over time and can be influenced by various factors, such as users’ schedules and preferences. On the other hand, static attributes represent more stable characteristics of users, such as their demographic information or home address. Combining these different types of attributes in a meaningful and effective way is crucial for developing accurate models of user predispositions. SUMMARY
[0007] In view of these challenges, there is a need for a system for determining user predispositions based on location data that can effectively integrate time-varying and static behavioral attributes while harnessing the power of machine learning techniques to automatically learn patterns and relationships in the data. The present invention enables the development of more accurate and robust systems to predict user predispositions to specific behaviors or experiences / conditions with higher precision and reliability while retaining the ability to adjust or re-adjust the models in a more efficient manner as needed to predict different behaviors, experiences, or conditions.
[0008] The present invention addresses this need by providing a method for determining user predispositions based on location data, the method comprising the steps of training at least one model to determine user predispositions to specific behaviors or experiences / conditions, and predicting individual user predispositions to specific behaviors or experiences / conditions. The method involves obtaining location data for a plurality of users, formatting the location data into trajectories, inferring activity data from the location data, converting the activity data and location data into time-varying and static behavioral attributes, and determining predispositions to specific behaviors or experiences / conditions by modeling the time-varying attributes, combining the modeled time-varying attributes with static attributes, and assigning quantified predispositions. The method further comprises adjusting model parameters based on results, obtaining trained models, and using the trained models to predict individual user predispositions based on their location data.
[0009] Location data can be collected from various sources such as GPS-enabled devices, Wi-Fi access points, and cell phone signal towers. This data can be used to create trajectories, which are sequences of geographic locations and timestamps representing a user's movement over time. By analyzing these trajectories, various types of activity data can be inferred, such as the type of places visited, the duration of visits, and the frequency of visiting specific locations.
[0010] The analysis of location data and activity data can reveal patterns and trends in user behavior, which can be used to model their predispositions to specific behaviors or experiences / conditions. These predispositions can be quantified and used to predict the likelihood of users engaging in certain activities or experiencing specific conditions in the future. Such predictions are valuable for a variety of applications, including personalized recommendations, targeted advertising, and public health interventions.
[0011] According to certain embodiments, the technology described herein relates to a method of modeling user predisposition based on location data, the method comprising: I. training at least one model to determine a user's predisposition to a particular behavior or experience / condition, the training comprising: A) obtaining location data for a plurality of users; B) formatting the location data into one or more trajectories for each of the plurality of users; C) inferring activity data from the location data; D) converting the activity data and the location data into time-varying and static behavioral attributes for each of the plurality of users; E) determining the predisposition to the particular behavior or experience / condition, including: for each user: modeling the time-varying attributes; combining the modeled time-varying attributes with the static attributes; and assigning a quantified predisposition for the particular behavior or experience / condition; and adjusting parameters based on results, thereby resulting in a trained model; II. predicting an individual user's predisposition to the particular behavior or experience / condition, comprising: A) obtaining location data for the individual user; B) inputting the location data into the at least one trained model resulting from the training step; and C) receiving from the trained model an assigned quantified predisposition of the individual user to the particular behavior or experience / condition.
[0012] According to certain aspects, the predisposition to the particular behavior or experience / condition is purchase intention; the predisposition to the particular behavior or experience / condition is hospitalization risk; the predisposition to the particular behavior or experience / condition is work engagement; the predisposition to the particular behavior or experience / condition is work change; the predisposition to the particular behavior or experience / condition is travel intention; the predisposition to the particular behavior or experience / condition is residential relocation intention; the predisposition to the particular behavior or experience / condition is healthcare risk; and / or the predisposition to the particular behavior or experience / condition is associated with a window of opportunity.
[0013] According to certain aspects, the location data is geospatial data of the user's hardware device over a period of time. According to certain aspects, the location data of the user's device is provided by a location provider / supplier. According to certain aspects, the location data of the user is identified without using one or more of: personally identifiable information (PII), demographic information, and socio-economic information about the user.
[0014] According to certain aspects, inferring activity data from the location data further comprises: i) mapping the location data to place types; ii) grouping the place types into activity groups based on functionality of the place types; and iii) converting the one or more trajectories of the user into one or more activity trajectories of the user using the activity groups. In some such aspects, the activity groups comprise one or more selected from the group of: hospital, health, essential shopping, fitness, public transportation, own transportation, religious, entertainment, travel, individual care, leisure shopping, unhealthy activity, restaurant, home, and work.
[0015] According to certain aspects, converting the activity data and the location data into behavior attributes for each of the plurality of users includes defining time-varying attributes; and defining static attributes.
[0016] In some such aspects, defining the time-varying attributes includes determining lifestyle attributes; determining activity attributes; and determining mobility attributes. In further such aspects, determining the lifestyle attributes includes using an unsupervised learning model that can identify similarities in activity patterns between users. In still further aspects, the unsupervised learning model includes a clustering and dimensionality reduction model, a hidden Markov model, or an LDA and topic model. The unsupervised learning model including an LDA and topic model can also include an author topic model (ATM), where users are authors, activities are words, activity periods are documents, and lifestyles are topics.
[0017] In other such aspects, defining the static attributes includes determining reachability attributes; and determining socio-demographic attributes.
[0018] In other such aspects, defining the time-varying attributes includes embedding the lifestyle attributes into a fixed-dimensional continuous vector representation with learnable weight parameters and tunable model hyperparameters; concatenating the embedded lifestyle attributes with attributes for activities and mobility; and converting the concatenated embedded lifestyle attributes and time-varying numerical attributes into a hidden representation with shared learnable weights and bias parameters and tunable model hyperparameters. In still further other aspects, converting the concatenated embedded lifestyle attributes and time-varying numerical attributes into a hidden representation includes using a supervised machine learning model that models spatial and temporal information from user trajectories. In still further other aspects, the supervised machine learning model includes a non-deep learning regression or classification model. The non-deep learning regression or classification model can include a decision tree-based model, a random forest-based model, or a gradient boosting model. Alternatively, the supervised machine learning model includes a deep learning model or a deep neural network-based model. The deep neural network-based model can include at least one of a convolutional neural network (CNN), a recurrent neural network (RNN), a long short-term memory (LSTM) model, a radial basis function network (RBFN), a Transformer-based attention model. The deep neural network-based model can also include a convolution long short-term memory (CLSTM) model that takes time-varying attributes as input.
[0019] In further such aspects, combining the modeled time-varying attributes with the static attributes includes embedding categorical static attributes into a fixed-dimensional continuous vector representation with learnable weight parameters and tunable model hyperparameters; and concatenating the embedded categorical static attributes, numerical static attributes, and hidden representations of the modeled time-varying attributes.
[0020] According to certain aspects, the model is adjusted via cross-validation.
[0021] According to certain aspects, the method further comprises outputting the assigned quantified propensity for a particular behavior or experience / condition of a particular user. In some such aspects, the outputting comprises providing a graphical representation of the assigned quantified propensity for the particular behavior or experience / condition of the particular user. In other such aspects, the outputting comprises providing a web service API that provides access to the assigned quantified propensity for the particular behavior or experience / condition of the particular user.
[0022] According to certain embodiments, a system for modeling user propensity based on location data is provided. The system comprises a data collection module configured to obtain location data for a plurality of users, a model training module configured to train a model using the location data to determine a propensity of the users for a particular behavior, and a prediction module configured to use the trained model and location data for an individual user to predict a propensity of the individual user for the particular behavior.
[0023] According to certain aspects, the model training module is further configured to infer activity data from the location data, convert the activity data and the location data into behavior attributes, and determine the propensity of each user for the particular behavior.
[0024] According to certain aspects, the prediction module is further configured to input the location data into the trained model and receive from the trained model an assigned quantified propensity for the particular behavior of the particular user.
[0025] The present invention provides a number of advantages over previous approaches. The present invention provides accurate predictions for a particular user, rather than predicting a type or group of users based on population demographics. The present invention provides a flexible model that can be adjusted or re-adjusted in a more efficient manner as needed to predict different behaviors, experiences or conditions. That is, iterative training of the model allows for the system to be adjusted to provide accurate predictions for a particular behavior, experience or condition, and then quickly and conveniently re-adjusted and re-trained to predict different behaviors, experiences and conditions, while using the same source of location data (broadly) as input. Furthermore, the present invention can also infer activities of users using publicly available location data, and thus does not require individual identifying information (PII) about the users to make predictions. BRIEF DESCRIPTION OF DRAWINGS
[0026] These and other features of the present invention will be more fully understood and appreciated by careful study of the following detailed description, taken together with the drawings, in which:
[0027] Figure 1 An illustrative environment in which the present invention is used in accordance with an embodiment of the present invention is depicted.
[0028] Figure 2 is an illustrative flow chart depicting the process of modeling user propensity based on location data according to embodiments of the present application.
[0029] Figure 3 is an illustrative flow chart depicting the process involved in inferring activity data according to embodiments of the present application.
[0030] Figure 4 is an illustrative flow chart depicting the process involved in converting activity data and location data into behavioral attributes according to embodiments of the present application.
[0031] Figure 5 is an illustrative flow chart depicting the process involved in determining propensity and adjustment according to embodiments of the present application.
[0032] Figure 6 is an illustrative dashboard showing user propensity in a region in a graphical manner according to embodiments of the present application.
[0033] Figure 7 is an illustrative dashboard showing user propensity in a specific region in a graphical manner according to embodiments of the present application.
[0034] Figure 8 is an illustrative dashboard showing real-time propensity modeling of a user according to embodiments of the present application.
[0035] Figure 9 is an illustrative dashboard showing potential uses of such real-time according to embodiments of the present application.
[0036] Figure 10 is a schematic diagram of a high-level architecture for implementing the processes according to embodiments of the present application.
[0037] Figure 11 is an illustrative probabilistic graphical model of an author topic model using board representation according to embodiments of the present application.
[0038] Figure 12 is an illustrative architecture diagram of a temporal deep learning model according to embodiments of the present application.
[0039] Figure 13 is an illustrative graphical model of a CLSTM and a concatenation layer according to embodiments of the present application.
[0040] Figure 14 shows an illustrative heat map of user activity according to embodiments of the present application. DETAILED DESCRIPTION
[0041] One illustrative embodiment of the present application relates to a technique that utilizes large amounts of crowd-sourced location data to identify lifestyles and quantify their association with future behaviors or experiences. When analyzing large amounts of location data from hundreds of thousands to millions of users, some key findings emerge. Individual lifestyle choices are more predictive of future behaviors and experiences than other factors typically thought to have a significant impact, such as accessibility of healthcare, socio-economic factors, age, hobbies, etc. For example, people with busy and varied work schedules and limited gym visits are 2.01 times more likely to be hospitalized within a year compared to the population average. The technique is able to predict a person's propensity for certain behaviors or experiences. It utilizes crowd-sourced location data and information indicative of a large number of users' location histories, and then correlates the data with various behaviors and experiences that those users have had. The correlations are then used to predict a particular person's future behaviors or experiences based on their real-time and / or recent location history. The applications of this technique are very broad. It can be applied to a variety of behaviors or experiences, such as a person's intent to purchase a product, the need for hospitalization, experiencing a job change, starting a particular travel adventure, moving, needing a particular healthcare, and other behaviors or experiences.
[0042] Figures 1 to 14 One or more example embodiments of a method and system for modeling user propensities based on location data according to the present application are shown, wherein like parts are denoted by like reference numerals throughout. While the present application will be described with reference to one or more example embodiments illustrated in the drawings, it is to be understood that numerous alternative forms can be devised for carrying out the present application. Those skilled in the art will further appreciate that the specific structural and functional details disclosed herein are merely representative in order to fully enable this present application and it will be appreciated that one of ordinary skill in the art will be able to devise numerous alternative arrangements that are within the scope of the present application.
[0043] Figure 1An illustrative environment 100 in which the present application is employed is depicted. Environment 100 includes a plurality of users 102, a server or cloud service 104, and a business 106. Each user 102 has a hardware or mobile device 108, such as a cell phone, connected to a network or the internet 110. Server or cloud service 104 is a system having a data collection module 112 configured to obtain location data for a plurality of users, a model training module 114 configured for training a model to determine a user's propensity for a particular behavior using the location data, and a prediction module 116 configured to use the trained model and location data for an individual user to predict the individual user's propensity for a particular behavior, prediction module 116 receives location data for a plurality of users 102 based on their respective hardware or mobile devices 108, which can be located via GPS or network connection. In some embodiments, the location data is collected by a vendor 118 that collects and sells location data and provided to server or cloud service 104. Using the received location data for users 102, server or cloud service 104 models the behavior of the users and provides a prediction of the behavior of users 102 to business 106. The business can be any type of entity whose business model relies on customers, such as a retail business, a hospital, etc. The business can then tailor advertising, marketing, and the experience provided based on this prediction.
[0044] Figure 2 A flowchart 200 is shown, which represents the various steps involved in a method of modeling user propensity based on location data as claimed in the claims of the present application. The flowchart consists of the following elements and their respective functions:
[0045] Training the model (step 202) involves training at least one model to determine a user's propensity for a particular behavior. The training process includes obtaining location data for a plurality of users (step 206), formatting the data into trajectories (step 208), inferring activity data (step 210), converting the activity data and location data into behavior attributes (step 212), and determining a propensity for a particular behavior or experience / condition (step 214).
[0046] In certain embodiments, the location data for a user is raw geospatial data collected from a user's hardware device, such as mobile device 108, over a period of time. In some such embodiments, the location data is provided by a location provider / vendor 118. In some embodiments, the location data is de-identified without using personally identifiable information (PII), demographic information, or socio-economic information about the user 102.
[0047] Formatting location data (step 208) involves processing the collected location data into one or more trajectories for each user. In certain embodiments, a trajectory can be a mathematical representation. For example, an individual's trajectory may be defined as a set of tuples ordered in time where is a location, where and are coordinates of a geographic location, and is a corresponding timestamp. Here, an individual corresponds to a user 102.
[0048] Inferring activity data (step 210) involves mapping geospatial location data to place types, grouping place types into activity groups, and converting trajectories into activity trajectories for each user. Figure 3 A depiction of the steps involved in inferring activity data (step 210) is shown. It includes mapping geospatial location data to place types (step 300), grouping place types into activity groups based on the functionality of the place types (step 302), and converting one or more trajectories of a user into one or more activity trajectories of the user using the activity groups (step 304).
[0049] Mapping location data to place types (step 300) can involve mapping the collected geospatial location data to specific place types, such as a hospital, a restaurant, or a shopping mall.
[0050] Grouping place types into activity groups (step 302) can involve grouping place types into activity groups based on the functionality of the place types. Examples of activity groups include: hospital, health, essential shopping, fitness, public transportation, own transportation, religion, entertainment, travel, individual care, leisure shopping, unhealthy activity, restaurant, home, and work.
[0051] Converting one or more trajectories of a user into one or more activity trajectories of the user using the activity groups (step 304) involves mathematical representations. For example, an individual's activity trajectory may be defined as mapping to activities that exhibit behavior, consumption, leisure patterns. is a set of tuples ordered in time where is an activity inferred by the individual from the locations closest to and and is a coarser timestamp of . Further, is represented as all time activities during a period of the ensemble. This step allows a more detailed analysis of the user behavior based on their activities and not only their location data.
[0052] As Figure 4 depicts the steps involved in converting the activity data and location data into behavior attributes for each of the users in the ensemble (step 212). It involves defining time-varying attributes (step 400) and defining static attributes (step 402).
[0053] Defining time-varying attributes (step 400) can also include determining lifestyle attributes (step 404), determining activity attributes (step 406), and determining mobility attributes (step 408).
[0054] Lifestyle attributes can be determined using an unsupervised learning model. This model identifies similarities in activity patterns between users. In certain embodiments, it can be a clustering and dimensionality reduction model, a Hidden Markov Model, or an LDA and topic model. In some particular embodiments, the LDA and topic model includes an Author Topic Model (ATM) where users are authors, activities are words, activity periods are documents, and lifestyles are topics.
[0055] Mobility attributes represent the mobility patterns of users based on their location data.
[0056] Defining static attributes (step 402) can also include determining reachability attributes (step 410) and determining socio-demographic attributes (step 412).
[0057] Figure 5 depicts the steps involved in determining the propensity for a particular behavior or experience / condition and making adjustments (step 214). For each user 102, it involves modeling time-varying attributes (step 500), combining the modeled time-varying attributes with static attributes (step 502), assigning a quantitative propensity for a particular behavior or experience / condition (step 504); and adjusting parameters based on the results, resulting in a trained model.
[0058] In certain embodiments, modeling time-varying attributes (step 500) involves embedding lifestyle attributes into a fixed-dimension continuous vector representation with learnable weight parameters and tunable model hyperparameters (step 508), concatenating the embedded lifestyle attributes with attributes for activities and mobility (step 510), and converting the concatenated embedded lifestyle attributes and time-varying numerical attributes into a hidden representation with shared learnable weights and bias parameters and tunable model hyperparameters (step 512).
[0059] In some such embodiments, converting the concatenated embedded time-varying categorical attributes and time-varying numerical attributes into a hidden representation (step 512) includes using a supervised machine learning model that models from spatial and temporal information of user trajectories. In particular embodiments, the supervised machine learning model includes a non-deep learning regression or classification model, such as a decision tree-based model, a random forest-based model, or a gradient boosting model. In other embodiments, the supervised machine learning model includes a deep learning model. In still other embodiments, the supervised machine learning model includes one or more of: a deep neural network-based model, such as a convolutional neural network (CNN), a recurrent neural network (RNN), a long short-term memory (LSTM) model, a radial basis function network (RBFN), and a Transformer-based attention model. In one particular embodiment, the deep neural network-based model includes a convolutional long short-term memory (CLSTM) model that takes time-varying attributes as input.
[0060] In certain embodiments, combining the modeled time-varying attributes with static attributes (step 502) involves: embedding categorical static attributes into a fixed-dimension continuous vector representation with learnable weight parameters and tunable model hyperparameters (step 514), and concatenating the embedded hidden representations of the embedded categorical static attributes, numerical static attributes, and modeled time-varying attributes (step 516).
[0061] Assigning a quantified propensity (step 504) involves making a probabilistic determination of a particular behavior or experience / condition. The propensity of the particular behavior or experience / condition can be any number of behaviors, experiences, or conditions. For example, the propensity of the particular behavior or experience / condition is a purchase intention, a hospitalization risk, a work participation, a job change, a travel intention, a residential relocation intention, or a healthcare risk.
[0062] In certain embodiments, the propensity of the particular behavior or experience / condition is associated with a window of opportunity, such as a number of months, weeks, days, or hours.
[0063] Adjusting parameters (step 506) allows the model to be tuned so that the assigned quantified propensity (step 504) reflects the desired behavior or experience / condition to be predicted. In certain embodiments, the model is tuned via cross-validation.
[0064] Referring back to Figure 2 Once the model has been trained (step 202), the trained model can be used to predict a propensity of an individual user for a particular behavior or experience / condition (step 204). This step includes: obtaining location data of the individual user (step 216), inputting the location data into the trained model (step 218), and receiving from the trained model an assigned quantified propensity of the user for the particular behavior or experience / condition (step 220).
[0065] In some embodiments, the preference is then output (step 222). In some such embodiments, the output involves providing a graphical representation of an assigned quantitative preference for a particular behavior or experience / condition of a particular user. In other embodiments, the output involves providing a web service API that provides access to the assigned quantitative preference for a particular behavior or experience / condition of a particular user.
[0066] Figures 6 to 9 Various examples of the graphical outputs that can be provided are depicted. Figure 6 It is an illustrative dashboard that graphically displays user preferences within a region. Figure 7 It is an explanatory dashboard, which has been zoomed in to graphically show user preferences in more specific areas. Figure 8 It is an illustrative dashboard that graphically displays real-time user preference modeling. Figure 9 This is an illustrative dashboard demonstrating such a potential real-time use according to an embodiment of the invention.
[0067] Any suitable and specially configured electronic or computing device may be used to implement the functions of this invention, including the hardware or mobile device 108 described herein, the server or cloud service 104 (including modules 112, 114, 116), the enterprise 106, and the supplier 118. Figure 10 This document describes an illustrative example of such an electronic or computing device 1000. The computing device 1000 is merely an illustrative example of a suitable computing environment and in no way limits the scope of the invention. Figure 10 As shown, "computing device" can include "workstation," "server," "laptop," "desktop," "device," "smart device," "tablet," "smartphone," "ECR," or other specially configured computing devices, as understood by those skilled in the art. Since computing device 1000 is depicted for illustrative purposes, embodiments of the invention can be implemented using any number of computing devices 1000 in any number of different ways. Therefore, those skilled in the art will understand that embodiments of the invention are not limited to a single computing device 1000, nor to a single type of implementation or configuration of the example computing device 1000.
[0068] The computing device 1000 may include a bus or network 1010 that may be directly or indirectly coupled to one or more of the following illustrative components: memory 1012, one or more processors 1014, one or more presentation units 1016, input / output ports 1018, input / output units 1020, and power supply 1024.
[0069] Those skilled in the art will appreciate that the bus 1010 can include one or more buses, such as an address bus, a data bus, a control bus, a network or any combination thereof. Those skilled in the art will further appreciate that many of the components in the example computing device 1000 can be implemented by a single device. Similarly, in some cases a single component can be implemented by multiple devices. Thus, Figure 10 Only an exemplary computing device capable of implementing one or more embodiments of the present application is shown and described, and this application is not limited to that particular computing device in any way.
[0070] The computing device 1000 can include various computer-readable media, which can be used for reading or writing data. For example, the computer-readable media can include random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other memory technologies, a CDROM, a digital versatile disk (DVD), a solid state drive (SSD), a cloud or other optical or holographic media, a magnetic cassette, a magnetic tape, a magnetic disk storage or other magnetic storage devices, or any other medium that can be used to encode information and be accessed by the computing device 1000.
[0071] The memory 1012 can include computer- storage media in the form of volatile and / or nonvolatile memory. The memory 1012 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include devices such as hard drives, solid-state memory, optical disc drives, and similar devices. The computing device 1000 can include one or more processors that read data from components such as the memory 1012, the various I / O components 1020, and the like. One or more presentation components 1016 present data indications to a user or other devices. Exemplary presentation components include a display device, a speaker, a printing component, a vibrating component, and the like.
[0072] The I / O ports 1018 can enable the computing device 1000 to be logically coupled to other devices such as the I / O components 1020, using a serial, parallel, or network and / or wireless communication protocol. Some of the I / O components 1020 can be built in to the computing device 1000. Examples of such I / O components 1020 include a microphone, a joystick, a recorder, a game pad, a satellite dish, a scanner, a printer, a wireless device, a network device, and the like. Illustrative Example: Purchase Intention
[0073] In the following example, the assigned quantified predisposition is purchase intention. Problem
[0074] Recognizing a customer's shopping intentions through their geolocation data can enable businesses to target specific areas, tailor promotional activities, and enhance the overall shopping experience. By analyzing patterns and trends in customer behavior, businesses can proactively engage potential customers and offer relevant products or services, ultimately increasing sales and customer engagement. While many businesses rely on traditional marketing and demographic data, integrating geolocation data can provide valuable insights into customer preferences, interests, and behaviors specific to their geographic location. This information can be used to develop location-based marketing campaigns, recommend nearby stores or services, and optimize inventory management based on the needs of different regions. As the retail industry evolves, the shift towards understanding customer intentions from geolocation data becomes increasingly important.
[0075] By leveraging predictive analytics and machine learning techniques, businesses can forecast customer demand, optimize resource allocation, and deliver personalized shopping experiences tailored to individual customers and their geographic locations. In summary, predicting customer shopping intentions from geolocation data is an emerging field that offers numerous benefits to businesses. By effectively utilizing these data, companies can gain a competitive advantage, improve customer satisfaction, and drive sales by delivering targeted marketing campaigns and personalized shopping experiences.
[0076] The present invention provides a solution to this problem. The technology operates on geolocation data collected from mobile phones in a privacy-preserving manner and utilizes proprietary cloud computing to predict customer visits to QSRs with higher efficiency than the current state-of-the-art. Proposed solution
[0077] Embodiments of the present invention relate to the proactive computer simulation of likely business visits when no user demographic data or data on user's previous behavior, such as spending on purchases or response to previous promotions, are available. The invention is an end-to-end process for creating a proxy of individual consumption factors to make this prediction by: (a) collecting geolocation data from mobile phones; (b) organizing and controlling these data; (c) computationally analyzing the data in a privacy-preserving cloud-based proprietary server; and (d) communicating the predictions to users and businesses in a timely manner through user applications and APIs. Framework
[0078] The proposed framework aims to extract a proxy of user's taste and consumption preferences from location data to simulate the user's choice of businesses. We now introduce the notation and state our research problem.
[0079] Definition 1 (Trajectory) Individual 's trajectory is defined as a set of tuples ordered in time where is a location, and is a coordinate of the geographic location (coordinates typically correspond to latitude and longitude), and is the corresponding timestamp.
[0080] Definition 2 (Activity Trajectory) of an individual is an activity trajectory is defined as mapping to activities that exhibit behavioral, consumption, leisure patterns. is a set of tuples ordered in time where is the activity inferred by the individual from the locations closest to and and is the coarse timestamp of (e.g., t i j = 9:33 AM is coarse-grained and represented as 9 AM to 11 AM). Furthermore, we denote as the full set of all temporal activities during .
[0081] Lifestyle (one of the proxies) captures the daily behavior of a consumer.
[0082] Definition 3 (Lifestyle) of an individual is a lifestyle is defined as a set of activities and their corresponding timestamps that represent the daily temporal activities of the individual during .
[0083] Next, we illustrate the conversion of a consumer trajectory ( ) to an activity trajectory ( ) (location and trajectory). We detail the identification of a lifestyle ( ) from a consumer location trajectory in Section 2.2.1. In Section 2.3, we discuss the methodology to simulate business visits from a proxy of , and other user consumption data (capturing consumer mobility, reachability, and social-demographic data).
[0084] Table 1 Activity Groups Location-to-activity traces
[0085] Mapping consumer locations to points of interest, such as place types - restaurants, grocery stores, or business types - Walmart, Five guys, etc., opens a new field of possibilities to understand and model the social and behavioral determinants of consumer consumption from a macro and micro perspective. For example, macro movements and temporal patterns of different competing brands inferred from such mappings have been used to decide the location of new franchises. In addition, micro, daily consumer-specific patterns, such as the number of visits to a place type or the time spent on different business types, can predict the next possible location of a consumer.
[0086] The mapping we designed aims to capture consumer activities based on sociologists' definition of lifestyles - an activity that exhibits a pattern of behavior, consumption, or leisure. To achieve this, we leverage publicly available resources to map locations to place types (see Table 1, second column), and use home, work definitions. Next, we group semantically similar place types (Table 1, first column) to form activity groups that represent consumer consumption (restaurants, unhealthy activities), leisure (entertainment, individual care, hotels, family, fitness), shopping (necessity shopping, leisure shopping), and commuting (public transportation, own transportation) behaviors. These 15 activity groups constitute the universe of all activities .
[0087] In addition, to abstract the exact time of day variation in daily activities, a coarser timestamp (granular timestamp associated with each consumer location), i.e. , associated with each activity: 12am-2am, 3am-5am, 5am-7am, 7am-9am, 9am-11am, 11am-2pm, 2pm-5pm, 5pm-7pm, 7pm-9pm, 9pm-12am. The tuples = produced in the consumer trajectories form the universe of all temporal activities (as defined in Def. 2 ). The mapping we made from locations to activity trajectories is explained in more detail next. Proxies from location and activity traces Lifestyle as a determinant of consumption
[0088] Lifestyles (e.g., work, life, travel, sports activities, diet, leisure, healthcare, consumption, etc.) can prove to have a strong relationship with individuals' economic preferences and future consumption behavior.
[0089] Given the large scale and high dimensionality, automatically discovering and modeling lifestyles from location data is a non-trivial problem. Furthermore, the variability in individual activity over several days, as well as the variability in activity between individuals, adds to the complexity. We employ an unsupervised topic modeling approach that has shown potential in revealing complex temporal and behavioral patterns to identify work, home, and consumption habits for smaller location datasets. Specifically, we leverage the concept of probabilistic author-topic modeling (ATM), designed for text documents, to model consumers' daily activities. Figure 11 As can be seen in the diagram, the probabilistic graphical model of this paper uses a plate notation. Leveraging our refined location data, we expand the scope of this literature by incorporating a broad set of fifteen activities representing patterns of behavior, consumption, and leisure.
[0090] Author-Topic Model (ATM): LDA is a probabilistic unsupervised learning model that uses a bag-of-words set and hidden discrete variables called topics. For text modeling, we can view each document as a mixture of various topics, where each topic is represented as a distribution of words. ATM incorporates LDA and assumes that document authors represent multinomial distributions over topics, where each topic is a probability distribution over words. The topic distribution of a document with multiple authors is a mixture of distributions associated with the authors. Figure 11 A graphical model of an ATM is shown. When generating a document, an author is randomly selected for each word in the document. This author selects a topic from a multinomial distribution over the topic, and then samples words from a multinomial distribution over words associated with that topic. We repeat this process for all words in the document. Formally, assume there are... Each theme individual authors One document and A unique word, word The probability is: ,in These are latent variables, indicating the topic from which the t-th word comes. The purpose of ATM inference is to determine each topic. Word distribution And each author Theme distribution . It is Dirichlet ( ), and P ( ) is Dirichley ( ),in and These are hyperparameters. We use the Gibbs approximation to estimate them as follows: (1) in and They are words and the author Assigned to topic The number of times. Similarly, These are the sums of words-topic and author-topic, respectively. Below is a more detailed introduction to ATM and Gibbs approximations. Next, we will delve into the consumer activity trajectory ( ATM-based lifestyle recognition in ( ).
[0091] Activity Tracking to Lifestyles: To identify lifestyles, we draw analogies between text documents and daily activities, authors, and consumers. We will Each activity in The mapped activity trajectory is regarded as a word We represent the daily activities of consumers (authors) as a bag-of-words document. We will consumers The multiple days are considered as unique documents by author a. Based on this, we estimate the parameters of the two ATM models using Equation 1. The equation represents each topic. Activity probability and each consumer Theme The probability of these probabilities. Given these probability distributions, we can rank activities for each discovered theme (lifestyle). We can also rank themes for consumers, thus discovering their primary lifestyles.
[0092] We represent each lifestyle as the top Y activities ranked by relevance (Sievert and Shirley 2014) – that is, the specific topic probability (first term in Equation 2) and lift (second term in Equation 2) for each activity. It is an event A convex combination of empirical distributions.
[0093] Next, we start with the estimated author-topic distribution. The most likely themes are assigned as consumers' primary lifestyles. This is combined with rankings based on relevance. For each activity, we can represent consumer i's lifestyle as... This completes the process from... Identifying individual lifestyles We detail the selection of various hyperparameters for the lifestyle recognition based on ATM in the following. Next, we discuss the proposed temporal deep learner to model the likelihood of future business visits. Business visit simulation
[0094] To assess whether a consumer visited a business, we overlay the consumer’s daily location trajectory on a publicly available repository of business facility locations.
[0095] To enrich our simulation signal, we augment the identified lifestyles with other healthcare proxies extracted from the location data. Proxies from location and activity traces
[0096] In Table 2, we describe the different aspects of the consumer attributes extracted from the location data, as well as the proxies we use to indicate the outcome of the consumer’s business visits.
[0097] Lifestyle: We use the discussed ATM-based technique to identify the workday and weekend lifestyles of the consumers from their respective activity trajectories. Table 2. Feature description for prediction
[0098] Activity: While the lifestyle captures the global daily activity of the consumer, the rich behavioral characteristics of the location data also enable us to capture the micro-level activity behavior of the day. We leverage the transformed activity trajectory (as defined above) to compute the daily visit frequency and dwell time of the consumer to each of the 15 activity groups as additional time-varying, numerical consumer attributes.
[0099] Mobility: Mobility indicators have been studied in the past and found to be associated with health outcomes. This set of attributes captures the daily mobility patterns of the consumer based on the locations visited in We also compute other richer mobility features such as entropy and turning radius. All of these are measured at the daily level and are time-varying numerical consumer attributes.
[0100] Accessibility: To incorporate this aspect of the consumer characteristics, we leverage the transformed activity trajectory (A ) and compute the accessibility of the consumer - the nearest distance to various public facilities such as hospitals, parks, gyms, pharmacies, public transportation, and work from the consumer’s home address. These attributes are static (non-time-varying), numerical.
[0101] Social Demographics: Based on the consumer home addresses from the converted activity traces and publicly available census data, we also compute several block-level social demographic data. These are static and include both categorical (e.g. employment poi, decor of consumer's workplace) and numerical (population count of consumer's census block) attributes. Modeling business visits
[0102] From Table 2 we can see that we have multiple types of consumer attributes - time-varying categorical (lifestyle, weekday / weekend) and numerical (work_freq, daily), static categorical (census_block_id) and numerical (commute_access). While the breadth of these attributes captures different aspects of the consumer that are relevant to the business visit outcome, this presents several modeling challenges. First, a joint representation of these unique types of attributes is needed to model the outcome. Second, this representation needs to account for the fact that interactions between these attributes can lead to better signal to predict the healthy outcome. For example, temporal correlations between different daily activity attributes, as opposed to daily trends in each attribute, can lead to a better prediction model. Third, given that these attributes capture multiple aspects of the consumer attributes, simply concatenating all representations can lead to a poor prediction model. For example, one simple approach to represent all attributes as time series can be to concatenate static categorical / numerical features with each timestamp, resulting in a highly parameterized, possibly overfit model.
[0103] We address these issues by separately learning representations of the time-varying and static features that account for the interactions between different types of attributes (temporal and static). Next, we combine these to allow for interactions between the two representations to learn a final joint representation of all attributes.
[0104] To achieve this, we represent time-varying attributes with contextual-LSTM (CLSTM) cells, which are a modification of traditional LSTM cells that are widely used for word translation and time series modeling. As the name suggests, CLSTM allows for the simultaneous incorporation of time-varying and static contextual features into a time series. In the original application, the time-varying contextual features considered are latent topics of words jointly represented with the words (each word of the time series is concatenated with an embedding of a topic) to predict the next possible word in a sentence. Extending this to our setting, we note that lifestyles (lifestyle attributes in Table 2) are latent topics learned from different activities. Hence, we treat these as the context of time-varying attributes (activities in Table 2) related to different activities. Furthermore, we observe that treating lifestyles as the context of other time-varying attributes (mobility in Table 2) gives better predictive performance empirically. Next, we concatenate these representations for a given time period with the embeddings of static categorical and numerical attributes (sociodemographics and reachability) to jointly learn representations of all consumer attributes that predict the consumer’s health outcome. The concatenation of multiple views of consumer attributes that form a unified representation is extensively studied in multi-modal learning.
[0105] Figure 12 An overview of the architecture of the proposed time-series deep learning model is presented in FIG. 1. The box in the top left corner of the figure shows the modeling of the temporal attributes of a consumer at the daily level with CLSTM cells (multiple days as a CLSTM layer), where lifestyles are treated as the context of activity and mobility attributes. The bottom left box shows the representation of static consumer attributes, which are later concatenated with the temporal representations to predict the consumer’s health outcome. Next, we formally detail the transformations performed by the layers in the learner proposed by us. Proposed learner
[0106] Let denote the tensor of time-varying numerical attributes (number of users x number of observation periods x number of days x number of time-varying numerical features), denote the tensor of time-varying categorical attributes (number of users x number of weeks x number of time-varying categorical features), the matrix and denote the static numerical and categorical consumer attributes, respectively. To simplify the notation, in the discussion below, we will focus on the transformation of the attributes of a single consumer represented by , , and and their transformation to the probability of a health outcome (health risk).
[0107] Embedding: The embedding layer transforms a one-hot encoded categorical attribute ( , ) to a fixed-dimension continuous vector representation. Formally, (3)
[0108] where, - the number of time-varying categorical attributes x , - the number of static categorical attributes x are learnable weight parameters, , are tunable model hyperparameters. Recall that in our setting, lifestyle (weekend vs. weekday) (Lifestyle ), both of which are represented by ten relevant activities (Activity , the full set of activities ). Thus, a consumer's weekday and weekend lifestyle can both be represented as a vector of length |, which means we learn two weight matrices of dimension W to compute . A similar procedure is followed to convert other static categorical attributes (employee_poi, census_block_id).
[0109] CLSTM layer: The CLSTM layer as shown in Figure 13 is composed of multiple CLSTM units, each of which acts on a different day of . Assuming corresponds to all the numerical, embedded categorical time-varying consumer attributes for an arbitrary day, the computation of depends on whether the day is a weekday or weekend, as our lifestyles are derived for weekdays / weekends rather than specific dates, each CLSTM unit performs the following transformation: (4) The above four equations detail the modification of a traditional LSTM unit, where and are the input, output, and forget gates, respectively, to incorporate the additional context . Rearranging the terms, we note that this is equivalent to considering a composite input because: (5) Each CLSTM unit transforms the concatenated input into a hidden representation with learnable shared weight and bias parameters and tunable hyperparameters (dimension: number of consumers x ). Thus, the result of the CLSTM layer is represented as where is the number of days in our observation period.
[0110] Concatenation: The concatenation layer does not contain any learnable parameters and simply serves to combine different intermediate representations. We perform two concatenations (see Figure 2 ). The first one is as discussed above, where we construct the composite input into the CLSTM layer (as shown in Figure 13 ). The next concatenation is performed on the hidden temporal representation obtained from the CLSTM layer ( ), the embedded static ( ) and the numerical attributes ( ). Note that , the hidden layer representation of the last day of observation, captures temporal relationships of the previous days due to the recursive nature of equation 5, we combine it with , to form , the final joint representation that includes both time-varying and static attributes.
[0111] Shopping intention: We pass the final representation into a fully connected dense layer that allows for interaction between temporal and static attributes and assigns a quantitative likelihood of the shopper's intention as: (6) where are learnable parameters. For a given binary health outcome, to learn the various weights (denoted as in equations 3, 5, 6), we minimize the binary cross-entropy loss between the observed shopping outcome (e.g., shopping_visit) and , the vector of shopping intention in the above equation. The remaining hyperparameters are tuned via cross-validation. Details will be discussed below. Empirical study Data
[0112] For analytical purposes, we combined several datasets: GPS location tracking data at the individual level, demographic data at the census block level from the American Community Survey (2016), and public datasets of hospitals, emergency medical services, and emergency care facilities from HILFD. For the location data, we partnered with a leading data collection agency that aggregates location data from hundreds of commonly used mobile applications, ranging from news to weather, map navigation, and fitness. Location data collection was performed within a framework that is GDPR and CCPA compliant. These data cover a quarter of the US population on both Android and iOS operating systems. Each row of data corresponds to a location recorded for an individual. Each row contains the following information: Individual ID: Anonymized unique identifier of the individual using the mobile application, Latitude, longitude, and timestamp of the visited location, Speed at which the location was captured.
[0113] In total, we obtained individual location data from Boston for five months (September to January) in 2019. We only considered consumers that appeared in all four months and were tracked for more than ten days in each month. Furthermore, we removed consumers for which we could not estimate the work and home address based on heuristic methods discussed in Appendix B. Our final dataset includes over 23,000 individuals. In Tables 3, 4, and Figure 4 we detail and provide summary statistics for different types of consumer attributes computed from the location traces, activity mappings, census tracts, and public medical facility data. We discuss them in detail next. Summary statistics
[0114] Location traces: In Table 3 (mobility row), we present summary statistics for the raw location data
[12] . On average, each consumer had about 31 locations per day over the four months, of which about 18 were unique. The average speed at which these locations were captured was 6.92 km / h. For the remaining measures, we removed locations that were captured at speeds exceeding 5 km / h and only considered stay locations - locations for which a consumer spent at least 5 minutes. The average great-circle distance between stay locations was 7.82 km, and the average time spent at these locations was about 2.16 hours. Overall, the location data summary statistics indicate fine-grained consumer observations for both cities.
[0115] Activity Trajectories: Table 3 (Activity and Reachability columns) and Table 4 detail the summary statistics of activity trajectories (see Section 2.1 for the conversion of locations to activity trajectories). From Table 3, we observe that, among the 15 predefined activity groups (see Table 1), home, work, and public transportation are the top three activity groups in terms of average number of occurrences and time spent per day. When broken down by weekdays and weekends (Table 4), as expected, we observe that the frequency of work occurrence is lower during weekends (0.79) compared to weekdays (4.27). In addition, home occurs more frequently during weekends (5.57) compared to weekdays (8.97). To accommodate the differences in the most common activities, we learn the lifestyles for weekdays and weekends separately for the empirical analysis. For weekdays, we have an average of 14.12 daily activities per consumer, based on the bag-of-words representation of consumer activities that we use for lifestyle identification (as discussed above), each document is converted to an average of 14 words. Table 3 Summary statistics of consumer attributes During weekends, this number remains similar. The average number of unique daily activities during weekends and weekdays are 10.16 and 9.82, respectively.
[0116] In Figure 14 , we plot the heatmaps of activity occurrences in activity trajectories during weekends and weekdays. For a given row (activity group), in Figure 14 a and Figure 14 b, lighter / darker red cells indicate lower / higher occurrence rates during the corresponding time slots. From Figure 14 a, we observe that the highest occurrence of work happens between 2pm and 5pm, while home occurs between 12pm and 3pm; most consumers are less likely to stay at home during weekdays (between 9am and 7pm). In contrast, during weekends ( Figure 14 b), we observe that consumers are more likely to stay at home during the same time period. In addition, other activities related to leisure, shopping, and consumption occur earlier (between 9am and 11am) during weekdays compared to weekends (between 12pm and 3pm). For example, unhealthy activities of the activity group are more likely to occur between 9pm and 11pm during weekdays compared to 12pm and 5pm during weekends.
[0117] Census Block Socio-Demographics: To assign a census block, the Haversine distance between an individual’s home address and the interpolated center latitude and longitude of the census block group is computed. Table 3 (Socio-Demographics) details the summary statistics of socio-demographic attributes.
[0118] The terms “comprise” and “comprising”, as used herein, are intended to be interpreted as having a broad meaning, rather than a limiting meaning. As used herein, the terms “exemplary”, “example” and “illustrative” are used as illustrative only and not to indicate or imply that a preferred or advantageous configuration is indicated or contemplated. The terms “about”, “generally” and “approximately”, as used herein, are intended to cover variations that can exist in the upper and lower limits of a range of subjective or objective values, such as variations in attributes, parameters, dimensions and dimensions. In one non-limiting example, the terms “about”, “generally” and “approximately” mean equal to or within 10% or less, or 10% or less. In one non-limiting example, the terms “about”, “generally” and “approximately” mean close enough to be considered within the scope of the relevant field by a person of ordinary skill in the art. As used herein, the term “substantially” refers to a full or nearly full range or degree of an action, characteristic, attribute, state, structure, item or result as would be understood by a person of ordinary skill in the art. For example, an object that is “substantially circular” means that the object is either a perfect circle within mathematically determinable limits, or an approximate circle as recognized or understood by a person of ordinary skill in the art. In some cases, the exact degree of allowable deviation from absolute perfection can depend on the particular circumstances. In general, however, the degree of closeness to achieve would be the same overall result as achieving or obtaining absolute and perfect completion. When used in a negative sense, the use of “substantially” is equally applicable to refer to a full or near full lack of an action, characteristic, attribute, state, structure, item or result as would be understood by a person of ordinary skill in the art.
[0119] Numerous modifications and alternative embodiments of the present application will be apparent to those skilled in the art in view of the foregoing description. Accordingly, this description is to be construed as illustrative only and is for the purpose of teaching those skilled in the art the best mode of carrying out the present application. Structural details can be substantially varied without departing from the spirit of the present application, and the exclusive use of all modifications falling within the scope of the appended claims is reserved. In this description, embodiments are described in a manner that enables a clear and concise description to be made, but it can be understood that the embodiments can be combined or separated differently without departing from the present application. The present application is limited only by the scope of the appended claims and the rules of applicable law.
[0120] It should also be understood that all general and specific features described herein are intended to be covered by the present application, as described herein, and all statements of the scope of the present application that can be linguistically considered to fall within the scope of the present application. CLAIM (AMENDED PURSUANT TO ARTICLE 19 OF THE TREATY) 1. A method of modeling user predisposition based on location data, the method comprising: I. training at least one model to determine a user’s predisposition to a particular behavior or experience / condition, the training comprising: A) obtaining location data for a plurality of users; B) formatting the location data into one or more trajectories for each user of the plurality of users; C) inferring activity data from the location data; D) converting the activity data and the location data into time-varying and static behavioral attributes for each user of the plurality of users; and E) determining a predisposition to a particular behavior or experience / condition, including, for each user: modeling the time-varying attributes; combining the modeled time-varying attributes with static attributes; assigning a quantified predisposition for the particular behavior or experience / condition; and adjusting parameters based on results, resulting in a trained model; and II. predicting an individual user’s predisposition to a particular behavior or experience / condition, comprising: A) obtaining location data for the individual user; B) inputting the location data into at least one trained model resulting from the training of at least one model; and C) receiving from the trained model an assigned quantified predisposition of the individual user to a particular behavior or experience / condition; wherein converting the activity data and the location data into behavioral attributes for each user of the plurality of users comprises defining time-varying attributes and defining static attributes; and wherein combining the modeled time-varying attributes with the static attributes comprises: embedding categorical static attributes into a fixed-dimension continuous vector representation with learnable weight parameters and tunable model hyperparameters; concatenating the embedded categorical static attributes, numerical static attributes, and a hidden representation of the modeled time-varying attributes. 2. The method of claim 1, wherein the predisposition to a particular behavior or experience / condition is purchase intent. 3. The method of claim 1, wherein the predisposition to a particular behavior or experience / condition is hospitalization risk. 4. The method of claim 1, wherein the predisposition to a particular behavior or experience / condition is work engagement. 5. The method of claim 1, wherein the predisposition to a particular behavior or experience / condition is job change. 6. The method of claim 1, wherein the predisposition to a particular behavior or experience / condition is travel intent. 7. The method of claim 1, wherein the predisposition to a particular behavior or experience / condition is residential relocation intent. 8. The method of claim 1, wherein the predisposition to a particular behavior or experience / condition is healthcare risk. 9. The method of claim 1, wherein the predisposition to a particular behavior or experience / condition is associated with a window of opportunity. 10. The method of claim 1, wherein the location data is geospatial data of a user hardware device over a period of time. 11. The method of claim 1, wherein the location data of a user device is provided by a location provider / supplier. 12. The method of claim 1, wherein the location data of a user is identified without using one or more of: personally identifiable information (PII), demographic information, and socio-economic information about the user. 13. The method of claim 1, wherein inferring the activity data from the location data further comprises: i) mapping the location data to place types; ii) grouping the place types into activity groups based on the functionality of the place types; and iii) using the activity groups to convert one or more trajectories of the user into one or more activity trajectories of the user. 14. The method of claim 13, wherein the activity groups comprise one or more selected from the group of: hospital, health, essential shopping, fitness, public transportation, own transportation, religious, entertainment, travel, individual care, leisure shopping, unhealthy activity, restaurant, home, and work. 15. The method of claim 1, wherein defining the time-varying attributes comprises: determining lifestyle attributes; determining activity attributes; and determining mobility attributes. 16. The method of claim 15, wherein determining the lifestyle attributes comprises using an unsupervised learning model that can identify similarities in activity patterns between the users. 17. The method of claim 16, wherein the unsupervised learning model comprises a clustering and dimensionality reduction model. 18. The method of claim 16, wherein the unsupervised learning model comprises a hidden Markov model. 19. The method of claim 16, wherein the unsupervised learning model comprises an LDA and topic model. 20. The method of claim 19, wherein the LDA and topic model comprises an author topic model (ATM), where the users are the authors, the activities are the words, the activity periods are the documents, and the lifestyles are the topics. 21. The method of claim 1, wherein defining the static attributes comprises: determining reachability attributes; and determining socio-demographic attributes. 22. The method of claim 15, wherein defining the time-varying attributes comprises embedding the lifestyle attributes into a fixed-dimension continuous vector representation with learnable weight parameters and tunable model hyperparameters, concatenating the embedded lifestyle attributes with the activity attributes and the mobility attributes, and converting the concatenated embedded lifestyle attributes and the time-varying numerical attributes into a hidden representation with shared learnable weights and bias parameters and tunable model hyperparameters. 23. The method of claim 22, wherein converting the concatenated embedded lifestyle attributes and the time-varying numerical attributes into a hidden representation comprises using a supervised machine learning model that models both spatial and temporal information from user trajectories. 24. The method of claim 23, wherein the supervised machine learning model comprises a non-deep learning regression or classification model. 25. The method of claim 24, wherein the non-deep learning regression or classification model comprises a decision tree-based model, a random forest-based model, or a gradient boosting model. 26. The method of claim 23, wherein the supervised machine learning model comprises a deep learning model. 27. The method of claim 23, wherein the supervised machine learning model comprises a deep neural network-based model. 28. The method of claim 27, wherein the deep neural network-based model comprises at least one of: a convolutional neural network (CNN), a recurrent neural network (RNN), a long short-term memory (LSTM) model, a radial basis function network (RBFN), a Transformer-based attention model. 29. The method of claim 28, wherein the deep neural network-based model comprises: a convolutional long short-term memory (CLSTM) model that takes the time-varying attributes as input. 30. The method of claim 1, wherein the model is tuned via cross-validation. 31. The method of claim 1, further comprising outputting the assigned quantitative predisposition of the particular user to a particular behavior or experience / condition. 32. The method of claim 31, wherein the outputting comprises providing a graphical representation of the assigned quantitative predisposition of the particular user to a particular behavior or experience / condition. 33. The method of claim 31, wherein the outputting comprises providing a web service API that provides access to the assigned quantified propensity of a particular behavior or experience / condition of the particular user. 34. A system for modeling user propensity based on location data, the system comprising: a data collection module configured to obtain location data for a plurality of users; a model training module configured to train a model using the location data to determine a propensity of a user for a particular behavior; and a prediction module configured to use the trained model and location data for an individual user to predict a propensity of the individual user for a particular behavior. 35. The system of claim 34, wherein the model training module is further configured to infer activity data from the location data, convert the activity data and the location data into behavioral attributes, and determine a propensity of each user for a particular behavior. 36. The system of claim 34, wherein the prediction module is further configured to input the location data into the trained model and receive from the trained module an assigned quantified propensity for a particular behavior of the particular user.
Claims
1. A method for modeling user preferences based on location data, the method comprising: I. Training at least one model to determine a user's tendency toward a specific behavior or experience / condition, the training comprising: A) acquiring location data of a plurality of users; B) formatting the location data into one or more trajectories for each of the plurality of users; C) inferring activity data from the location data; D) converting the activity data and the location data into time-varying and static behavioral attributes for each of the plurality of users; and E) determining a tendency toward a specific behavior or experience / condition, comprising: for each user: modeling the time-varying attributes; combining the modeled time-varying attributes with static attributes; assigning a quantitative tendency toward the specific behavior or experience / condition; and adjusting parameters based on the results to obtain a trained model; and II. Predicting an individual user's tendency toward a specific behavior or experience / condition, comprising: A) acquiring location data of the individual user; B) inputting the location data into at least one trained model generated by training at least one model; and C) receiving from the trained model the assigned quantitative tendency of the individual user toward the specific behavior or experience / condition.
2. The method according to claim 1, wherein, A tendency toward a particular behavior or experience / situation is a purchase intention.
3. The method according to claim 1, wherein, A predisposition to a particular behavior or experience / condition is a risk factor for hospitalization.
4. The method according to claim 1, wherein, A tendency toward a particular behavior or experience / situation is a characteristic of work participation.
5. The method according to claim 1, wherein, A tendency toward a particular behavior or experience / situation is a job change.
6. The method according to claim 1, wherein, A tendency toward a particular behavior or experience / situation constitutes travel intention.
7. The method according to claim 1, wherein, A tendency toward a particular behavior or experience / situation is a residential relocation intention.
8. The method according to claim 1, wherein, A tendency toward a particular behavior or experience / condition is a healthcare risk.
9. The method according to claim 1, wherein, A tendency toward a particular behavior or experience / situation is associated with an opportunity window.
10. The method according to claim 1, wherein, The location data refers to the geospatial data of the user's hardware device over a period of time.
11. The method according to claim 1, wherein, The location data for user equipment is provided by the location provider / supplier.
12. The method according to claim 1, wherein, User location data is identified without using one or more of the following: Individually Identifiable Information (PII), demographic information, and socioeconomic information about the user.
13. The method according to claim 1, wherein, Inferring the activity data from the location data further includes: i) mapping the location data to location types; ii) grouping the location types into activity groups based on the functionality of the location types; and iii) using the activity groups to convert one or more trajectories of the user into one or more activity trajectories of the user.
14. The method according to claim 13, wherein, The activity groups include one or more selected from the following groups: hospitals, health, essential shopping, fitness, public transportation, private transportation, religion, recreation, travel, personal care, leisure shopping, unhealthy activities, restaurants, home, and work.
15. The method according to claim 1, wherein, Converting the activity data and the location data into behavioral attributes for each of the multiple users includes: defining the time-varying attributes; and defining the static attributes.
16. The method according to claim 15, wherein, The time-varying attributes are defined as follows: determining lifestyle attributes; determining activity attributes; and determining mobility attributes.
17. The method according to claim 16, wherein, Determining the lifestyle attributes involves using an unsupervised learning model that can identify similarities in activity patterns among the users.
18. The method according to claim 17, wherein, The unsupervised learning models include clustering and dimensionality reduction models.
19. The method of claim 17, wherein, The unsupervised learning model includes the Hidden Markov Model.
20. The method of claim 17, wherein, The unsupervised learning models include LDA and topic models.
21. The method according to claim 20, wherein, The LDA and topic model include the Author-Topic Model (ATM), where users are authors, activities are words, activity periods are documents, and lifestyles are topics.
22. The method according to claim 15, wherein, The definition of the static attributes includes: determining the reachability attributes; and determining the socio-demographic attributes.
23. The method according to claim 16, wherein, The definition of the time-varying attribute includes: embedding the lifestyle attribute into a fixed-dimensional continuous vector representation with learnable weight parameters and adjustable model hyperparameters; concatenating the embedded lifestyle attribute with the activity attribute and the mobility attribute; and converting the concatenated embedded lifestyle attribute and the time-varying numerical attribute into a hidden representation with shared learnable weights and bias parameters and adjustable model hyperparameters.
24. The method according to claim 23, wherein, Converting the spliced, embedded lifestyle attributes and the time-varying numerical attributes into a hidden representation involves using a supervised machine learning model that models both spatial and temporal information from user trajectories.
25. The method according to claim 24, wherein, The supervised machine learning models include non-deep learning regression or classification models.
26. The method of claim 25, wherein, The non-deep learning regression or classification models include decision tree-based models, random forest-based models, or gradient boosting models.
27. The method according to claim 24, wherein, The supervised machine learning model includes a deep learning model.
28. The method according to claim 24, wherein, The supervised machine learning model includes models based on deep neural networks.
29. The method according to claim 28, wherein, The deep neural network-based model includes at least one of the following: Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM) model, Radial Basis Function Network (RBFN), and Transformer-based attention model.
30. The method according to claim 29, wherein, The deep neural network-based model includes a convolutional long short-term memory (CLSTM) model that takes the time-varying properties as input.
31. The method according to claim 15, wherein, Combining the modeled time-varying attributes with the static attributes includes: embedding the categorical static attributes into a fixed-dimensional continuous vector representation with learnable weight parameters and adjustable model hyperparameters; and concatenating the hidden representations of the embedded categorical static attributes, numerical static attributes, and modeled time-varying attributes.
32. The method according to claim 1, wherein, The model was adjusted through cross-validation.
33. The method of claim 1, further comprising outputting a quantitatively assigned tendency of the particular user toward a particular behavior or experience / situation.
34. The method according to claim 33, wherein, The output includes a graphical representation of the assigned quantitative tendencies for a particular user’s specific behavior or experience / condition.
35. The method according to claim 33, wherein, The output includes providing a web service API that provides assigned, quantified preferences for specific behaviors or experiences / conditions of the particular user.
36. A system for modeling user preferences based on location data, the system comprising: The data collection module is configured to acquire location data from multiple users; A model training module is configured to train a model using the location data to determine a user's tendency toward a specific behavior; The prediction module is configured to use a trained model and the individual user's location data to predict the tendency of the individual user to engage in specific behaviors.
37. The system according to claim 36, wherein, The model training module is also configured to infer activity data from location data, convert the activity data and the location data into behavioral attributes, and determine each user's tendency toward specific behaviors.
38. The system according to claim 36, wherein, The prediction module is also configured to input the location data into a trained model and receive, from the trained module, an assigned quantitative tendency for a specific behavior of the particular user.