Methods and system for generation of audience data for advertising
The system addresses the inefficiencies of conventional data processing methods by using a processor and database to generate audience data from geolocation information, enabling efficient and cost-effective targeting of consumer audiences for advertising.
Patent Information
- Application Number
- PCT/CA2024/051577
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-27
- Filing Date
- 2024-11-27
- Publication Date
- 2025-06-05
AI Technical Summary
Conventional methods for collecting and analyzing consumer-related data for advertising are time-consuming and expensive, and they struggle to efficiently target specific consumer audiences.
A system and method for generating audience data by processing geolocation data from consumers, using a processor and database with geofiles, geofence, and point-of-interest datasets to create a partitioned combined dataset, and generating audiences on demand based on query parameters.
The system enables efficient and cost-effective generation of high-quality consumer audiences, reducing the time and expense associated with conventional data processing methods while providing accurate targeting for advertising purposes.
Smart Images

Figure CA2024051577_05062025_PF_FP_ABST
Abstract
Description
METHODS AND SYSTEM FOR GENERATION OF AUDIENCE DATA FOR ADVERTISINGRELATED APPLICATION
[0001] The present application claims priority to or benefit of United States provisional patent application No. 63 / 602,885, filed November 27, 2023, which is incorporated herein by reference in its entirety.TECHNICAL FIELD
[0002] The present disclosure relates to processing large volumes of consumer-related data and specifically to a method and a system for generation of audience data related to consumers for advertising.BACKGROUND
[0003] In order to efficiently advertise products and services over computer networks, enormous volumes of data about consumers’ behavior and location need to be collected and analyzed. Although collecting consumer-related data became easy with wide-spread use of electronic devices of all sizes, the volumes of data have become enormous to process. Nevertheless, vendors and marketers are preoccupied with proper targeting audiences of consumers to increase sales, thus requiring customization and narrowing down the collected consumer-related data. In view of immense volumes of data that can be collected, conventional methods for collection of consumer-related data and analysis of such data may be time-consuming and expensive.SUMMARY
[0004] It is an object of the present disclosure to provide a system and methods for generation of data related to consumers. The system and methods provide collecting, storing, and generating data related to consumers, and in particular consumer audiences, on demand by processing data related to consumers and their locations, as well as their behaviors with respect to the locations, such as, for example, time and periods of time spent and detected at the location. The system and methods use geolocation data related to locations of consumers and generate various audiences and audience extensions on demand. The system and methods described herein may be used for advertising purposes, such as, for example, and without limitation, providing data for use by a Demand Side Platform (DSP) and / or for marketing attribution.
[0005] According to one aspect of the disclosed technology, there is provided a method for generation of data related to consumers to be executed by a system comprising a processor and a database comprising a geofiles database, a geofence dataset (also referred to herein as a “geofence database”) and a point-of-interest database. In at least one embodiment, the method comprises: receiving, from a plurality of electronic devices, location data points related to the consumers, each location data pointcomprising, for each electronic device of each consumer, a mobile advertiser identifier (also referred to herein as “MAID” or a “unique device identifier”), a timestamp, a physical location identification of the electronic device, and an internet protocol (IP) address of the electronic device; generating a partitioned combined dataset by: mapping, by a mapping routine, the mobile advertiser identifier to a geofiles data in the geofiles database and to a geofence data in the geofence dataset based on the physical location identification for each location data point of the location data points; mapping, by the mapping routine, the mobile advertiser identifier to a company data in the point-of-interest database corresponding to the mobile advertiser identifier to generate a combined dataset; splitting the combined dataset with respect to geohashes and time periods to generate combined dataset partitions; compressing each one of the combined dataset partitions to generate the partitioned combined dataset; and storing the partitioned combined dataset based on the corresponding geohash identification and the corresponding time periods. The method may further comprise assigning additional characteristics to the mobile advertiser identifier of the location data point, the additional characteristics being at least one of consumer home location and consumer work location.
[0006] In at least one embodiment, the method may comprise: receiving, from a user device, a query data comprising requested location address and an audience refining parameters; based on the requested location address, generating a query-related polygon and generating requested polygon addresses located within the query-related polygon; mapping the requested polygon addresses to building polygons using the geofence database; mapping the requested polygon addresses located within the query-related polygon to at least one corresponding query-related geohash to determine at least one query-related combined dataset partition of the combined dataset partitions based on the query-related geohash and a query-related time period; searching the at least one query-related combined dataset partition to identify polygon mobile advertiser identifiers (also referred to herein as “polygon MAIDs”) related to the query-related polygon and the audience refining parameters; and generating an initial audience comprising the polygon mobile advertiser identifiers. The query-related time period may be provided in the query data or determined by the processor based on the query data. The query-related time period may be a daytime period, a nighttime period or a visitor time period which may comprise shorter term periods of time compared to the daytime period or the nighttime period. The initial audience may be a residential initial audience when the query-related time period is the nighttime period, a work initial audience when the query-related time period is the daytime period, or a visitor initial audience when the query-related time period is the visitor time period.
[0007] In at least one embodiment, the initial audience is a residential initial audience, initial residential MAIDs being the polygon mobile advertiser identifiers, and the method comprises: for each initial residential MAID, using the partitioned combined dataset, determining a work building geofence related to a location of the initial residential MAID during an extended day-time period; extracting, from the partitioned combined dataset, work-extended MAIDs, different from the initial residential MAIDs, detectedwithin the work building geofence of the initial residential MAIDs; and generating an extended work-based audience having the work-extended MAIDs and the initial residential MAIDs of the residential initial audience. In at least one embodiment, the initial audience is a work initial audience, initial work MAIDs being the polygon mobile advertiser identifiers, and the method further comprises: for each initial work MAID, using the partitioned combined dataset, determining a residential building geofence related to a location of the initial work MAID during an extended night-time period; extracting, from the partitioned combined dataset, home-extended MAIDs, different from the initial work MAIDs, detected within the residential building geofence of the initial work MAIDs; and generating an extended home-based audience having the home-extended MAIDs and the initial work MAIDs of the work initial audience.
[0008] In at least one embodiment, the method may comprise: for each initial MAID of the initial MAIDs, using the partitioned combined dataset, determining a building geofence related to a location of the initial MAID during a first pre-determined time period, the initial MAIDs being the polygon mobile advertiser identifiers; extracting, from the partitioned combined dataset, social-extended MAIDs, different from the initial MAIDs, detected within the building geofence of the initial MAIDs during a second pre-determined time period; and generating an extended social-based audience having the social-extended MAIDs and the initial MAIDs of the initial audience.
[0009] The method may comprise generating a targeted extended audience by reducing the initial audience based on a building attendance frequency determined by the processor and related to each one of the polygon mobile advertiser identifiers, the building attendance frequency being within a predetermined building attendance frequency range.
[0010] In at least one embodiment, the method may comprise: mapping keywords to locations of interest of point-of-interest locations of the point-of-interest database; mapping point-of-interest locations to buildings using the geofence dataset; based on the point-of-interest location geohashes of the point-of- interest locations, extracting audience attendance data; and generating a syndicated audience comprising syndicated audience MAIDs. In at least one embodiment, the method may comprise: for each syndicated audience MAID of the syndicated audience MAIDs, using the partitioned combined dataset, determining a building geofence related to a location of the syndicated audience MAID during a first predetermined time period; extracting, from a location data point dataset partition corresponding to a geohash of the building geofence, social-extended MAIDs, different from the syndicated audience MAIDs, detected within the building geofence of the syndicated audience MAIDs during a second pre-determined time period; and generating an extended point-of-interest audience having the extended MAIDs and the syndicated audience MAIDs of the syndicated audience. The at least one requested polygon address located within the query-related polygon may be determined by searching a geofence dataset. The query data may comprise a requested audience type and the polygon mobile advertiser identifiers are extracted from the query-related combined dataset partition based on the requested audience type and geohash mapping to a requested building geohash.
[0011] According to another aspect of the disclosed technology, there is provided a system for modeling of data related to consumers and locations. In at least one embodiment, the system comprises: a database comprising a geofiles database, geofence dataset and point-of-interest dataset, and a processor configured to: receive, from a plurality of electronic devices, location data points related to the consumers, each location data point comprising, for each electronic device of each consumer, a mobile advertiser identifier, a timestamp, a physical location identification of the electronic device, and an IP address of the electronic device; generate a partitioned combined dataset by: mapping, by a mapping routine, the mobile advertiser identifier to a geofiles data in the geofiles database and to a geofence data in the geofence dataset based on the physical location identification for each location data point of the location data points; mapping, by the mapping routine, the mobile advertiser identifier to a company data in the point-of-interest database corresponding to the mobile advertiser identifier to generate a combined dataset; splitting the combined dataset with respect to geohashes and time periods to generate combined dataset partitions; compressing each one of the combined dataset partitions to generate the partitioned combined dataset; and storing the partitioned combined dataset based on the corresponding geohash identification and the corresponding time periods. The processor may be further configured to assign additional characteristics to the mobile advertiser identifier of the location data point, the additional characteristics being at least one of consumer home location and consumer work location.
[0012] The processor may be further configured to: receive, from a user device, a query data comprising requested location address and an audience refining parameters; based on the requested location address, generate a query-related polygon and generating requested polygon addresses located within the query-related polygon; map the requested polygon addresses to building polygons using the geofence database; map the requested polygon addresses located within the query-related polygon to at least one corresponding query-related geohash to determine at least one query-related combined dataset partition of the combined dataset partitions based on the query-related geohash and a query-related time period; search the at least one query-related combined dataset partition to identify polygon mobile advertiser identifiers related to the query-related polygon and the audience refining parameters; and generate an initial audience comprising the polygon mobile advertiser identifiers. The query-related time period may be provided in the query data or determined by the processor based on the query data.
[0013] In at least one embodiment, the initial audience may be a residential initial audience, initial residential MAIDs being the polygon mobile advertiser identifiers, and the processor may be further configured to: for each initial residential MAID, using the partitioned combined dataset, determine a work building geofence related to a location of the initial residential MAID during an extended day-time period; extract, from the partitioned combined dataset, work-extended MAIDs, different from the initial residential MAIDs, detected within the work building geofence of the initial residential MAIDs; and generate an extended work-based audience having the work-extended MAIDs and the initial residential MAIDs of the residential initial audience.
[0014] In at least one embodiment, the initial audience may be a work initial audience, initial work MAI Ds being the polygon mobile advertiser identifiers, and the processor may be further configured to: for each initial work MAID, use the partitioned combined dataset, determine a residential building geofence related to a location of the initial work MAID during an extended night-time period; extract, from the partitioned combined dataset, home-extended MAIDs, different from the initial work MAIDs, detected within the residential building geofence of the initial work MAIDs; and generate an extended home-based audience having the home-extended MAIDs and the initial work MAIDs of the work initial audience. The processor may be further configured to: for each initial MAID of the initial MAIDs, using the partitioned combined dataset, determine a building geofence related to a location of the initial MAID during a first predetermined time period, the initial MAIDs being the polygon mobile advertiser identifiers; extract, from the partitioned combined dataset, social-extended MAIDs, different from the initial MAIDs, detected within the building geofence of the initial MAIDs during a second pre-determined time period; and generate an extended social-based audience having the social-extended MAIDs and the initial MAIDs of the initial audience.
[0015] The processor may be further configured to generate a targeted extended audience by reducing the initial audience based on a building attendance frequency determined by the processor and related to each one of the polygon mobile advertiser identifiers, the building attendance frequency being within a pre-determined building attendance frequency range.
[0016] The processor may be further configured to: map keywords to locations of interest of point-of- interest locations of the point-of-interest database; map point-of-interest locations to buildings using the geofence dataset; based on the point-of-interest location geohashes of the point-of-interest locations, extract audience attendance data; and generate a syndicated audience comprising syndicated audience MAIDs.
[0017] The processor may be further configured to: foreach syndicated audience MAID of the syndicated audience MAIDs, using the partitioned combined dataset, determine a building geofence related to a location of the syndicated audience MAID during a first pre-determined time period; extract, from a location data point dataset partition corresponding to a geohash of the building geofence, social-extended MAIDs, different from the syndicated audience MAIDs, detected within the building geofence of the syndicated audience MAIDs during a second pre-determined time period; and generate an extended point-of-interest audience having the extended MAIDs and the syndicated audience MAIDs of the syndicated audience. The at least one requested polygon address located within the query-related polygon may be determined by searching a geofence dataset. The query data may comprise a requested audience type and the polygon mobile advertiser identifiers are extracted from the query-related combined dataset partition based on the requested audience type and geohash mapping to a requested building geohash.
[0018] According to one aspect of the disclosed technology, there is provided a method for modeling of data related to consumers and locations, the method comprising: collecting location data points related to the consumers comprising, for each electronic device of each consumer, a unique device identifier (also referred to herein as the “mobile advertiser identifier” or “MAID”), a timestamp, a latitude and a longitude, and an IP address of an electronic device of each consumer; generating a combined dataset by: mapping location data points to a geohash, and a business and / or a building based on the latitude and the longitude for each electronic device by using geofiles; mapping a point-of-interest database to the building and / or a brand name; assign characteristics to the customers and / or buildings, characteristics being at least one of home location, work location, and visitation frequency; and compressing the resulting dataset to generate the combined dataset; and storing the combined dataset in a database using the corresponding geohash identification.
[0019] According to one aspect of the disclosed technology, there is provided a system and methods for modeling of data related to consumers and locations and for generating various consumer audiences based on collected and processed data and a user query. In at least one embodiment, the method comprises: receiving location data points related to the consumers; generating a partitioned combined dataset by: mapping a mobile advertiser identifier to a geofiles data in a geofiles database and to a geofence data in a geofence dataset based on a physical location identification for each location data point; mapping the mobile advertiser identifier to a company data in the point-of-interest database corresponding to the mobile advertiser identifier to generate a combined dataset; splitting the combined dataset with respect to geohashes and time periods; generating a partitioned combined dataset and storing it based on the corresponding geohash identification and the corresponding time periods.BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Further features and advantages of the present disclosure will become apparent from the following detailed description, taken in combination with the appended drawings, in which:
[0021] Figure 1A is a schematic block diagram of a system, in accordance with at least one embodiment of the present disclosure;
[0022] Figure 1 B is another schematic block diagram of the system of Figure 1A, in accordance with at least one embodiment of the present disclosure;
[0023] Figure 2 illustrates a method for modeling of data related to consumers and locations, in accordance with at least one embodiment of the present disclosure;
[0024] Figure 3 illustrates a method for modeling of data related to consumers and locations, in accordance with at least one other embodiment of the present disclosure;
[0025] Figure 4 illustrates other steps of the method for modeling of data related to consumers and locations, in accordance with at least one embodiment of the present disclosure;
[0026] Figure 5 illustrates another method for modeling of data related to consumers and locations, in accordance with at least one embodiment of the present disclosure; and
[0027] Figure 6 schematically illustrates initial audiences and a syndicated audience and corresponding extended audiences, in accordance with at least one embodiment of the present disclosure.
[0028] It will be noted that throughout the appended drawings, like features are identified by like reference numerals.DETAILED DESCRIPTION
[0029] Various aspects of the present disclosure generally address one or more of the problems of data collection, storage, search and retrieval of the data related to consumers. The present description provides a system and a method for manipulating the data related to consumers.
[0030] To advertise a product or a service to consumers, venders need to determine their target consumers and determine how to contact such target consumers. One of the methods of contacting the target consumers may be by using device identifications of the consumer devices (such as, for example, mobile phones, computer tablets, etc.).
[0031] In conventional methods of determining potential customer audiences and their corresponding device identifications, an area covered by a radius based a provided address in the middle is usually used. Such conventional approach does not distinguish among residential locations, work- related locations or just visiting locations for the consumer. In addition, conventionally, to enlarge or fin- tune the potential customer audience, a simple increase or decrease in the radius of the area is applied. Such conventional approach does not provide accurate and useful data and needs to be improved.
[0032] Methods and systems described herein may be used, for example, to determine the device identifications of the consumer devices as requested by a user (who may be, for example, the vender). The user may specify various criteria that may help define (shape) target consumers, for example, based on the consumer’s locations during specific periods of time. For example, customers that have visited a particular building during a weekend may be of particular interest to the advertiser. Methods and systems described herein may be also used for attribution.
[0033] The system as described herein partitions the data during the execution of the methods. In addition, the system and method implement extension technologies that is a more effective way to increase the scale of an audience with minimum loss of quality.
[0034] The system and methods as described herein generate higherquality of consumer audiences compared to the conventional methods known in the art. In the present technology described herein, transformation of data provided at a number of steps. Using past data, partitioning of the data based on the dates and geohash, and using the geohash to extract the data, help to significantly improve the quality of the consumer audience and the speed of its generation by the system and methods described herein.
[0035] Fig. 1A illustrates a system 100 for modeling of data related to consumers 127 and locations 101 , in accordance with at least one embodiment of the present disclosure. The system 100 comprises a server 110 and a database 115, and the server 110 has one or more processors 112.
[0036] The system 100 is configured to gather information about presence of consumer devices 125 of consumers 127 in various geolocations at various times, store this information and use this information on demand to generate various so-called “consumer audiences” for a user 107 (Fig. 1 B) based on a query data 108 received from a user device 105. Each generated consumer audience 129 may comprise a list of device identifications (also referred to herein as “device IDs”) of consumers 127 that respect the query data 108. A consumer device 125 (also referred to herein as “consumer’s electronic device 125”) may be a phone or, for example, an iPad or a personal computer or a similar electronic device. A user device 105 may be a phone or, for example, an iPad or a personal computer or a similar electronic device.
[0037] The server 110 is configured to receive the location data points 120 from the consumer devices 125. In at least one embodiment, each location data point 120 received by the server 110 for each consumer device 125 comprises the following data:
[0038] - a mobile advertiser identifier 161 (which may be also referred to as a “MAID 161” ora “maid161” or a “mobile advertiser ID 161” or an “advertiser ID 161” or a “unique device identifier 161”) that may be used to identify and target the consumer device 125 or a consumer 127;
[0039] - a location data point timestamp 162 (which may be also referred to as a “point timestamp”) which reflects the time when the location data point 120 was registered;
[0040] - a physical location identification 163 (also referred to herein as an “identification of the physical location” and “physical location ID 163”) where the consumer device 125 was seen (detected) within a pre-determined accuracy, which may be transmitted, for example, as latitude and longitude, and which may have, for example and without limitation, 5 meters to 15 meters accuracy; and
[0041] - an internet protocol (IP) address 164 (also referred to herein as an “IP address”) of the consumer device 125 at the moment when the location data point 120 is transmitted to the server 110.
[0042] For example, the MAID 161 may be used for identification and addressing or targeting the consumers 127 with advertisement through the third-party Demand Side Platform (DSP). For example, the MAIDs 161 may be different for the same consumers 127 and for the same consumer device 125.
[0043] The physical location identification 163, comprising geolocation data of the consumer device 125, which indicates geolocation data of the consumer 127, may be identified at and by the consumer device 125 and then transmitted to the server 110 as a part of location data points 120. Such physical location identification 163 may be generated, for example, using information of a global positioning system (GPS), as well as barometers, accelerometers and / or other sensors of the consumer device 125. Each physical location identification 163 comprises data about the geolocation of the consumer device 125. The physical location identification 163 may be expressed, for example, as a geographical position (which is expressed at least using longitude and latitude) of the consumer device 125, and may alsocomprise an altitude information, which may allow to distinguish between different levels of the same buildings, for example.
[0044] For example, the server 110 may receive, on average, 11 billion (B) location data points 120 per day from across Canada and the United States. For example, 11 B location data points 120 per day may result in 1 Terabytes (TB) of data per day. For example, the accumulated data by the server 110 may be more than 1.6 Petabytes (PB) and all this data need to be stored on the servers 110 and / or databases.
[0045] Referring to Figs. 1A and 1 B, the location data points 120 are stored in a location data point dataset 122 of the database 115. To determine the consumer audiences when the query data 108 is received from the user device 105 (Fig. 1 B), the processor 112 needs to search, on demand, through the data (or at least a portion of the data) of the location data point dataset 122 that is stored in the database 115 about the consumers 127. When extracting data from any database using conventional methods and systems, the cost associated with such an operation revolves around the amount of data the processor needs to “look at” (find and analyse) prior to extracting the information the operator or the user is looking for. This cost can be considered at an average of 5$ / TB of data the system has to crawl. This means that the cost of crawling through the data and, for example, extracting the MAIDs 161 related to the consumers 127 that were at, for example, a major sports center for the past four years would be so high that the value of the accumulated data would become completely worthless if the data searching and the storage of the data is provided by the conventional methods. The system 100 (Figs. 1A, 1 B) and methods 200, 300, 400 (Figs. 2-4) of storing the data, partitioning the data and extracting the data in a cost-effective way are provided herein.
[0046] In at least one embodiment, the following parameters may be used by the system 100 when searching: the size of the partitioned combined dataset 138 described herein; the query data 108 (which comprises the requested location address 109a and the audience refining parameters 109b) provided to the system 100 by the user 107; and the number of files the system 100 needs to search through (the partitioned combined dataset 138), based on the storage technology. To take into account all these parameters, the system 100 determines a data pipeline (in other words, a set of steps to be executed). For example, to implement the methods as described herein, the system 100 may use, for example, technological stacks such as, for example, Amazon Athena™ (which is a serverless, interactive analytics service built on open-source frameworks, supporting open-table and file format) and Google BigQuery™ (which is a serverless, cost-effective and multi-cloud data warehouse designed to help turn big data into valuable business insights). For example, the system may use a series of queries either in Athena™ or BigQuery™ and, in some embodiments, transfer data between these two services during the implementation of the methods as described herein.
[0047] To reduce search times in response to the query data 108 and reduce the cost of generating audiences, the received location data points 120 are processed by the processor 112. The processor 112generates another dataset which is referred to herein as a “combined dataset 130”. The combined dataset 130 provides additional data describing, for example, at which particular building, business or organization the consumer 127 and the consumer device 125 were located based on their physical location identification 163. Referring to Figs. 1A-1 B, the combined dataset 130 stores, for each location data point 120, links to the location data points 120 of the location data point dataset 122, and links to their corresponding geofiles data from a geofiles database 131 , corresponding geofence data from a geofence dataset 133, and corresponding company data of a point-of-interest dataset 135. For example, the combined dataset 130 may be a set of links (also referred to herein as “references”) to the MAI Ds 161 of the location data point dataset 122, to their corresponding geofiles data from the geofiles database 131 , corresponding geofence data from the geofence dataset 133, and corresponding company data of the point-of-interest dataset 135.
[0048] In at least one embodiment, the processor 112 uses portions of the data to be searched, referred to herein as “combined dataset partitions 137” (Fig. 1A), in order to be able to effectively query and search later, on the request of the user 107, through all the collected and recorded data related to the location data points 120 and through the combined dataset 130 at a reasonable cost. To significantly reduce the search time in the collected and recorded data later, when the user query 108 is received, the processor 112 of the system 100 assigns a geohash (preferably, a three-character geohash) to each location data point 120.
[0049] The system 100 uses a geohash method to express the location of a building or a location (geographical position) of the consumer device 125. The geohash method is an encoding method used to encode geographic coordinates (latitude and longitude) into an alphanumeric string of digits and letters. Each geohash corresponds to one location cell (or area or quadrant) on a geographical map. In at least one preferred embodiment, the geohash (also referred to herein as a “geohash label”) used in the present disclosure has three characters (in other words, the geohash has a precision or accuracy of three (3), or, in other words, the geohash has a length of 3, or a “geohash length” is 3) and is referred to herein as a “three-character geohash”. One three-character geohash corresponds to an area of approximately 156.5 kilometers (km) by 156 km. Each additional character of the geohash adds precision to the location cell. The three-character geohash may be preferably used in the execution of the methods as described herein because the three-character geohash has been shown to provide enough precision while providing enough data for the extension modeling as described herein. Expressing a location using the geohash method refers herein to any geographical location. For example, any building in Montreal may be represented (expressed) with “f25” as this is the three-characters geohash that represents the city and it’s surrounding.
[0050] Referring to Fig. 2, at step 202, the location data points 120 are received by the system 100. At step 204, after the system 100 receives the location data points 120, the location data points 120 are assigned a corresponding geohash based on the physical location identification 163.
[0051] In at least one embodiment, the processor 112 generates a combined dataset 130 which comprises references to the location data points 120 in the location data point dataset 122 and to data corresponding to each location data point 120 in various databases, such as, a geofiles database 131 , a geofence dataset 133, point-of-interest dataset 135. In at least one embodiment, the location data points 120 or references (links) to the location data points 120 in the combined dataset 130 may be sorted (split) with reference to different geohashes of the pre-determined precision (preferably, three-character geohashes).
[0052] In at least one embodiment, the processor 112 of the system 100 uses the geofiles database 131 to generate the combined dataset 130. The geofiles database 131 has the geofiles data which comprises census divisions (which may be, for example, census block groups in the US or a census dissemination area as in Canada, or similar divisions in other jurisdictions), corresponding zip code(s), a province or a state, and the corresponding country. For example, the geofiles database 131 may be received from Statistics Canada and / or United States Census. Using the geofiles data of the geofiles database 131 , the processor 112 is configured to link each MAID 161 of each particular location data point 120 to a corresponding geofiles data of the geofiles database 131 which is related to that particular location data point 120 based on the corresponding physical location identification 163 of the location data point 120. In other words, the geofiles database 131 is used by the server 110 to link the MAID 161 of the location data point 120 to the corresponding geofiles data based on the physical location identification 163. In the combined dataset 130, the processor 112 stores links to the MAIDs 161 of the location data point dataset 122 and links to geofiles data of the geofiles database 131 , which correspond to each one of MAIDs 161 . The link to the geofiles data may thus help to identify, on demand, the census division, the zip code, the province or state and the country which correspond to a particular MAID 161 .
[0053] To generate the combined dataset 130, the processor 112 of the system 100 also uses a geofence dataset 133 (which may be also referred to as a “buildings’ contours dataset 133”) which comprises building geofences and corresponding building references for each building in a concerned area. The building geofence is a polygon (so-called “geofence”) representing geographical building boundaries. The building reference may comprise a name of a company occupying the building and an address of that building. For example, the geofence dataset 133 may be generated by the server 110 from satellite imagery. For example, the concerned area may be the north America including the US and Canada. Thus, a building geofence may be determined (generated) and stored in the geofence dataset 133.
[0054] The geofence dataset 133 provides contours of various buildings and permits the server 110 to map the location data points 120 to corresponding building geofences and therefore to the corresponding building references. Using the data in the geofence dataset 133, the processor 112 on the server 110 may map each MAID 161 and the corresponding location data point 120 to a particular building using the physical location identification 163 of the location data point 120 and the corresponding buildinggeofence from the geofence dataset 133. The processor 112 may then determine the corresponding building reference and map it too to the MAID 161. In other words, the server 110 (the processor 112) using the geofence dataset 133, is configured to map (attach or link) each location data point 120 received with MAID 161 to the particular building of a set of buildings provided for the concerned area in the geofence dataset 133.
[0055] To generate the combined dataset 130, the processor 112 also uses the point of interest (POI) dataset 135. After the processor 112 of the server 110 has mapped the MAID 161 to the particular corresponding building, the server 110 uses company data in the POI datasets 135 to attach that particular corresponding building to a company data of the point-of-interest dataset 135 which may comprise, for example, a brand name or a company name. For example, the point-of-interest dataset 135 may comprise points of interest such as, for example, stores, restaurants, companies, public buildings, etc.
[0056] For example, the server 110 may identify that the consumer 127 visited the particular building and therefore the consumer device 125 was localized in that particular building as described above, and then the processor 112 on the server 110 determines that that the particular building has a business located therein (such as, for example, a franchise of one of well-known fast-food chains) by mapping the building reference associated to (already mapped to) the MAID 161 to the corresponding company data in the POI dataset 135. For each reference to the MAID 161 in the location data point dataset 122, the processor 112 generates, in the combined dataset 130, a link (a reference) to the corresponding company data in the point-of-interest dataset 135.
[0057] Referring to Fig. 2, at step 204, geofiles database 131 is used to map the location to a building. The processor 112 maps the physical location ID of the location data points 120 with a business ID provided in the geofiles database 131.
[0058] At step 206, a point-of-interest (POI) database 135 is used to map the building that was already determined to a company name and, in some embodiments, a brand name and / or a business. The POI database 135 comprises a latitude and longitude. In addition, the POI dataset 135 comprises data having stores, restaurants, companies, public buildings, etc. The processor 112 maps this point on a map to a building that corresponds to a building polygon.
[0059] In at least one embodiment, the geofiles database 131 , the geofence dataset 133, the point- of-interest dataset 135 are partitioned prior to the execution of other steps of the methods as described herein, and the partitioning is based on the geohash, preferably the three-character geohash. This significantly helps to search the data later when a request with the query data 108 from the user 107 is received, for example.
[0060] In at least one embodiment, the combined dataset 130 comprises links (references) to the MAIDs 161 in the location data point dataset 122, and, in some embodiments, links to the geofiles data in geofiles database 131 corresponding to the MAIDs 161 , geofence data in the geofence dataset 133corresponding to the geofiles data, and the company data in the point-of-interest dataset 135 corresponding to the geofence data. Each link may comprise a database-specific key that links all of these datasets 122, 131 , 133, 135: location data point dataset 122, geofiles database 131 , geofence dataset 133, point-of-interest dataset 135. If any additional information may be needed later for a particular MAID 161 from any of these datasets, the processor 112 may use the database-specific key to access the corresponding database. In at least one embodiment, the database-specific key provides a link between corresponding data points in each dataset of the datasets 122, 131 , 133, 135.
[0061] In at least one embodiment, the system 100 partitions the data received for the location data points 120 in the location data point dataset 122 based on the three-character geohash and the date that is based on the timestamp 162 of each location data point 120.
[0062] In at least one embodiment, the processor 112 partitions the data in the combined dataset 130 into a set of combined dataset partitions 137 (also referred to as a whole as “partitioned combined dataset 138”) each corresponding to one three-character geohash and a date or another pre-determined partition-related period of time. For example, the combined dataset 130 may be partitioned (split) into a first combined dataset partition (related to a first geohash and a first day), a second combined dataset partition (related to a second geohash and the first day), ... up to a dataset partition related to n-th geohash and m-th day. In other terms, the processor 112 splits data in the combined dataset 130 with respect to geohashes and time periods to generate a set of the combined dataset partitions 137. The processor 112 then compresses each one of the combined dataset partitions 137 to generate a plurality of partitioned combined datasets and stores the partitioned combined dataset 138 in a database using (based on) the corresponding geohash identification. This partitioned combined dataset is essential for matching the buildings, POI locations and MAID 161 and storing the relationship between all these different entities. By splitting the data in the combined dataset 130 at step 208 with respect to quadrants identified as the geohashes, the system 100 greatly reduces the amount of data it needs to search within when query data 108 is received and helps to effectively query the data. In at least one embodiment, the processor 112 of the system 100 separately goes over a maximum of 100 geohashes at the time, in order to not write more than 100 partitions per query.
[0063] Still referring to Fig. 2, at step 210, the file is compressed (concentrated) to generate a partitioned dataset to reduce the number of files the system needs to go through for each query. At step 212, the processor 112 of the system 100 deletes all temporary files, and the partitioned datasets (dataset partitions) are prepared for data modelling and further processing.
[0064] In at least one embodiment, the processor 112, prior to splitting the combined dataset 130, assigns additional characteristics to the consumer device 125 related to the location data points 120 collected. The additional characteristics may be, for example, a customer home location, a customer work location. The additional characteristic may be, for example, a visitation frequency which corresponds toa frequency that the same consumer device 125 has been detected in the same building and / or business (company), for example.
[0065] The system 100 and the combined dataset 130 generated by the system 100 may be used as follows. As illustrated in Fig. 1 B, in response to receiving a query data 108 which comprises a requested location address 109a of a requested location and the audience refining parameters 109b received from a user 107 of the system 100, the system 100 generates an extended consumer audience dataset 142 (which may be, for example, a list of device IDs of consumers 127 that respect the query data 108) as described herein.
[0066] First, the system 100 matches the requested location address 109a to one geolocation polygon (which may be also referred herein as a “query-related polygon 172”) using the buildings and geolocation data of the geofence dataset 133 as discussed above. In other words, the processor 112 maps or determines (for example, by a mapping routine 152), a query-related polygon 172 based on the requested location address 109a. In at least one embodiment, the requested location address 109a may comprise a set of addresses or geolocations of corners of the query-related polygon 172. The requested location address 109a may comprise a central address of the query-related polygon 172 and a user- defined radius (provided by the user 107 or selected from suggested radius choices by the user 107 when the query data 108 is entered by the user 107) of the query-related polygon 172 for the processor 112 to determine input polygon boundaries of the query-related polygon 172 on its own. In at least one embodiment, the user 107 may trace input polygon boundaries of the query-related polygon 172 on a portion of a map (also referred to herein as a “map portion”) displayed on the user device 105. The processor 112, after receiving the input polygon boundaries of the query-related polygon 172 from the user device 105 may determine additional points of query-related polygon boundaries of the query-related polygon 172.
[0067] In at least one embodiment, the processor 112 may import or trace the determined query- related polygon 172 related to (mapped to) the requested location address(es) 109a of the requested location. A set of requested polygon addresses may be generated based on the geofiles database 131 , preferably based on pre-partitioned geofiles database 131. The processor 112 may map the requested location address 109a to a query-related polygon 172. For example, the processor 112 of the system 100 may use an automated method to extract a set of building contours related to the query-related polygon 172 based on a satellite image.
[0068] Based on the determined query-related polygon 172 and corresponding query-related polygon boundaries, the system 100 may extract all the MAIDs 161 (also referred herein as “polygon MAIDs”) of the consumer devices 125 that were detected (in other words, identified or determined to be located) inside the query-related polygon 172 during a predetermined time period. Such pre-determined time period may be provided in the query data 108 (for example, in the audience refining parameters 109b provided by the user 107 or selected by the user 107) when received from the user device 105 ormay be pre-determined by the processor 112 or an operator of the system 100. In at least one embodiment, the polygon MAIDs are related to the query-related polygon 172 and the audience refining parameters.
[0069] In at least one embodiment, the processor 112 of the system 100 may also determine a list of geolocations within the query-related polygon 172 (which may comprise, for example, zip addresses and census locations) using the geofiles database 131 and extract all MAIDs 161 related to consumers (which may be referred to herein as “concerned consumers”) living within the query-related polygon 172 (having a home address and / or being located during night time hours).
[0070] In at least one embodiment, the processor 112 may first generate an existing audience 140 (also referred to herein as an “initial audience 140”) of consumers 127 having the consumer device 125 detected within the query-related polygon 172 during the predetermined time period or at any time provided by the user 107 and being recorded in the location data point dataset 122. For example, the initial audience 140 may correspond to the polygon MAIDs which may be all MAIDs 161 of the consumer device 125 that were detected in a building related to a fast-food chain locations.
[0071] In at least one embodiment, the initial audience 140 may comprise MAIDs 161 of the consumer devices 125 which were detected within the query-related polygon 172 during a pre-determined part of the day (for example, daytime or nighttime) and pre-determined query-related time period (for example, pre-determined or user-specified number of hours and / or minutes), and such record may be found in the location data point dataset 122. For example, a query-related time period may be a day-time period, a nighttime period or a visitor time period which may comprise shorter term periods of time compared to the daytime period or the nighttime period. The initial audience 140 may be a residential initial audience 140a when the query-related time period is the nighttime period, a work initial audience 140b when the query-related time period is the daytime period, or a visitor initial audience 140c when the query-related time period is the visitor time period. The query-related time period may be provided in the query data 108 or determined by the processor 112 based on the query data 108.
[0072] In at least one embodiment, based on the initial audience 140 generated previously, such as the residential initial audience 140a, work initial audience 140b, and / or visitor initial audience 140c, the processor 112 may generate the extended audience 142.
[0073] In at least one embodiment, the processor 112 of the system 100 may extend (increase) such initial audience 140 using an extension determining method (which may also be referred to as “extension model”) to generate a so-called extended audience 142. For example, the extended audience 142 may be generated based on the initial audience 140 and the audience refining parameters 109b. For example, the 109b may comprise a requested audience type and additional audience data. For example, the processor 112 may generate additional audience conditions based on the audience refining parameters 109b. The audience refining parameters 109b as referred to herein comprise parameters that determine whether a home-extension routine, a work-extension routine, or a social-extension routine areused, the time periods, and the pre-determined attendance frequency and attendance percentage. In at least one embodiment, based on the audience refining parameters 109b, the processor 112 determines whether to use a home-extension routine, a work-extension routine, or a social-extension routine and the additional audience conditions. In at least one embodiment the audience refining parameters 109b may be determined by the user 107. In at least one embodiment, the audience refining parameters 109b may be provided on a screen of the user device 105 for the user 107 to select and be transmitted to the server 110 as a part of the 108.
[0074] Referring to Fig. 6, modelling so-called “extensions” for the initial audience 140 having initial MAIDs 161 i may comprise generating a home extension corresponding to an extended home-based audience 260a, a work extension corresponding to an extended work-based audience 260b, and / or a social extension corresponding to an extended social-based audience 260c. In at least one embodiment, the extended audience 142 may comprise, in addition to the MAIDs 161 of the initial audience 140 (initial MAIDs 161 i), additional MAIDs which are identified by the processor 112 as follows. For example, in the combined dataset 130 or partitioned combined dataset 138, the system 100 has data not only where, when and for how long the MAIDs 161 were detected during the day, but also where the MAIDs 161 were detected during the night, and therefore in addition to home initial audience 140a and work initial audience 140b, the work-extension routine and the home-extension routine, respectively, may determine the location of a consumer work address or a consumer residence address, respectively, based on the amount and the period of time spent in locations other than locations used to generate the home initial audience 140a and work initial audience 140b, respectively. The work-extension routine and the homeextension routine may then determine extended MAIDs which may be work-extended MAIDs 261 b or home-extended MAIDs 261a, respectively, by retrieving them from the partitioned combined dataset 138 or the combined dataset 130 based on the work consumer address or the residential consumer address, respectively.
[0075] In at least one embodiment, based on the residential initial audience 140a or the work initial audience 140b, the social-extension routine is configured to generate the extended social-based audience 260c which comprises, in addition to the initial residential MAIDs 161a or initial work MAIDs 161 b, social-extended MAIDs 261 c. The social-extended MAIDs 261c may correspond to MAIDs of the combined dataset 130 or partitioned combined dataset 138, other than the initial residential MAIDs 161a or initial work MAIDs 161 b, respectively, which were detected at the same location and at the same moment as the initial residential MAIDs 161a or initial work MAIDs 161 b, respectively.
[0076] In at least one embodiment, for each initial MAID of the initial MAIDs, using the partitioned combined dataset 138, the social-extension routine of the processor 112 may determine a building geofence related to a location of the initial work MAID during a first pre-determined time period, the initial MAIDs being the polygon mobile advertiser identifiers; extract, from the partitioned combined dataset 138, social-extended MAIDs, different from the initial MAIDs, detected within the building geofence of theinitial MAIDs during a second pre-determined time period; and generate an extended social-based audience having the social-extended MAIDs and the initial MAIDs of the initial audience.
[0077] In at least one embodiment, using the MAIDs 161 (initial MAIDs 161 i) of the work initial audience 140b, the processor 112 may determine consumer residence addresses (or consumer residence buildings) of the consumers 127 corresponding to the initial MAIDs 161 i by determining, from the combined dataset 130 or the partitioned combined dataset 138, a location where the initial MAIDs 161 i were detected during, for example, a night time.
[0078] For example, when the work initial audience 140b relates to one or more ski stations, the extended home-based audience 260a (also referred to herein as “home-extension audience 260a”), in addition to initial work MAIDs 161 b of the work initial audience 140b, also comprises so-called “home- extended MAIDs 261a”. Each home-extended MAID 261a corresponds to a MAID 161 from the combined dataset 130 or the corresponding partitioned combined dataset 138, other than the initial work MAIDs 161 b, that were detected at the location within the same residential building geofence as the residential building geofence of the corresponding work MAID 161 b. The residential building geofence may be determined for each initial work MAID 161 b. The residential building geofence corresponds to a location where the initial work MAID 161 b was detected during an extended night-time period (which may be predetermined and may be, for example, between 7PM and 7AM).
[0079] In at least one embodiment, for each initial work MAID 161 b, using the partitioned combined dataset 138, a residential building geofence related to a location of the initial work MAID 161 b during an extended night-time period may be determined. Then, from the partitioned combined dataset 138, home- extended MAIDs 261 b, different from the initial work MAIDs 161 b, and detected within the residential building geofence of the initial work MAIDs 161 b may be extracted. The generated extended home-based audience 260a may thus comprise the home-extended MAIDs 261a and the initial work MAIDs 161 b of the work initial audience 140b.
[0080] Similarly, when the residential initial audience 140a relates to one or more dwellings, the extended work-based audience 260b, in addition to the initial residential MAIDs 161a of the residential initial audience 140a, also comprises so-called “work-extended MAIDs 261 b”. Each work-extended MAID 261 b corresponds to a MAID 161 from the combined dataset 130 or the partitioned combined dataset 138, other than the initial residential MAID 161a, that were detected at the location within the same work building geofence as the work building geofence of the corresponding residential MAID 161a. The work building geofence may be determined for each initial residential MAIDs 161 a. The work building geofence corresponds to a location where the initial residential MAID 161 a was detected during an extended daytime period (which may be pre-determined and may be, for example, between 7AM and 7PM).
[0081] In at least one embodiment, for each initial residential MAID 161 a, using the partitioned combined dataset 138, a work building geofence related to a location of the initial residential MAID during an extended day-time period may be determined. Then, from the partitioned combined dataset 138, work-extended MAIDs 261 b, different from the initial residential MAIDs 161 a, and detected within the work building geofence of the initial residential MAIDs 161a may be extracted. The generated extended workbased audience 260b may thus comprise the work-extended MAIDs 261 b and the initial residential MAIDs 161 a of the residential initial audience 140a.
[0082] Forexample, the initial residential MAIDs 161a correspond to the MAIDs 161 of the consumer device 125 which were detected within the query-related polygon 172 during a pre-determined part of the day (for example, daytime or nighttime) and pre-determined time period (for example, pre-determined or user-specified number of hours and / or minutes), and such record may be found in the location data point dataset 122 and therefore a link to such MAID 161 may be found in the combined dataset 130.
[0083] In at least one embodiment, to determine the extended audience 142, the processor 112 executes a home-extension routine (which may be also referred to as an “address-based model”) which, in order to generate the extended audience 142 based on home-extension requirements, determines and adds to the data in the initial audience 140 all MAIDs 161 that correspond to the consumer devices 125 and the consumers 127 that were detected at the same address (visiting consumers) at the specific location during a nighttime (or another specified period of time), and that were previously received by the processor 112. Such extended audience 142 may be referred to as the extended home-based audience 260b. In at least one embodiment, MAIDs 161 are collected and used for determination of the MAIDs visited the address, as the MAIDs 161 may be changed for the same consumer device 125 and / or the same consumers 127.
[0084] In at least one embodiment, to determine the extended audience 142, the processor 112 executes a work-extension routine which, in order to generate the extended audience 142, increases the initial audience 140 by adding to the initial audience 140 all MAIDs 161 that correspond to the consumer devices 125 detected at an address within the query-related polygon 172 during the working hours (which may be pre-determined or provided with the query data 108).
[0085] In at least one embodiment, to determine the extended audience 142, the processor 112 executes a social-extension routine which increases the initial audience 140 by generating and adding, to the data provided in the initial audience 140, a social-extension audience which comprises of acquaintances devices IDs of acquaintances devices 126 of acquaintances 128 of the customers 127 identified in the initial audience 140.
[0086] The system 100 may also adjust (in other words, customize) the generated (determined) extended audience 142 (or, in some embodiments, the initial audience 140) to generate a targeted extended audience 145 by using a building attendance frequency (also referred to herein as “attendance frequency”) which is the frequency of attending one or more buildings by the customers 127 (members of the initial audience 140). The building attendance frequency corresponds to the number of times the consumer device 125 of the consumer 127 was detected at that building, based on the data collected previously by the processor 112 (and stored, for example, in the partitioned combined datasets 138).Taking into account the building attendance frequency may help to reduce an audience volume (audience size). In other words, the amount of mobile advertiser identifiers in the targeted extended audience 145 (so-called “targeted audience volume”) generated by the processor 112 may be smaller than the amount of mobile advertiser identifiers in the extended audience 142 (so-called “extended audience volume”) and / or an amount of mobile advertiser identifiers in the initial audience 140 (so-called “initial audience volume”), and thus, such that the volume of data in the generated targeted extended audience 145 (which may be also referred to as a frequency-adjusted audience) is smaller than the volume of data in the initial audience 140. The building attendance frequency may be determined and compared to a pre-determined building attendance frequency range to determine whether the building attendance frequency is withing the pre-determined building attendance frequency range. The targeted extended audience may be generated by reducing the initial audience 140 based on a building attendance frequency determined by the processor 112 and related to each one of the polygon mobile advertiser identifiers determined by the processor 112.
[0087] In some embodiments, the extended audience 142 (or, in some embodiments, the initial audience 140) may be adjusted (and, in most cases, reduced) by taking into account not only the building attendance frequency (frequency of attending the building or the location which is located in the query- related polygon 172), but also a building attendance percentage which is a percentage of time per day (and / or per month and / or year) spent by each consumer at a particular location.
[0088] Using the abovementioned methods, the volume of the initial audience and / or extended audience generated by the processor 112 may be, on one hand side, increased, and, on another hand side, refined and scaled down and narrowed when the targeted extended audience 145 is generated according to (in other words, using or based on) the audience refining parameters 109b received from the user device 105 and selected or provided by the user 107 earlier. In other words, the targeted audience volume of the targeted extended audience 145 may be larger or smaller than the extended audience volume of the extended audience 142, depending on the audience refining parameters 109b.
[0089] In order to optimize the query process, the processor 112 builds partitions in the data that corresponds to the first three characters of a geohash. The geofiles dataset 131 and geofence 133 may be also partitioned in a similar way with reference to locations of the buildings. In at least one embodiment, the label corresponding to the three-character geohash of the geographical location of the building may be mapped to geographical locations of various datasets, such as, for example, geofiles database 131 , geofence dataset 133, point-of-interest dataset 135 and the combined dataset 130. In at least one embodiment, these datasets 131 , 133, 135, 130 may be partitioned based on the geohash labels as described herein.
[0090] The methods and system as described herein permit finding and unwrapping (unpacking) the portion of the data that is needed in order to find the information needed to generate the combined dataset 130, initial audience 140, extended audience 142, targeted extended audience 145. There is no need tounpack all the available data to find the information. Every day, home-extension routine may be executed by the processor 112 of the system 100 to generate a home-extension audience. The processor 112 of the system 100 thus generates a combined dataset partitions that is a combined dataset partitioned with respect to three-character geohashes and time periods and that allows generating a specific targeted extended audience 145 and transmit to the user 107 in response to the user query 108 at a much faster and cheaper rate than using the conventional systems and methods known in the art.
[0091] When the system 100 needs to extract data that was previously collected and processed by the system 100, only data related to a particular partitioned dataset (also referred to herein as a “query- related combined dataset partition”) related to a particular requested polygon address provided in the query data 108 or a set of requested polygon addresses (related, for example, to a particular business or a brand or a keyword) located within the query-related polygon 172, and therefore to the particular geohash(es) corresponding for those addresses, is extracted. The processor 112 of system 100 as described herein does not need to extract all the data collected previously by the processor(s) 112, but only the data related to geohash(es) related to the addresses (locations) of the query-related polygon 172 of the user’s query 108.
[0092] The system 100 and the methods 200, 300, 400 as described herein use the three-character geohash (geohash3) to store and retrieve the geolocation data. The system 100 and the methods 200, 300, 400 combine the satellite imagery data, used to generate the geofiles database 131 , to trace buildings and POI data in order to add context, such as company names and brands, to these buildings. The system 100 and the methods as described herein permit to reduce the numberof files and the storage size needed to store the data collected due to compressing (packing) and concatenating (linking together in a chain or series) the initial audiences 140, extended audiences 142, and / or target audiences 145 generated in response to the previous queries and query data 108 received.
[0093] Fig. 3 illustrates a method 300 for raw data ingestion, in accordance with at least one embodiment of the present disclosure. At step 302, the system 100 receives location data points 120 daily. The processor 112 system 100 receives the raw data from the consumer device 125 or, in some embodiments, another system 100 which collected data. The processor 112 of the system 100 may wait, for example, until all the data was received to start its daily processing.
[0094] At step 304, the datasets, such as location data point dataset 122 are partitioned based on geohashes (for example, the three-character geohash which is also referred to herein as “geohash3”) that correspond to each physical location ID 163 of the location data points 120 to generate location data point dataset partitions 123 (Fig. 1 B). In other words, the location data point dataset 122 is partitioned (separated) at the geohash level to generate smaller, partitioned datasets that are more convenient to be identified based on the geohash and searched. Such partitioning of the location data point dataset 122 helps improving the efficiency of processing and searching the data later, in other words to improve the efficiency of the “processing pipeline”. In at least one embodiment, the data of the location data pointdataset 122 may be then transferred to a data warehouse (which may be, for example, BigQuery™ enterprise data warehouse which is a fully managed and completely serverless enterprise data warehouse of Google Cloud™) at step 306, where it is later processed to generate various audiences, such as, for example, extended audiences, syndicated audiences, and POI extended audiences.
[0095] At step 308, the location data points 120 collected for a pre-determined period of time (such as a day, for example) is mapped to specific buildings available from the geofiles database 131. Every location data point 120that is related to the consumer device 125 that was detected to be located inside one building, are attached to that one building, along with the number of times detected to be located in that building and timestamps related to each detection of the consumer device 125 in that building.
[0096] At step 310, the mapping associations between the location data points 120 and unique building identifications are stored. For example, the processor 112 may store in the combined dataset 130, for each location data point 120, a unique building identification, a MAID and a corresponding quality score. The quality score as referred to herein is a score measured on a scale of 1 to 100 and calculated based on a number of location data points 120 related to one MAID 161 detected, along with a percentage of their time spent in one location. For example, the quality score may be determined as an average of location data points 120 per day per identifier (MAID 161) in a specific country, and increased or reduced (adjusted) depending on whether that one specific MAID 161 was detected in the specific building more or less frequently than that average (of the location data points 120 per day per MAID 161 in the specific country). For example, the more frequently the MAID 161 was detected within one specific building, the higher may be the number of points 120 related to one MAID 161 , the higher may be the quality score.
[0097] In at least one embodiment, the processor 112 may store links or references to the MAID 161 in the location data point dataset 122, the corresponding unique building identification in the geofence database 133, and the corresponding quality score. At step 310, the system 100 optimizes and stores the results of the step 308 into a temporary table.
[0098] For example, buildings may be assigned different types based on time periods when each MAID 161 was detected in that building. For example, a type of the building may be determined as a residential type of the building, an office type of the building or a point-of interest location building.
[0099] At step 312, home location labels may be assigned to each building of the buildings where the consumer device 125 were detected earlier as provided in the location data point dataset 122, based on a building time period corresponding to the time period spent by the consumer 127, and therefore time period (duration) of the consumer device 125 being detected continuously (during each visit of the consumer 127) in each building. Each visit of consumer 127 to a particular building corresponds to the “device detection instance” and then each device detection instance may have a corresponding “building time period” (that may be obtained by using a location time period minus an initial time of the visit). Based on the timestamp and duration of the building time period, and a type of the building, the assigning routine 153 of the processor 112 of the system 100 associates the MAID 161 with a home location and adds ahome location label, and stores this association along with a quality score in a database, such as the combined dataset 130. The type of the building may be determined, for example, by comparing it with the point-of-interest dataset 135. For example, if the processor 112 can associate the building to a POI location, such as building is determined as a POI building, otherwise, the building is determined to be a residential building or for other purposes.
[0100] At step 314, work location label may be assigned to each building based on the time spent by the MAID 161 and therefore the consumer consumers 127 and / or consumer device 125 in each building. Based on the timestamp and the type of the building related to the location of the MAID 161 and therefore the consumers 127 and / or consumer device 125, the MAID 161 is associated with (mapped to), by the processor 112, a home location label and this association is stored along with a quality score in the combined dataset 130. For example, a link or a reference may be stored in the combined dataset 130
[0101] At step 316, visitation labels are assigned based on the visits of the MAID 161 , and therefore the consumers 127 and / or consumer device 125, to each building. For example, to generate the visitation labels, the processor 112 may analyse number of visits, time of visits, number of signals per visit, where the visit is the instance when the MAID 161 was detected in the corresponding location, and, for each MAID 161 , generate a quality score of the association between the building and the MAID 161.
[0102] The system 100 stores every association of the MAID 161 to the corresponding building that is stored within the point-of-interest (POI) dataset 135 and stores the results in the combined dataset 130. This combined dataset 130 may then be used to generate the syndicated audiences.
[0103] At step 318, the system 100 deletes temporary data and stores only associations in the combined dataset 130. The system 100 deletes every temporary table generated to complete the steps 306-316 as these temporary tables are not needed for further processing.
[0104] In at least one embodiment steps 322-334 may be also executed, as illustrated in Fig. 3. At step 322, data received at step 302 is transferred to an interactive query service such as, for example, Amazon Athena (which is an interactive query service that makes it easy to analyze data directly in Amazon S3 using standard Structured Query Language (SQL)). The data may be loaded into the interactive query service (for example, Amazon Athena) as, for example, a temporary table in order to store it for custom audience generation later in response to receiving the query data 108.
[0105] At step 324, the system 100 generates empty partition tables that are then used for data processing later. These temporary tables are then used to execute steps 326-334.
[0106] At step 326, the processor 112 processes over 1000 geohashes with an increment of 100 and writes the location data points 120 into the partition tables generated at the previous step (step 324). In at least one embodiment, the system 100 can only insert 100 partitions in the partitioned combined dataset with every single query. For this reason, the system 100 processes data for over 1000 geohash across multiple queries to overcome this limitation.
[0107] At step 328, the system 100 combines and compresses the files generated at step 326. At step 326, the processor 112 may generate over 40000 files. The system 100 reduces the number of files and compresses the data to reduce the cost of extraction. This is done by combining the files, converting them to a parquet format and then compressing them using snappy format. The generated compressed data files are then stored and ready to be discovered by a crawler.
[0108] At step 330, the system 100 runs the crawler (in other terms, executes a crawler routine 151) to identify the new partitions generated at step 328. The crawler routine 151 analyses the combined and compressed partition tables to identify these newly generated partitions.
[0109] At step 332, the processor 112 of the system 100 executes a check routine 155 to determine whether the number of points received in the raw file (at step 302) matches the number of points in the final processed files. The processor 112 of the system 100 then executes a validation routine 156 to verify (validate) that the same number of points have been received in the initial raw data with the location data points 120 and in the output of the steps described above. At step 334, all temporary data is deleted. Once the validation is successful and the validation routine 156 has been successfully executed, the system 100 removes all the temporary tables and resources generated, in order to reduce the storage costs. In at least one embodiment, the processor 112 may then proceed to step 304 as illustrated in Fig. 3.
[0110] Fig. 4 illustrates a method 400 for data extraction and generation of a custom audience, in accordance with at least one embodiment of the present disclosure. At step 401 , the user 107 requests a new custom audience and a query data 108, which comprises the requested location and the audience refining parameters, is collected from the user device 105 and received by the server 110. For example, the audience refining parameters may comprise a frequency and date ranges. For example, the user 107 may request such custom audience through a self-serve platform or via an email. For example, the server 110 may receive such an email and determine the query data 108 automatically from that email.
[0111] The query data 108 may comprise the requested location address 109a, the audience refining parameters 109b, such as, for example, the audience type, which may be received through the data processing pipeline. The audience type may comprise, for example, a request for a simple audience extraction and therefore generation of the initial audience 140, an extraction based on home or work locations and therefore generation of the extended audience 142. The initial audience 140 and the extended audience 142 comprise corresponding lists of MAIDs 161.
[0112] At step 403, based on the requested audience type and other audience refining parameters 109b received, the processor 112 determined a data pipeline to be used based on pre-determined rules. For example, the pipeline may specify the queries and routines to run and depends on the type of audiences. Different pipelines relate, for example, to a simple extraction, home or work extraction to generate the extracted audiences, and execute different set of steps in order to generate each type of the extracted audience.
[0113] Such pre-determined rules may comprise, for example, rules for generating the query-related polygon 172 and determining the polygon borders of the query-related polygon 172. Since the system 100 may generate (build) a variety of audience types as described herein, such as, for example, the initial audience 140, the extended audience 142, the targeted extended audience 145 and their variations such as, for example, the home-extended audience, the work-extended audience, etc. There may be various scripts using the same methods as described herein that may be called upon based on the predetermined (pre-defined) rules.
[0114] At step 404, for the requested location address 109a provided by the user 107, a query- related geohash (such as three-character geohash described above) is determined, and then the requested location address 109a is mapped to a corresponding building related to that requested location address 109a and by using the geofence dataset 133 and, for example, by using the google geocoding application programming interface (API). The query-related geohash and query-related time period is used to determine one or more query-related combined dataset partition(s) 137a (Fig. 1A) which corresponds to one or more combined dataset partition(s) of the set of the combined dataset partitions which the processor 112 needs to search in order to identify polygon mobile advertiser identifiers related to the query-related polygon 172 and the audience refining parameters 109b. Only one or more specific combined dataset partitions, such as the previously determined query-related combined dataset partition 137a based on the query-related geohash and the query-related time period is (are) searched at step 405 to identify polygon mobile advertiser identifiers related to the query-related polygon 172 and the audience refining parameters 109b. Searching a limited number of partitions, and preferably one partition, which has much less data than the combined dataset 130, allows to process the data and generate the consumer audiences on demand in a cheaper and faster way compared to conventional methods.
[0115] Once a requested polygon address is mapped to the respective query-related building using the geofence dataset 133, the coordinate of the query-related building may be converted to the corresponding geohash (three-character geohash), also referred herein as a “requested building geohash”. The processor 112 later searches the query-related combined dataset partition 137a and extracts the MAI Ds 161 related to the query-related polygon 172 without having to search through the whole combined dataset 130 but rather to search and extract the MAIDs 161 those related to the determined requested building geohash, the query-related combined dataset partition. Due to searching within the query-related combined dataset partition 137a related to the geohash, the cost related to the extraction of the data related to the query data may be significantly reduced.
[0116] At step 406, polygon mobile advertiser identifiers (additional MAIDs 161) are extracted from the query-related combined dataset partition 137a based on the requested audience type (requested by the user earlier at step 401), and in particular, MAIDs 161 are extracted from the query-related combined dataset partition 137a corresponding to the requested audience type and based on the geohash mapping and therefore the requested building geohash determined earlier, and the audience type in the query data108 that was submitted by the user 107. In other words, in at least one embodiment, the query-related combined dataset partition 137a corresponds to the requested building geohash.
[0117] The extended audience 142 is generated using the additional MAIDs 161. Based on the extended audience 142, the processor 112 identified the associated MAIDs 161 at step 407 using the extended audience and the quality scores. In at least one embodiment, the quality score is defined by two factors in the case of extensions. The first one is the link between the MAID 161 in the initial audience and a building, and the second quality score matching this building to other MAIDs 161 also attached to these buildings.
[0118] At step 408, the extended audience 142 is generated and then transmitted to the user 107 (for example, through the platform’s interface). The information about the extended audience volume and / or targeted extended audience 145 (where the option of the home, work or visiting extensions has been requested by the user 107) are then generated and transmitted to the user 107. The extended audience 142 may comprise, for example, a set of MAIDs 161 , as well as brand studies of over-index and under-index brands and sociodemographic of the MAIDs 161 found in the audience. For example, the initial audience 140, the extended audience 142 and / or targeted extended audience 145 may be transmitted by email and / or through the platform.
[0119] The initial audience 140, the extended audience 142 and the targeted extended audience 145 each comprises a list of mobile advertiser identifiers (MAIDs) that correspond to the query data 108 based on which that particular audience was generated by the processor 112.
[0120] In at least one embodiment, the processor 112 may execute a method 500 for generating a new syndicated audience 147 illustrated in Fig. 5, in accordance with at least one embodiment of the present disclosure. At step 501 , the processor 112 receives a set of locations of interest of the point-of- interest dataset 135 that may be previously identified, for example, using data related to market demand. The set of locations of interest comprises locations (preferably, all locations) that may be considered as “interesting” for the new syndicated audience 147. In some embodiments, the processor 112 may extract keywords based on the locations of interest or, vice versa, the processor 112 may determine the locations of interest based on the entered keywords. The processor 112 then maps (matches) the keywords related to the locations of interest to the POI locations (preferably all POI locations of the POI dataset 135, previously partitioned based on the geohash) at step 502. At step 503, POI locations are mapped (matched) to buildings using the geofence dataset 133, for example, based on the POI geographical position (such as, for example, latitude and longitude) of the POI locations. The POI locations are thus linked to an actual building based on a point on a map that was provided by the POI geographical location. In other words, the POI locations are mapped to the buildings using the geofence dataset 133.
[0121] For each POI location, the processor 112 determines a POI geohash which corresponds to the POI geographical location. At step 504, based on the geohash of the POI locations (“POI geohash”), a syndicated audience 147 (Fig. 1 B) may be generated by extracting the MAIDs 161 of the location datapoint dataset partitions 123 which correspond to each one of the POI geohashes determined earlier for the POI locations. The syndicated audience 147 is a list of MAIDs 161 that attended the POI buildings is extracted.
[0122] Based on the syndicated audience 147, extensions may be generated at step 505. As discussed herein for the extensions of the initial audience 140, the generated so-called “extensions” for the syndicated audience 147 may also comprise a home extension corresponding to the home-extension audience, a work extension corresponding to the work-extension audience, and / or the social-extension corresponding to a social extension audience. The extensions generated for the initial audience or the syndicated audience 147 to obtain the extended audiences, help to enrich the initial audience or the syndicated audience 147, respectively, in a meaningful way rather than increasing radius around a specific location.
[0123] Based on all the MAIDs 161 extracted at the execution of method 400, the processor 112 of the system 100 extracts all the associated MAIDs 161 using the syndicated audience 147 databases and the quality scores to generate a POI extended audience 148 (Fig. 1 B), which may be a home-extension POI audience, the work-extension POI audience, and / or the social-extension POI audience. At step 506, the audiences, such as the syndicated audience 147 and / or the POI extended audience 148 generated by the system 100 are sent to DSP for activation. The system 100 then transmits these audiences to DSP in order for the users of these DSP to use them in their advertisement campaigns.
[0124] In at least one embodiment, keywords are mapped to locations of interest of point-of-interest locations of the point-of-interest database 135. Then the point-of-interest locations are mapper to buildings using the geofence dataset 133. Based on the point-of-interest location geohashes of the point- of-interest locations, audience attendance data may be then extracted and the syndicated audience 147 comprising syndicated audience MAIDs may be generated.
[0125] In at least one embodiment, for each syndicated audience MAID of the syndicated audience MAIDs of the syndicated audience 147, the partitioned combined dataset 138 may be used to determine a building geofence related to a location of the syndicated audience MAID during a first pre-determined time period, from a location data point dataset partition 123 corresponding to a geohash of the building geofence, social-extended MAIDs may be then extracted, different from the syndicated audience MAIDs, detected within the building geofence of the syndicated audience MAIDs during a second pre-determined time period. An extended point-of-interest audience having the extended MAIDs and the syndicated audience MAIDs of the syndicated audience may be then generated.
[0126] A method and a system for modeling of data related to consumers and locations are provided. In at least one embodiment, the method comprises: collecting location data points related to the consumers comprising, for each electronic device of each consumer, a mobile advertiser identifier, a timestamp, a latitude and a longitude, and an IP address of an electronic device of each consumer; generating a combined dataset by: mapping location data points to a geohash, and a business and / or abuilding based on the latitude and the longitude for each electronic device by using geofiles; mapping a point-of-interest (POI) database to the building and / or a brand name; assign characteristics to the customers and / or buildings, characteristics being at least one of home location, work location, and visitation frequency; and compressing the resulting dataset to generate the combined dataset; and storing the combined dataset in a database using the corresponding geohash identification.
[0127] The steps of the method embodiments as described herein are executed by various routines executed by one or more processors 112 located on the server 110 (Fig. 1A). In addition to the crawler routine 151 described above, the processor(s) of the server 110 is (are) configured to execute other routines, such as, for example, mapping routine 152 (executing, for example, steps 204, 206, 308 of the method embodiments described above), assigning routine 153 (executing, for example, steps 310, 312, 314), geohash routine 154 (executing, for example, steps 208, 304), compressing routine 151 (executing, for example, step 210). Each one of these specific routines executes one or more steps of the method embodiments as described herein. While executing the routines described herein, the at least one processor 112 of the system 100 is configured to execute the instructions stored on a computer readable memory of the server 110. The computer executable instructions, when executed by the processor(s) 112 (server(s)) 110 as described herein, are configured to perform the steps of the methods as described herein.
[0128] The geohash routine 154 processes data using the geohash (e.g. three-character geohash, geohash3). In at least one embodiment, when extracting the data, the processor 112 is configured to uncompress the data from the parquet format. The processor 112 also is configured to combine all files of any database (dataset) described herein that are part of the same geohash3 (are related to the same geohash). In at least one embodiment, these files are recompressed in the parquet format with snappy compression.
[0129] The server 110 may be a combination of several servers, located remotely to each other. The server 110 may communicate with one or more other servers (and therefore processors) to execute one or more steps of the method described above. For example, in some embodiments, other servers (and / or services, such as, for example, Amazon Athena) may execute the instructions generated by (or stored on) and transmitted to them by and from the server 110 (such as, for example, when executing some of the steps 322-334 using Amazon Athena or another similar platform / service).
[0130] While preferred embodiments have been described above and illustrated in the accompanying drawings, it will be evident to those skilled in the art that modifications may be made without departing from this disclosure. Such modifications are considered as possible variants comprised in the scope of the disclosure.
Claims
CLAIMS:1 . A method for generation of data related to consumers to be executed by a system comprising a processor and a database comprising a geofiles database, a geofence dataset and a point- of-interest database, the method comprising: receiving, from a plurality of electronic devices, location data points related to the consumers, each location data point comprising, for each electronic device of each consumer, a mobile advertiser identifier (MAID), a timestamp, a physical location identification of the electronic device, and an internet protocol (IP) address of the electronic device; generating a partitioned combined dataset by: mapping, by a mapping routine, the MAID to a geofiles data in the geofiles database and to a geofence data in the geofence dataset based on the physical location identification for each location data point of the location data points; mapping, by the mapping routine, the MAID to a company data in the point-of-interest database corresponding to the MAID to generate a combined dataset; splitting the combined dataset with respect to geohashes and time periods to generate combined dataset partitions; compressing each one of the combined dataset partitions to generate the partitioned combined dataset; and storing the partitioned combined dataset based on the corresponding geohash identification and the corresponding time periods.
2. The method of claim 1 , further comprising assigning additional characteristics to the MAID of the location data point, the additional characteristics being at least one of consumer home location and consumer work location.
3. The method of claim 1 or 2, further comprising: receiving, from a user device, a query data comprising requested location address and an audience refining parameters; based on the requested location address, generating a query-related polygon and generating requested polygon addresses located within the query-related polygon; mapping the requested polygon addresses to building polygons using the geofence database; mapping the requested polygon addresses located within the query-related polygon to at least one corresponding query-related geohash to determine at least one query-related combined dataset partition of the combined dataset partitions based on the query-related geohash and a query-related time period; searching the at least one query-related combined dataset partition to identify polygon MAIDs related to the query-related polygon and the audience refining parameters; andgenerating an initial audience comprising the polygon MAIDs.
4. The method of claim 3, wherein the query-related time period is provided in the query data or determined by the processor based on the query data.
5. The method of claim 3 or 4, wherein the initial audience is a residential initial audience, initial residential MAIDs being the polygon MAIDs, and the method further comprises: for each initial residential MAID, using the partitioned combined dataset, determining a work building geofence related to a location of the initial residential MAID during an extended day-time period; extracting, from the partitioned combined dataset, work-extended MAIDs, different from the initial residential MAIDs, detected within the work building geofence of the initial residential MAIDs; and generating an extended work-based audience having the work-extended MAIDs and the initial residential MAIDs of the residential initial audience.
6. The method of claim 3 or 4, wherein the initial audience is a work initial audience, initial work MAIDs being the polygon MAIDs, and the method further comprises: for each initial work MAID, using the partitioned combined dataset, determining a residential building geofence related to a location of the initial work MAID during an extended night-time period; extracting, from the partitioned combined dataset, home-extended MAIDs, different from the initial work MAIDs, detected within the residential building geofence of the initial work MAIDs; and generating an extended home-based audience having the home-extended MAIDs and the initial work MAIDs of the work initial audience.
7. The method of claim 3 or 4, further comprising: for each initial MAID of the initial MAIDs, using the partitioned combined dataset, determining a building geofence related to a location of the initial MAID during a first predetermined time period, the initial MAIDs being the polygon MAIDs; extracting, from the partitioned combined dataset, social-extended MAIDs, different from the initial MAIDs, detected within the building geofence of the initial MAIDs during a second pre-determined time period; and generating an extended social-based audience having the social-extended MAIDs and the initial MAIDs of the initial audience.
8. The method according to any one of claims 3 to 7, further comprising generating a targeted extended audience by reducing the initial audience based on a building attendance frequency determined by the processor and related to each one of the polygon MAIDs, the building attendance frequency being within a pre-determined building attendance frequency range.
9. The method according to any one of claims 3 to 8, further comprising:mapping keywords to locations of interest of point-of-interest locations of the point-of- interest database; mapping point-of-interest locations to buildings using the geofence dataset; based on the point-of-interest location geohashes of the point-of-interest locations, extracting audience attendance data; and generating a syndicated audience comprising syndicated audience MAIDs.
10. The method of claim 9, further comprising: for each syndicated audience MAID of the syndicated audience MAIDs, using the partitioned combined dataset, determining a building geofence related to a location of the syndicated audience MAID during a first pre-determined time period; extracting, from a location data point dataset partition corresponding to a geohash of the building geofence, social-extended MAIDs, different from the syndicated audience MAIDs, detected within the building geofence of the syndicated audience MAIDs during a second predetermined time period; and generating an extended point-of-interest audience having the extended MAIDs and the syndicated audience MAIDs of the syndicated audience.
11. The method according to any one of claims 3 to 10, wherein the at least one requested polygon address located within the query-related polygon is determined by searching a geofence dataset.
12. The method according to any one of claims 3 or 11 , wherein the query data comprises a requested audience type and the polygon MAIDs are extracted from the query-related combined dataset partition based on the requested audience type and geohash mapping to a requested building geohash.
13. A system for modeling of data related to consumers and locations, the system comprising: a database comprising a geofiles database, geofence dataset and point-of-interest dataset, and a processor configured to: receive, from a plurality of electronic devices, location data points related to the consumers, each location data point comprising, for each electronic device of each consumer, a mobile advertiser identifier (MAID), a timestamp, a physical location identification of the electronic device, and an internet protocol (IP) address of the electronic device; generate a partitioned combined dataset by: mapping, by a mapping routine, the MAID to a geofiles data in the geofiles database and to a geofence data in the geofence dataset based on the physical location identification for each location data point of the location data points; mapping, by the mapping routine, the MAID to a company data in the point-of-interest database corresponding to the MAID to generate a combined dataset;splitting the combined dataset with respect to geohashes and time periods to generate combined dataset partitions; compressing each one of the combined dataset partitions to generate the partitioned combined dataset; and storing the partitioned combined dataset based on the corresponding geohash identification and the corresponding time periods.
14. The system of claim 13, wherein the processor is further configured to assign additional characteristics to the MAID of the location data point, the additional characteristics being at least one of consumer home location and consumer work location.
15. The system of claim 13 or 14, wherein the processor is further configured to: receive, from a user device, a query data comprising requested location address and an audience refining parameters; based on the requested location address, generate a query-related polygon and generating requested polygon addresses located within the query-related polygon; map the requested polygon addresses to building polygons using the geofence database; map the requested polygon addresses located within the query-related polygon to at least one corresponding query-related geohash to determine at least one query-related combined dataset partition of the combined dataset partitions based on the query-related geohash and a query-related time period; search the at least one query-related combined dataset partition to identify polygon MAIDs related to the query-related polygon and the audience refining parameters; and generate an initial audience comprising the polygon MAIDs.
16. The system of claim 15, wherein the query-related time period is provided in the query data or determined by the processor based on the query data.
17. The system of claim 15 or 16, wherein the initial audience is a residential initial audience, initial residential MAIDs being the polygon MAIDs, and the processor is further configured to: for each initial residential MAID, using the partitioned combined dataset, determine a work building geofence related to a location of the initial residential MAID during an extended daytime period; extract, from the partitioned combined dataset, work-extended MAIDs, different from the initial residential MAIDs, detected within the work building geofence of the initial residential MAIDs; and generate an extended work-based audience having the work-extended MAIDs and the initial residential MAIDs of the residential initial audience.
18. The system of claim 15 or 16, wherein the initial audience is a work initial audience, initial work MAIDs being the polygon MAIDs, and the processor is further configured to:for each initial work MAID, use the partitioned combined dataset, determine a residential building geofence related to a location of the initial work MAID during an extended night-time period; extract, from the partitioned combined dataset, home-extended MAIDs, different from the initial work MAIDs, detected within the residential building geofence of the initial work MAIDs; and generate an extended home-based audience having the home-extended MAIDs and the initial work MAIDs of the work initial audience.
19. The system of claim 15 or 16, wherein the processor is further configured to: for each initial MAID of the initial MAIDs, using the partitioned combined dataset, determine a building geofence related to a location of the initial MAID during a first predetermined time period, the initial MAIDs being the polygon MAIDs; extract, from the partitioned combined dataset, social-extended MAIDs, different from the initial MAIDs, detected within the building geofence of the initial MAIDs during a second predetermined time period; and generate an extended social-based audience having the social-extended MAIDs and the initial MAIDs of the initial audience.
20. The system according to any one of claims 15 to 19, wherein the processor is further configured to generate a targeted extended audience by reducing the initial audience based on a building attendance frequency determined by the processor and related to each one of the polygon MAIDs, the building attendance frequency being within a pre-determined building attendance frequency range.
21. The system according to any one of claims 15 to 20, wherein the processor is further configured to: map keywords to locations of interest of point-of-interest locations of the point-of-interest database; map point-of-interest locations to buildings using the geofence dataset; based on the point-of-interest location geohashes of the point-of-interest locations, extract audience attendance data; and generate a syndicated audience comprising syndicated audience MAIDs.
22. The system of claim 21 , wherein the processor is further configured to: for each syndicated audience MAID of the syndicated audience MAIDs, using the partitioned combined dataset, determine a building geofence related to a location of the syndicated audience MAID during a first pre-determined time period; extract, from a location data point dataset partition corresponding to a geohash of the building geofence, social-extended MAIDs, different from the syndicated audience MAIDs,detected within the building geofence of the syndicated audience MAIDs during a second predetermined time period; and generate an extended point-of-interest audience having the extended MAIDs and the syndicated audience MAIDs of the syndicated audience.
23. The system according to any one of claims 15 to 22, wherein the at least one requested polygon address located within the query-related polygon is determined by searching a geofence dataset.
24. The system according to any one of claims 15 to 23, wherein the query data comprises a requested audience type and the polygon MAIDs are extracted from the query-related combined dataset partition based on the requested audience type and geohash mapping to a requested building geohash.
Citation Information
Patent Citations
Systems, methods, and apparatuses for providing content according to geolocation
US10932118B1
Listener SDK-Based Enrichment of Indoor Positioning
US20200273071A1
Systems, methods and program products for distributing personalized content to specific electronic non-personal public or semi-public displays
US20200273072A1