Methods and systems for anonymously tracking and / or profiling individuals based on biometric data

By generating anonymous identifier bias measurements and using biometric data, the contradiction between individual anonymity and group flow data collection is solved, and anonymous tracking and analyzing individual flows between different subject states is achieved, meeting the requirements of law and public opinion, and providing effective statistical data support.

CN114766020BActive Publication Date: 2025-08-19INDIVID CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080083086.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-04
Filing Date
2020-08-06
Publication Date
2025-08-19
Estimated Expiration
2040-08-06

AI Technical Summary

Technical Problem

The prior art is difficult to collect and analyze group flow data while maintaining individual anonymity, especially in the process of conversion between individual identity tracking and group statistics. Traditional pseudo-anonymization methods cannot meet the legal and public opinion requirements of the right to anonymity.

Method used

By generating anonymous identifier bias metrics, using biometric data to estimate individual flow between different subject states, using anonymization methods such as hashing and noise masking, combining decorrelation modules and bias metrics, anonymous identifiers are generated to maintain individual anonymity, and at the same time, group flow metrics are calculated.

Benefits of technology

It realizes that without storing personal data, it is possible to anonymously collect and analyze individual flowing data between different subject states, meet the legal and public opinion requirements of anonymity, and provide statistical data support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114766020B_ABST
    Figure CN114766020B_ABST
Patent Text Reader

Abstract

Methods and systems are provided for anonymously tracking and / or analyzing the flow or movement of individual subjects and / or objects using biometric data. Specifically, a computer-implemented method is provided for anonymously estimating the amount and / or flow of individual subjects and / or objects, referred to as individuals, in a population, moving and / or coinciding between two or more subject states using biometric data. The method comprises the following steps: receiving (S1) identification data from two or more individuals, wherein the identification data includes and / or is based on biometric data; generating (S2) an anonymous identifier for each individual online and via one or more processors; and storing (S3): the anonymous identifier for each individual and data representing the subject state; and / or a deviation measure for such anonymous identifier.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates generally to the problem of anonymity in technical applications and to technical aspects of data collection and data / population statistics based on biometric data, and more particularly to the technical field of estimating or measuring population flows and / or to methods and systems and computer programs for implementing such population flow estimation based on biometric data. Background Art

[0002] Legislation and public opinion are increasingly driving the move toward anonymity in technology. This conflicts with the need to collect data on group mobility to automate or optimize processes and social processes. Retailers want to collect statistics on their visitors to improve their operations. Smart cities need data to optimize quality of life and energy efficiency. Public transportation systems need to collect data on travel patterns to reduce travel times and optimize costs.

[0003] There is a pressing need for technologies that can collect data for statistical purposes while preserving the anonymity of individuals. Specifically, tracking the movement of people from one point in time to another is problematic, as re-identifying an individual at a later time is generally considered a violation of their anonymity. This means that the whole idea of anonymously tracking groups is slightly counterintuitive, as it is often nearly impossible at the individual level.

[0004] Current privacy-enhancing methods for tracking people, based on pseudonymization and unique identifiers, clearly fail to meet these needs, meaning companies avoid collecting data on group movements altogether. Any system that can collect data on such group movements without violating anonymity is highly desirable. In particular, profiling is widely considered to threaten the fundamental rights and freedoms of individuals. In some cases, encryption with minimal information damage has been used, making it possible to re-identify individuals with a sufficiently high probability (typically an error rate of tens of thousands of times) that any misidentification can be completely ignored. However, such pseudonymization techniques, regardless of whether they are actually reversible, are considered incompatible with legislative interpretations of anonymization or with public opinion, as the possibility of re-identification is itself a defining property of personal data. Summary of the Invention

[0005] A general object is to provide a system for providing anonymity when computing group statistics based on biometric data.

[0006] A specific object is to provide a system and method for maintaining anonymity when estimating or measuring an individual's flow between two or more spatiotemporal locations, computer system states interacting with users, and / or health states and health monitoring states of a subject (collectively or individually referred to as subject states) based on biometric data.

[0007] Another object is to provide a system for anonymously tracking and / or analyzing the transitions and / or flows and / or movements of individual subjects and / or objects (referred to as individuals) based on biometric data.

[0008] Yet another object is to provide a monitoring system comprising such a system.

[0009] Yet another object is to provide a computer-implemented method for estimating the amount or number of individuals in a population that overlap between two or more subject states based on biometric data.

[0010] Another object is to provide a method for generating a measure of the flow or movement of individual subjects and / or objects (referred to as individuals) between subject states based on biometric data.

[0011] Yet another object is to provide a computer program and / or a computer program product configured to perform such a computer-implemented method.

[0012] These and other objects are met by the embodiments as defined herein.

[0013] According to a first aspect, there is provided a system comprising:

[0014] - one or more processors;

[0015] an anonymization module configured to, via the one or more processors, receive, for each of a plurality of individuals comprising individual subjects and / or objects in a population of individuals, identification information representing an identity of the individual, wherein the identification information representing the identity of the individual includes and / or is based on biometric data, and generate an anonymous identifier deviation metric based on the identification information of the one or more individuals;

[0016] - a memory configured to store at least one anonymous identifier deviation metric based on at least one of the generated identifier deviation metrics;

[0017] an estimator configured to receive, via the one or more processors, a plurality of anonymous identifier deviation metrics, at least one identifier deviation metric for each of at least two subject states of an individual from the memory and / or directly from the anonymization module, and generate, based on the received anonymous identifier deviation metrics, one or more population flow metrics associated with individuals passing from one subject state to another subject state.

[0018] According to a second aspect, there is provided a system for anonymously tracking and / or analyzing the flow or movement of individual subjects and / or objects (referred to as individuals) between subject states based on biometric data.

[0019] The system is configured to determine an anonymous identifier for each individual in a group of multiple individuals using identification information representing the identity of the individual as input, wherein the identification information representing the identity of the individual includes and / or is based on biometric data. Each anonymous identifier corresponds to any individual in a group of individuals whose identification information produces the same anonymous identifier with a probability such that no individual generates the anonymous identifier with a probability greater than the sum of the probabilities of all other individuals generating the identifier.

[0020] The system is further configured to keep track of deviation metrics, one deviation metric for each of the two or more subject states, wherein each deviation metric is generated based on anonymous identifiers associated with the corresponding individuals associated with a particular corresponding subject state.

[0021] The system is further configured to determine at least one population flow metric representing a number of individuals passing from a first subject state to a second subject state based on the deviation metrics corresponding to the subject states.

[0022] According to a third aspect, there is provided a monitoring system comprising the system according to the first aspect or the second aspect.

[0023] According to a fourth aspect, there is provided a computer-implemented method for anonymously estimating the amount and / or flow of movement and / or overlap of individual subjects and / or objects (referred to as individuals) in a population between two or more subject states based on biometric data. The method comprises the following steps:

[0024] - receiving identification data from two or more individuals, wherein the identification data of each individual includes and / or is based on biometric data;

[0025] - generating an anonymous identifier for each individual online and by one or more processors; and

[0026] - Storing: an anonymous identifier for each individual and data representing the subject's status; and / or a measure of deviation of such anonymous identifiers.

[0027] According to a fifth aspect, there is provided a computer-implemented method for generating a measure of the flow or movement of individual subjects and / or objects (referred to as individuals) between subject states based on biometric data. The method comprises the following steps:

[0028] - configuring one or more processors to receive anonymous identifier deviation metrics generated based on biometric-based identifiers from an individual's visits to each of two subject states and / or the individual's presence in each of the two subject states, wherein each identifier represents an identity of the individual and includes and / or is based on biometric data;

[0029] - generating, using the one or more processors, a population flow metric between two subject states by comparing the anonymous identifier deviation metrics between the subject states;

[0030] - Storing said population flow metric in a memory.

[0031] According to a sixth aspect, there is provided a computer program comprising instructions which, when executed by at least one processor, cause the at least one processor to perform the computer-implemented method according to the fourth and / or fifth aspects.

[0032] According to a seventh aspect, there is provided a computer program product comprising a non-transitory computer readable medium having such a computer program stored thereon.

[0033] According to an eighth aspect, a system for executing the method according to the fourth aspect and / or the fifth aspect is provided.

[0034] In this way, anonymity can be effectively provided while allowing data collection and the calculation of statistics about groups of individuals based on biometric data.

[0035] Specifically, the proposed technique is able to preserve anonymity while estimating or measuring an individual's flow between two or more subject states based on biometric data.

[0036] In particular, the proposed invention allows linking of data points collected at different times based on biometric data for statistical purposes without the need to store personal data.

[0037] Generally speaking, the present invention provides improved techniques for achieving and / or protecting anonymity in connection with data collection and statistics based on biometric data.

[0038] Other advantages provided by the present invention will be understood when reading the following description of embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The present invention, together with further objects and advantages thereof, may best be understood by reference to the following description taken in conjunction with the accompanying drawings, in which:

[0040] Figure 1Ais a schematic diagram illustrating an example of a system according to an embodiment.

[0041] Figure 1B is a schematic flow chart illustrating an example of a computer-implemented method for achieving anonymous estimation of the amount and / or flow of movement and / or coincidence of individual agents and / or objects (referred to as agents) in a population between two or more agent states.

[0042] Figure 1C is a schematic flow chart illustrating another extended example of a computer-implemented method for achieving anonymous estimation of volume and / or flow of individual subjects and / or objects.

[0043] Figure 1D is a schematic flow chart illustrating an example of a computer-implemented method for generating a measure of the flow or movement of individual subjects and / or objects (referred to as subjects) between subject states.

[0044] Figure 2 is a schematic diagram showing an example of microaggregating a population into groups.

[0045] Figure 3 is a diagram showing another example of microaggregating populations into groups, including the concept of deviation measures.

[0046] Figure 4 is a diagram showing how each group of individuals can be associated with a set of agent states N, each agent state for a set of time points.

[0047] Figure 5 is a diagram showing examples of subject status such as spatiotemporal position data and useful identification biometric information (ID).

[0048] Figure 6 is a schematic diagram illustrating an example of a monitoring system.

[0049] Figure 7 is a schematic flow chart illustrating an example of a computer-implemented method for estimating the amount or number of individuals in a population that coincide between two or more spatiotemporal locations.

[0050] Figure 8 is a schematic flow chart illustrating another example of a computer-implemented method for estimating the amount or number of individuals in a population that overlap between two or more spatiotemporal locations.

[0051] Figure 9 is a schematic diagram illustrating an example of the movement or flow of one or more individuals from location A to location B.

[0052] Figure 10is a diagram illustrating an example of movement or flow of a user from one virtual location, such as an IP location, to another virtual location.

[0053] Figure 11 is a schematic diagram illustrating an example of a computer implementation according to an embodiment.

[0054] Figure 12 is a schematic flow chart illustrating an example of a computer-implemented method for generating a measure of the flow or movement of individual subjects and / or objects (referred to as individuals) between locations in space and time.

[0055] Figure 13 is a diagram illustrating an example of how an identifier bias metric can be anonymized by adding noise at one or more times and how this can generate a bias compensation term.

[0056] Figure 14 is an example demonstrating noise masking anonymization. DETAILED DESCRIPTION

[0057] Throughout the drawings, the same reference numerals are used for similar or corresponding elements.

[0058] To better understand the proposed technique, it may be useful to start with a brief analysis of the technical problem.

[0059] Careful analysis by the inventors has revealed that personal data can be made anonymous by storing partial identities (i.e., partial information about an individual's identity that is not itself personal data). Furthermore, and perhaps surprisingly, it is possible to build a system capable of measuring group mobility using such anonymous data even when the data is based on factors that are not directly related to group mobility and / or its distribution. Importantly, the proposed invention also works if the factors used are unrelated to group mobility and / or if any estimate of their prior distribution is not feasible. Thus, the present invention is applicable to general groups using almost any identifying factor (i.e., data type), without requiring further knowledge of the underlying distribution.

[0060] The present invention provides systems and methods for anonymously estimating population mobility. Three specific anonymization methods and systems suitable for achieving these objectives are also provided. Briefly, two anonymization methods, hashing and noise masking, are based on anonymizing identifying information about each access to a subject's state in an anonymization module, while the third method is based on anonymizing desired stored data (i.e., identifier deviation metrics). These methods can also be used in combination with one another.

[0061] The present invention also provides a way to use the present invention without first estimating the underlying distribution by using a decorrelating hash module and / or a decorrelating module and / or a decorrelating deviation metric.

[0062] In the following, reference will be made to Figures 1A to 11 The exemplary schematic diagrams depict non-limiting examples of the proposed technology.

[0063] Figure 1A is a schematic diagram illustrating an example of a system according to an embodiment. In this particular example, the system 10 basically comprises one or more processors 11 , an anonymization module 12 , an estimator 13 , an input / output module 14 and a memory 15 having one or more deviation metrics 16 .

[0064] According to a first aspect of the present invention, there is provided a system 10 comprising:

[0065] - one or more processors 11, 110;

[0066] an anonymization module 12 configured to perform the following operations via one or more processors 11, 110: for each of a plurality of individuals comprising individual subjects and / or objects in a population of individuals, receive identification information representing an identity of the individual, wherein the identification information representing the identity of the individual includes and / or is based on biometric data; and generate an anonymous identifier deviation metric based on the identification information of the one or more individuals;

[0067] - a memory 15, 120 configured to store at least one anonymous identifier deviation metric based on at least one of the generated identifier deviation metrics;

[0068] - an estimator 13, which is configured to perform the following operations via the one or more processors 11, 110: receive a plurality of anonymous identifier deviation measures, at least one identifier deviation measure for each of at least two subject states of an individual from the memory and / or directly from the anonymization module; and generate one or more group flow measures related to the individual passing from one subject state to another subject state based on the received anonymous identifier deviation measures.

[0069] For example, each identifier deviation metric is generated based on two or more identifier density estimates and / or based on one or more values generated based on the identifier density estimates.

[0070] For example, each identifier deviation metric represents the deviation of identification information of one or more individuals compared to an expected distribution of such identification information in a population.

[0071] In a particular example, the identifier bias metric of the anonymization module is based on a group identifier that represents a large number of individuals.

[0072] For example, the identifier deviation metric may be based on a visit counter.

[0073] For example, the identifier deviation metric is generated based on the identification information using a hash function.

[0074] For example, the anonymization module 12 may be configured to generate a group identifier based on the individual's biometric information by using a Locality Sensitive Hash (LSH) function.

[0075] As an example, the one or more population flow metrics include the number and / or rate of visitors passing from one spatiotemporal location / venue to another spatiotemporal location / venue.

[0076] For example, at least one of the one or more population flow metrics is generated based at least in part on a linear transformation of counter information of two or more access counters.

[0077] Optionally, the anonymization module 12 and / or the identification information representing the identity of the individual is random, and wherein the randomness of the identification information and / or the anonymization module 12 is taken into account when generating the linear transformation.

[0078] For example, when generating population flow metric(s), a baseline corresponding to the expected correlation from two independently generated populations is subtracted.

[0079] For example, each identifier bias metric may be generated using a combination of an identifier and noise, such that contributions to the identifier bias metric are rendered anonymous due to a sufficient level of noise that access to subject state is not attributable to a particular identifier.

[0080] As an example, the identifier deviation metric may be based on two or more identifier density estimates.

[0081] In a particular example, the anonymization module is configured to generate at least one identifier deviation metric based on (multiple) anonymous identifier deviation metrics stored in the memory; and provide anonymity by adding sufficient noise to the anonymous identifier deviation metrics stored in the memory at one or more time instants such that the total contribution from any single identifier cannot be determined.

[0082] Optionally, information about the generated noise sample(s) is also stored and used to reduce the variance of the group flow metric.

[0083] For example, identification information representing the identity of an individual may include and / or be based on at least one of the following non-limiting examples of biometric data: iris image, facial image, feature vector, body image, fingerprint, and / or gait.

[0084] In other words, identification information can be considered as biometric information that represents the identity of an individual.

[0085] For example, these subject states include spatiotemporal location, computer system state interacting with a user, and / or health state and health monitoring state of the subject.

[0086] For example, biometric vectors are based on neural networks that extract representations of possible biometric data from images containing biometric information.

[0087] For example, in addition to biometric data, identification data may also contain, encode and / or represent additional identification data, such as image data or feature vectors based on image data, which also contains clothing and / or other non-biometric data as well as, for example, the face.

[0088] In a specific example, which will be elaborated upon later, these subject states are spatiotemporal positions and / or locations, and

[0089] The anonymization module 12 is configured to generate a group identifier based on the identification information of the individuals to effectively micro-aggregate the population into corresponding groups;

[0090] The memory 15, 120 is configured to store a visit counter for each of the two or more group identifiers from each of the two or more spatiotemporal locations or places associated with the corresponding individual; and

[0091] The estimator 13 is configured to receive counter information from the at least two visit counters and generate one or more group flow metrics related to individuals passing from one spatiotemporal location to another spatiotemporal location.

[0092] For example, the anonymization module may be configured to generate a group identifier based on identification information of an individual by using a hash function.

[0093] For example, the system 10, 100 includes an input module 14, 140, which is configured to perform the following operations through the one or more processors 11, 110: for each of a large number of individuals, receive location data representing a spatiotemporal location; and match the spatiotemporal location of the individual with a visit counter corresponding to a group identifier associated with the individual, and each visit counter of each group identifier also corresponds to a specific spatiotemporal location.

[0094] According to a second aspect, there is provided a system 10, 100 for anonymously tracking and / or analyzing the flow or movement of individual subjects and / or objects (referred to as individuals) between subject states based on biometric data.

[0095] The system 10, 100 is configured to determine an anonymous identifier for each individual in a group of multiple individuals using identification information representing the identity of the individual as input, wherein the identification information representing the identity of the individual includes and / or is based on biometric data. Each anonymous identifier corresponds to any individual in a group of individuals whose identification information produces the same anonymous identifier with a probability such that no individual generates the anonymous identifier with a probability greater than the sum of the probabilities of all other individuals generating the identifier.

[0096] The system 10 , 100 is configured to keep track of deviation metrics, one deviation metric for each of two or more subject states, wherein each deviation metric is generated based on anonymous identifiers associated with the corresponding individuals associated with a particular corresponding subject state.

[0097] The system 10 , 100 is further configured to determine at least one population flow metric representing the number of individuals passing from a first subject state to a second subject state based on the deviation metrics corresponding to the subject states.

[0098] These anonymous identifiers are, for example, group identifiers and / or noise-masked identifiers.

[0099] In a specific non-limiting example, the system 10 , 100 is configured to determine a group identifier for each individual in a group of multiple individuals based on a hash function using information representing the identity of the individual as input.

[0100] Each group identifier corresponds to a group of individuals whose identity information results in the same group identifier, thereby effectively micro-aggregating the population into at least two groups.

[0101] In this example, the subject states are spatiotemporal locations or places and the deviation metrics correspond to visit data, and the system 10, 100 is configured to keep track of visit data for each group, which represents the number of visits to two or more spatiotemporal locations by individuals belonging to the group.

[0102] The system 10 , 100 is further configured to determine at least one group flow metric representing a number of individuals passing from the first spatiotemporal location to the second spatiotemporal location based on the visit data for each group identifier.

[0103] For example, the system 10, 100 includes processing circuitry 11, 110 and memory 15, 120, wherein the memory includes instructions that, when executed by the processing circuitry, cause the system to anonymously track and / or analyze the movement or activity of individuals.

[0104] For example, the anonymization module 12 may be configured to generate a group identifier and / or a noise-masked identifier based on identification information of an individual by using a hash function.

[0105] Figure 1B is a schematic flow chart illustrating an example of a computer-implemented method for achieving anonymous estimation of the amount and / or flow of movement and / or coincidence of individual subjects and / or objects (referred to as individuals) in a population between two or more subject states based on biometric data.

[0106] The method comprises the following steps:

[0107] - receiving (S1) identification data from two or more individuals, wherein the identification data of each individual comprises and / or is based on biometric data;

[0108] - generating (S2) an anonymous identifier for each individual online and by one or more processors; and

[0109] - Storing (S3): an anonymous identifier for each individual and data representing the subject's state; and / or a measure of deviation of such an anonymous identifier.

[0110] For example, the anonymous identifier may be an anonymous identifier deviation metric or other anonymous identifier that is not actually correlated with population flows.

[0111] For example, the deviation metric may be decorrelated and / or identifying data that is somehow related to group flows, and wherein the anonymous identifier is generated using a decorrelation module and / or a decorrelation hash module.

[0112] In a particular example, the anonymous identifier is an anonymous deviation metric and the anonymous deviation metric is generated based on stored anonymous identifier deviation metrics to which noise has been added at one or more time instances.

[0113] As an example, anonymous identifiers may be generated by adding noise to identification data.

[0114] For example, a compensation term to be added to the group flow estimate and / or necessary information for generating such a group flow estimate is calculated based on one or more generated noise samples used by the method.

[0115] For example, any two stored anonymous identifiers or identifier deviation measures are unlinkable to each other, ie, there are no pseudo-anonymous identifiers that link states in the stored data.

[0116] In a specific example, the anonymous identifier is a group identity, and the group identity for each individual is stored along with data representing the subject state; and / or a counter for each subject state and group identity.

[0117] For example, the subject state may be a spatiotemporal location, a state of a computer system interacting with a user, and / or a health state and / or health monitoring state of the subject.

[0118] Optionally, activity data representing one or more actions or activities of each individual is also stored along with the corresponding group identity and data describing the subject's state.

[0119] Optionally, the method may further comprise the step of generating (S4) a population flow measure between two subject states, such as Figure 1C Indicated schematically in FIG.

[0120] Figure 1D is a schematic flow chart illustrating an example of a computer-implemented method for generating a measure of the flow or movement of individual subjects and / or objects (referred to as subjects) between subject states based on biometric data.

[0121] The method comprises the following steps:

[0122] - configuring (S11) one or more processors to receive anonymous identifier deviation metrics generated based on biometric-based identifiers from visits by an individual to each of two subject states and / or the individual's presence in each of the two subject states, wherein each identifier represents an identity of the individual and includes and / or is based on biometric data;

[0123] - generating (S12) a population flow measure between two subject states by comparing the anonymous identifier deviation measures between the subject states, using the one or more processors;

[0124] - Storing (S13) said group flow metric in a memory.

[0125] These subject states are, for example, a spatiotemporal position, a computer system state interacting with a user, and / or a health state and / or health monitoring state of the subject.

[0126] For example, these anonymous identifier deviation metrics can be counters of group identities.

[0127] Typically, a single visitor present in one subject state cannot be re-identified with high probability in another subject state using these anonymous identifier deviation measures, for example, he / she cannot be linked through pseudonymization and / or through a single entry in a database.

[0128] For example, generating step S12 is not based on data that already contains some measure of population flow between locations at the individual level and / or micro-aggregate level.

[0129] For example, these anonymous identifier bias measures are not actually correlated with group mobility.

[0130] Optionally, group flow estimates are generated based on a linear mapping from these anonymous identifier deviation measures.

[0131] For example, group flow metrics may also be generated based on information about noise samples used to anonymize the data.

[0132] As an example, configuring step S11 includes configuring one or more processors to receive counters of anonymous and approximately independently distributed group identities derived from individual visits to each of two subject states; and generating step S12 includes using a linear correlation between the group identity counters of each of the two subject states to generate a group flow metric between the two subject states.

[0133] For example, the subject states may be spatiotemporal locations, and a group flow metric between two spatiotemporal locations may be generated using a linear correlation between the group identity counters of each of the two subject states.

[0134] Optionally, the anonymous identifier or identifier bias metric for each subject state may be based on two or more identifier density estimates.

[0135] Figure 2 The figure illustrates an example of microaggregating a population into groups. For example, a population of subjects / objects under study can be microaggregated into groups using a suitable one-way hash. In short, the basic concept is to use identification information representing the identity of each of a large number of individuals (e.g., ID#1, ID#2, ..., ID#Y) and generate group identifiers (Group ID#1, ..., Group ID#X) based on the individual identification information to effectively microaggregate the population into corresponding groups (Group#1, ..., Group#X).

[0136] Figure 3 1 is a diagram illustrating another example of microaggregating a population into groups, including the concept of visit counters. For each of two or more group identifiers from each of two or more spatiotemporal locations or places associated with a corresponding individual, there is a visit counter 16. In other words, each of at least two groups (with corresponding group identifiers) has a plurality (K, L, M) of visit counters for maintaining visit counts from each of two or more spatiotemporal locations or places associated with a corresponding individual of the group under consideration.

[0137] The estimator 13 (also referred to as a group flow estimator) may then be configured to receive counter information from the at least two visit counters and generate one or more group flow metrics relating to individuals passing from one spatiotemporal location to another spatiotemporal location.

[0138] Figure 4 is a diagram showing how each group of individuals can be associated with a set of spatial locations N, each for a set of time points.

[0139] Optionally, the system 10 includes an input module 14 configured to perform the following operations via one or more processors: for each of a plurality of individuals, receive location data representing a spatiotemporal location, and match the spatiotemporal location of the individual with a visit counter 16 corresponding to a group identifier associated with the individual.

[0140] For example, each access counter 16 for each group identifier also corresponds to a specific spatiotemporal location.

[0141] For example, one or more population flow metrics include the number and / or rate of visitors passing from one spatiotemporal location to another spatiotemporal location.

[0142] In a particular example, at least one of the one or more population flow metrics is generated based at least in part on a linear transformation of counter information of two or more access counters.

[0143] For example, the anonymization module 12 and / or the identification information representing the identity of the individual may be random, and the randomness of the identification information (identifier) and / or the anonymization module 12 may be considered when generating the linear transformation.

[0144] As an example, the linear transformation can be based at least in part on a correlation between two access counters and subtracting from the correlation a baseline corresponding to an expected correlation from two independently generated populations.

[0145] Figure 5 is a diagram showing examples of subject status such as spatiotemporal position data and useful identification biometric information (ID).

[0146] For example, in addition to the temporal aspect (ie related to time), spatiotemporal location data may relate to physical locations such as streets, shops, subway stations or any other suitable geographical locations and / or virtual locations such as IP addresses, domains, frames, etc.

[0147] Non-limiting examples of identification information (also referred to as identifiers) representing an individual's identity based on his / her biometric attributes may include and / or be based on at least one of: an iris image, a facial image, a feature vector, a body image, a fingerprint, and / or a gait.

[0148] This means one or more of the above-mentioned information items and / or combinations thereof.

[0149] In certain examples, the anonymization module is configured to operate based on a random table, a pseudo-random table, a cryptographic hash function, and / or other similar functions that are not actually relevant to the aspect of interest that the system is designed to study.

[0150] As an example, a hashing process may be non-deterministic.

[0151] For example, it may be considered important that data on at least two individuals are collected or expected to be collected for each unique group identifier (when such a unique group identifier is used). In other words, data on at least two individuals are collected or expected to be collected for each unique hash. Alternatively, if the criterion is slightly weaker, it may be important that at least two individuals are expected to exist in some group that could reasonably be expected to have access to the subject's state, for example, individuals in the city or country of interest for which the data is being collected. The criterion for anonymity should be the range of reasonable identities, not the range of reasonable identifiers. For example, the number of possible physical characteristics will generally be greater than the range of actual physical characteristics of a country or otherwise defined group.

[0152] More generally, to handle noise-based anonymization cases with similar criteria, it may be important, for example, that the probability of correctly identifying an individual should not be higher than 50%, with optional exceptions for cases where the probability is negligible. For example, it may also be important that the probability of identifying an individual is no higher than 50% given known subject states and / or reasonably available information about the existence of such subject states for a particular individual. This knowledge may also be probabilistic. Such probabilities can be calculated in a straightforward manner by a skilled person using analytical or Monte Carlo methods.

[0153] When using noise-masked identifiers, for example, it may be important that no noise-masked identifier value can be linked to any single person with a higher probability than the probability that the identifier value belongs to any other person in the group. Therefore, the probability that the noise-masked identifier belongs to any of the n-1 remaining individuals in a group of n people should ideally be higher than 0.5. In other words, the probability of identifying an individual should not be higher than 0.5 and in many cases is much lower, with similar protection provided by anonymization for some k=2 or higher. In other words, each of the large number of identifiers should have a probability of generating a given noise-masked identifier value that is less than the sum of the probabilities of generating the noise-masked identifier from each of the other identifiers. If the noise level is too low, the collected data allows the creation of profiles and the method is no longer anonymous due to insufficient data collection.

[0154] As an example, for four different received identifiers, the probabilities of generating a particular noise-masked identifier might be 0.6, 0.4, 0.3, and 0.4, with a maximum probability of 0.6 / 1.7 correctly assigning the data to a specific individual, thus achieving an anonymity greater than 0.5. Often, it's reasonable to assume that the prior probabilities are uniform across a population. In other cases, for example, if people are being identified through facial images and it's known a priori that certain types of faces are more likely to appear in a given population, then a prior distribution needs to be considered. In practice, this is often difficult to estimate. In such cases, it's desirable to instead use a decorrelation module and / or a distribution with sufficient distribution to leave a sufficient margin for uncertainty in the prior probabilities. A perfectly uniform distribution across all possible noise-masked identifier values, regardless of the received identifier, is impractical because this would clearly remove any desired bias in the data caused by the specific set of identifiers used to generate the noise-masked identifiers. In other words, choosing an appropriate noise distribution becomes a balance between estimation accuracy and the anonymity provided. However, there are often a variety of options that can provide a high degree of anonymity and reasonable accuracy.

[0155] It should be noted that the one or more criteria of anonymity do not simply include the fact that the original identifier can no longer be recreated with high probability, for example to prevent the re-creation of recognizable facial images, etc. This weaker property is true for some salted hashes, temporary random identifiers, and a wide range of other similar identifiers (referred to as pseudo-anonymity). In contrast, the present invention targets a significantly stricter level of anonymization by making it impossible for an attacker to use stored identifiers to link two or more data points at the individual level (while still being able to link at the aggregated statistical level), while also preventing, for example, the linking of data into profiles. This is also the definition of anonymization that is common in modern and stricter definitions provided by modern scientific and legal definitions of anonymity. In contrast, any availability or possibility of non-anonymous data (e.g., non-anonymous identifiers) that is linkable at the individual level would render achieving the goals described in the present invention unimportant and meaningless.

[0156] For example, one particular effect of the anonymization described herein may be to effectively prevent or significantly hinder any potential profiling of individuals by third parties using data stored in the system.

[0157] As an alternative to the methods of the present invention, data can be anonymized after collection while maintaining group mobility metrics in various ways, such as by microaggregating groups and storing group mobility for each group. However, such anonymization requires one or more non-anonymous data collection steps. Therefore, such a system and / or method for group mobility metrics would not be anonymous, as personal data would need to be collected and stored from each individual at least during the intervals between visits to the corresponding subject's state. This issue is also important enough to be explicitly recognized in legislation, for example:

[0158] "In order to show the movement of traffic in a particular direction over a period of time, an identifier is needed to link the location of individuals over a specific time interval. If anonymized data is used, this identifier will be lost and this movement cannot be shown."

[0159] These conclusions clearly did not anticipate the present invention and clearly demonstrate the perceived impossibility of achieving the stated goals with conventional methods while maintaining adequate anonymity.

[0160] Such non-anonymous data is incompatible with the data collection contemplated by the present invention, as the lack of anonymity in both its collection and storage makes this type of data incompatible with the goal of anonymously tracking and / or analyzing the movements of individual subjects.

[0161] The original identifiers may be unevenly distributed. This is the case, for example, with localized geographic biases in biometrically relevant phenotypes within a population. In such cases, the required uniform noise level may be prohibitively high. An improved and appropriate noise level for maintaining anonymity may need to depend on the identifiers themselves—for example, adding more noise to identifiers that are more likely to have few neighbors. This, however, requires estimating the underlying distribution of identifiers. This distribution estimation can be very difficult in practice and may also be subject to estimation errors that threaten anonymity.

[0162] For this case, an optional additional decorrelation module is proposed which aims to effectively remove any relevant correlations in the anonymous identifiers. For example, the optional additional decorrelation module uses a cryptographic hash and / or similar decorrelation function before adding noise to the resulting decorrelated identifiers in the anonymization module. The role of the decorrelation module is to remove any patterns and / or any large-scale patterns in the distribution, which will make the identifier density uniform, while the anonymity is provided by the noise in the anonymization module rather than decorrelation. In contrast to the hash function used to generate the group identifier, the decorrelation module itself does not need to provide the anonymous identifier. Therefore, the decorrelation module can also be truly reversible or potentially reversible, such as a reversible mapping or salted hash that allows data linking and / or reconstruction of the original identifier with some probability. Further description of the decorrelation aspects and possible uses of locality sensitive hashing in the decorrelation module follows the guidelines provided in the following related examples.

[0163] In an alternative example embodiment of the decorrelation module, a decorrelation function is applied to the noise instead. This means that a normally well-behaved noise source, such as Gaussian noise, is transformed into decorrelated noise, i.e. decorrelated noise having a probability distribution that effectively lacks large-scale continuous patterns, for example by applying a hash function to the well-behaved noise source. This decorrelated noise from such a decorrelation module can then be used to simultaneously anonymize and decorrelate the identification data, for example by adding the decorrelated noise and then applying a modulo rspan operation, where rspan is the image range of the noise source. Care needs to be taken when setting the numerical resolution of the noise and / or designing the hashing method used so that the noise is not completely uniformly distributed, as a non-uniform distribution is required to create the necessary identifier correlation bias used by the present invention.

[0164] As an alternative to the decorrelation module, a decorrelation deviation metric can be used. For example, this can be any deviation metric that does not exhibit large-scale patterns that may be related to the physical system, for example by being based on a function such as a random initialization table and / or a function that is weighted by correlation of effectively random identifiers and / or a function that only maintains small-scale patterns that are unlikely to cause significant correlation, such as a modular operation. The necessary considerations when designing a decorrelation deviation metric are very similar to those when designing a decorrelation module and will be apparent to the skilled person.

[0165] Decorrelation of identification data should be interpreted in the context of a deviation metric. If the deviation metric can be influenced by existing access probability patterns in the identification data, for example, if identifiers that affect a particular identifier density metric are on average more likely to access subject states than other identifiers in the population, then the access frequencies of the identification data can be considered to be correlated (with the shape of the deviation metric). Thus, the correlation can be broken by changing the deviation metric and / or the anonymous identifiers to break their correlation, while the access frequencies of each subject state and identifier can be considered to be given values of the measurement system. For example, since the probability of two completely random functions and / or distributions being significantly correlated is low, any choice of random mapping is sufficient to cause them to decorrelate with high probability.

[0166] In short, the theoretical reason for the effectiveness of decorrelation is related to the fact that data originating from the physical world and / or functions used to model this physical world (such as the most common and named functions used in engineering) form infinitesimal and specific subsets of all possible functions and have a relatively high probability of being similar and showing spurious correlations, especially for large patterns. Small-scale physical patterns are often at least partially chaotic and effectively random. Further details on this property can be found in earlier works published by the inventors (such as "Mind and Matter: Why It All Makes Sense"). In contrast, the probability of a function / distribution effectively randomly selected from all possible functions / distributions showing such correlations with functions of physical origin and / or other randomly selected functions is much lower, usually zero or negligible. The avalanche effect gives a different but similar perspective on decorrelation. For example, bent functions and / or functions that meet strict avalanche criteria may be suitable as functions for decorrelation purposes, while functions that are considered to be particularly well behaved and / or have low-valued derivatives are generally less suitable because their approximate linearity is related to the approximate linearity inherent in most physical systems and models at some scale. Cryptographic hash functions and random mappings such as random tables benefit from these properties, but for the purposes of the present invention, many other functions also possess and / or approximate (e.g., LSH) related properties. Suitable alternatives should be apparent to those skilled in the art of hashing, cryptography, and compression theory.

[0167] Note that the general application of adding noise in this paper as any random mapping does not necessarily rely on adding a noise term to the identifier. For example, multiplicative noise can also be used. From an information theoretic perspective, this can still be viewed as adding noise to the information encoded in the data, regardless of the form of this encoding.

[0168] The selection of a specific hash and / or noise mask identifier may vary between subject states and may also depend on other factors. For example, certain identifiers may be assigned to hashes and other identifiers may be assigned to noise-based masks. The noise may be identifier-dependent and / or subject-state-dependent.

[0169] In some cases, some accessible identification data is considered an identifier and other potential identification data is considered additional data that is unknown to the attacker. For example, precise location data in a public place cannot be used to identify an individual unless the attacker may have location data with the same timestamp. If such data is potentially available to an attacker, it may be appropriate to anonymize any additional data along with the identifier. The present invention can be used in any such combination. For example, a facial image can be used as both an identifier and an anonymous identifier stored by the present invention. Location data is stored along with the anonymous identifier to facilitate analysis of travel patterns. This additional location data can then be anonymized separately, for example by quantizing location and time into intervals large enough to render it anonymous. The resolution may be different in residential areas and in public places such as retail locations.

[0170] In general, the proposed invention can be applied to any sufficiently identifying portion of identification data (i.e., the identification itself) and the additional identification data can be anonymized by a separate method. These subject states can then be statistically linked by those identifiers processed by the invention, while the remaining identification data can be anonymized in a way that does not allow such statistical linking.

[0171] According to another aspect, a system for anonymously tracking and / or analyzing the flow or movement of individual subjects and / or objects (referred to as individuals) is provided.

[0172] In this non-limiting example, the system is configured to use information representing the identity of the individual as input and determine a group identifier for each individual in a group of multiple individuals based on a hash function. Each group identifier corresponds to a group of individuals whose identity information produces the same group identifier, thereby effectively micro-aggregating the group into at least two groups.

[0173] A noise masked identifier performs the same function by adding random noise with a distribution such that every possible noise masked identifier value is achievable by a large number of identifiers.

[0174] The system is further configured to keep track of visit data for each group, the visit data representing the number of visits to two or more spatiotemporal locations by individuals belonging to the group. More generally, the system is configured to keep track of a deviation measure of the two or more subject states.

[0175] The system is further configured to determine at least one group flow metric (for the entire group) of the plurality of individuals passing from the first spatiotemporal location to the second spatiotemporal location based on the visit data for each group identifier.

[0176] More generally, the system is configured to determine at least one population flow metric (for the entire population) of a plurality of individuals passing from a first subject state to a second subject state based on the deviation metric.

[0177] Exemplary references Figure 1A and / or Figure 11 The system may include processing circuitry 11, 110 and memory 15, 120, wherein the memory 15, 120 includes instructions that, when executed by the processing circuitry 11, 110, cause the system to anonymously track and / or analyze the flow or movement of individuals.

[0178] According to yet another aspect, the proposed technology provides a monitoring system 50 comprising the system 10 as described herein, such as Figure 6 Schematically shown in .

[0179] Figure 7 is a schematic flow chart illustrating a specific, non-limiting example of a computer-implemented method for achieving an estimate of the amount or number and / or flow of individuals in a population moving and / or coinciding between two or more spatiotemporal locations.

[0180] Basically, the method includes the following steps:

[0181] S21: receiving identification biometric data from two or more individuals (wherein the identification biometric data includes and / or is based on biometric data);

[0182] S22: generating, by one or more processors, a group identity (and / or noise masking identifier) for each individual that is not actually associated with group flow; and

[0183] S23: Store: the group identity (or more generally, a measure of deviation for each subject state) and data describing the spatiotemporal position; and / or a counter for each spatiotemporal position and group identity.

[0184] For example, the group identity may be generated by applying a hash function that effectively removes any pre-existing correlation between the identification data and trends at one or more spatiotemporal locations.

[0185] Optionally, the noise masking anonymization includes a decorrelation step that effectively removes correlations in the identifier space.

[0186] For example, the population of measured interview individuals may be an unknown sample from a larger population, where the larger population is large enough that the expected number of individuals in the larger population to be assigned to each group identity and / or noise-masked identifier is two or more.

[0187] The group of interview individuals may, for example, be considered a representative sample from the larger group, which may also be measured implicitly and / or explicitly through data collected from the interview group.

[0188] Optionally, the generation of group identities may be partially random on each application.

[0189] For example, for each individual, the identification data may include information representing the identity of the individual based at least in part on a biometric attribute of the individual. Non-limiting examples of such biometric information may include and / or be based on at least one of the following: an iris image, a facial image, a feature vector, a body image, a fingerprint, and / or a gait.

[0190] Figure 8 is a schematic flow chart illustrating another specific, non-limiting example of a computer-implemented method for estimating the amount or number of individuals in a population that coincide between two or more spatiotemporal locations.

[0191] In this particular embodiment, the method further comprises the steps of:

[0192] S24: Generate a group flow metric between the two spatiotemporal locations using the counter of the group identity at each of the two spatiotemporal locations.

[0193] For example, the generation of group flows can be based on a linear transformation of visit counters.

[0194] Alternatively, the linear transformation may include a correlation between a vector describing the group flow of each group identity in the first location and a vector describing the group flow of each group identity in the second location.

[0195] As an example, a baseline is subtracted from the correlation corresponding to the expected correlation between two vectors.

[0196] For example, the number of individuals in a group may be two or more per group identity.

[0197] Optionally, activity data representing one or more actions or activities of each individual may also be stored together with the corresponding group identity and data describing the spatiotemporal location, so that not only the spatiotemporal aspects but also the individual actions or activities can be analyzed and understood.

[0198] Figure 9is a schematic diagram illustrating an example of the movement or flow of one or more individuals from location A to location B. For example, this may involve individual subjects and / or objects moving from one location to another and being identified, for example, by a camera or other means, such as a person being identified by facial recognition, fingerprint and / or iris scan, and / or other biometric information.

[0199] Figure 10 A diagram illustrating an example of the movement or flow of a user from one virtual location, such as an IP location, to another virtual location. This may be an individual user moving from one internet domain to another, such as from IP location A to IP location B, and being identified, for example, by facial recognition, fingerprint and / or iris scan and / or other biometric information.

[0200] For example, biometric information may be obtained, for example, by using well-accepted techniques for extracting fingerprints, facial data, and / or iris data via laptops, personal computers, smartphones, tablet computers, and the like.

[0201] Figure 12 is a schematic flow chart illustrating an example of a computer-implemented method for generating a measure of the flow or movement of individual subjects and / or objects (referred to as individuals) between spatiotemporal locations based on biometric data.

[0202] Basically, the method includes the following steps:

[0203] S31: configuring one or more processors to receive a counter of anonymous and approximately independently distributed group identities derived from an individual's visits to each of two spatiotemporal locations, wherein the counter is based on biometric data;

[0204] S32: Using the one or more processors, generate a group flow metric between two spatiotemporal locations using a linear correlation between the counters of group identities for each of the two spatiotemporal locations; and

[0205] S33: Storing the group flow metric in a memory.

[0206] For a better understanding, various aspects of the proposed technology will now be described with reference to some basic key features followed by non-limiting examples of some optional features.

[0207] The present invention receives some identifying biometric data that can uniquely identify an individual and / or personal items of an individual with a high probability. The identifying biometric data can optionally be continuous data, such as biometric measurements. The identifying biometric data can also be any combination and / or function of such data from one or more sources.

[0208] In a preferred example, the present invention comprises an anonymization module comprising an (anonymous) hash module and / or a noise-based anonymization module.

[0209] Example - Hash Module

[0210] Some aspects of this invention relate to hashing modules. In our view, a hashing module is a system that retrieves identification data and generates some data about an individual's identity—data sufficient to identify an individual within a group much smaller than the entire population, but not small enough to uniquely identify the individual. This effectively partitions the population into groups of one or more individuals, i.e., performs automated online micro-aggregation of the population. Ideally, but not necessarily, these groups should be independent of the population flow being studied to simplify measurement. In other words, the goal is to partition the groups so that the expected flow for each group is approximately the same. Specifically, the variances of any pair of groups should be approximately independently distributed. In other words, it is desirable to be able to treat the groups as effectively random subsets of the population in statistical estimation. This can be achieved, for example, by applying cryptographic hashes or other hashes with the so-called avalanche effect. If local sensitivity is undesirable, a specific example of a suitable hash is a subset of the bits of a cryptographic hash such as SHA-2, sized to represent the desired number of groups corresponding to the desired number of individuals per group. In this example, a constant set of bits can be used for padding to achieve the necessary message length. However, this particular hashing example introduces some overhead in the computational requirements and a hashing module more suitable for this specific purpose could be designed, as the application herein does not require all cryptographic requirements.

[0211] Preferably, any correlation (whether linear or of another type) that could significantly bias the metrics generated from the system should be effectively removed by the hash module. As an example, a good approximation of a random mapping (such as a system based on a block cipher, a chaotic system, or pseudo-random number generation) can achieve this goal. At the extreme of minimalism, if it is considered unlikely that correlated identities will be created, a simple modular operation may be sufficient.

[0212] If the identifiers do not contain such correlation, for example if they are randomly assigned, then the hash will not benefit from decorrelation since any group assignment will be effectively random even without this decorrelation.

[0213] In some aspects of the invention, depending on the anonymity requirements, the number of groups can be set so that an expected two or more individuals from the group whose data has been retrieved, or two or more individuals from some larger group (which is effectively a random sample of the larger group), are expected to be assigned to each group. The invention allows for efficient unbiased estimation in both cases, as well as more extreme anonymous hashing schemes with a large number of individuals per group.

[0214] The hash keys representing the group identities can be stored explicitly (e.g. as numbers in a database) or implicitly (e.g. by having a separate list per hash key).

[0215] In other words, the hash module takes some identification data of the group and also generates randomly sampled subgroups from the entire group, for example, efficiently (i.e., a good enough approximation for the purposes herein). The hash module as described herein has several potential purposes: ensuring / guaranteeing decorrelation of data from group flows (i.e., using group identities that may be different from the identification data and not actually related to group flows) and anonymizing the data by micro-aggregating the data while maintaining some limited information about the identity of each individual. In some embodiments of the invention, as described in more detail below, the hash module may also maintain limited information about the data itself by using locality sensitive hashing.

[0216] For these aspects of the invention, the statistics collected for each group identity help to generate group flow statistics for the (entire) studied population comprising a plurality of such groups. It is not an object of the invention per se to measure differences between groups, and particularly if decorrelation intentionally generates rather meaningless group subdivisions by effectively removing any potential correlations between group members.

[0217] As an example of a suitable hash module, in a preferred embodiment, it is inappropriate to divide the group into groups based on the continuous range of one or more of the many meaningful variables (such as annual income, home location, IP range or height) because this may result in different expected group flow patterns for each group, which will require an estimate of the overall group flow to be measured. On the other hand, a limited number of bits from a cryptographic hash or a random mapping of any of these (multiple) criteria from an initial grouping to a sufficiently small range can be used to aggregate the effective random selection of small groups of such continuous ranges into larger groups. In other words, the identifier is divided into many small continuous ranges and the group is defined as some effective random selections of such continuous ranges so that each continuous range belongs to a single group. In this way, the group is divided into a set of groups that are actually indistinguishable from a random subset of the entire group because any large-scale pattern is effectively removed. Alternatively, a cookie can be saved on the user's computer. The cookie is a pseudo-randomly generated number within a certain range that is small enough that several users can expect to get the same number. Alternatively, these continuous ranges may also be replaced, for example, with further defined continuous n-dimensional ranges and / or non-uniquely mapped to specific groups with a similar effect for the purpose of the present invention, ie creating a suitable locality sensitive hash.

[0218] Random group assignment does not prevent the use of hashing methods and can also add meaningful additional anonymity. Biometric data typically contains a certain level of noise due to measurement errors and / or other factors, which makes any subsequent group assignment based on this data a random mapping as a function of identity. Random elements can also be added intentionally. For example, the system could simply roll a dice and assign individuals to groups according to a deterministic mapping 50% of the time, and assign individuals to completely random groups the other 50% of the time. As long as the distribution of this random assignment is known and / or can be estimated, the data can still be used in the system. Furthermore, in addition to the anonymity already provided by grouping, the simple dice strategy described above would be roughly equivalent to k-anonymity with k=2.

[0219] Example - Noise-based Anonymization

[0220] Some aspects of the present invention include a noise-based anonymization module. A noise-based anonymization module generates a new noise-masked identifier based on identification data. This module uses a random mapping where the output is irreversible due to added noise, rather than by limiting the amount of information stored. In other words, the signal remains below the identification limit, even if the total amount of information used to store the signal and noise is assumed to be greater than the limit. Any random mapping can be used such that linking the noise-masked identifier to a specific identity is impossible. Compared to a hash module, a noise-masked anonymization module produces an output with sufficient information content to uniquely identify an individual. However, a certain portion of this information is pure noise added by the anonymizer, and the actual information about the individual's identity falls below the threshold required to link data points with high probability at the individual level. While a hash module is preferred in most cases, a noise-masked identifier may more naturally match various noise identifiers and may also prevent certain deanonymization scenarios where an attacker knows an individual has been recorded.

[0221] Noise can be any external source of information that can be considered noise in the context of the present invention and does not imply a true noise source. For example, timestamps or values from some complex process, chaotic system, complex system, various pseudo-random numbers, media sources, and similar sources whose patterns are not reversible can be used. From the perspective of anonymity, it is important that the noise cannot be easily recreated and / or reversed, and the statistical purpose of the present invention further requires that the noise can be described by a certain distribution and does not introduce significant undesirable correlations that change the statistics.

[0222] Figure 13This figure illustrates an example of how identifier bias metrics can be anonymized by adding noise at one or more times, and how this can generate a bias compensation term. In this example, visit counters are used for subject states A and B, respectively. These population counters are randomly initialized, for example, before data collection begins. The bias compensation term is calculated by estimating the population flow from A to B due to spurious correlations in the initialization. These spurious correlations can then be removed from the population flow estimate to reduce the variance of the estimate. To further mask the initialization, additional small noise can optionally be added to the compensation term at the expense of a slightly increased variance in the population flow.

[0223] Figure 14 An example of noise masking anonymization is shown. It shows the probability density function of a noise masked identifier given a certain identifier. The probability density functions of two different identifiers are shown, and in this example, the probability density functions are approximately normally distributed around the identifiers. Not all possible input values correspond to individuals in the population and / or memory. In the event that the probability density functions from different identifiers overlap, the original identity that generated the noise masked identifier may not be known with certainty. Re-identification using a particular noise masked identifier becomes less likely as more overlap in the probability density functions from various identifiers is provided for that particular noise masked identifier, for example by having more identifiers in the population and / or memory.

[0224] Example - Anonymous Identifier

[0225] For example, anonymous identifiers are considered herein to be group identifiers and / or noise-masked identifiers.

[0226] In other words, an identifier in this context is, in a general sense, a specific instance of any type of identifying data and is not necessarily enumerable as a more narrow definition of the concept would suggest.

[0227] For example, persons assigned to the same group by the hash module can be considered as a hash group.

[0228] Example - Deviation Metrics

[0229] For example, data deviation in this context refers to how some particular data is distributed compared to the expectation from the generated distribution. A deviation metric is some information that describes the deviation of the collected data. In other words, the present invention measures how the actual identifier distribution differs from the expected identifier distribution, for example, the distribution when all individuals are equally likely to access two subject states. It is typically encoded as one or more floating point or integer values. The purpose of the deviation metric is to later compare between subject states in order to estimate how much of the deviation is common between the two subject states. A large number of different deviation metrics will be apparent to the skilled person. In fact, any deviation metric can be used in the present invention, although some deviation metrics retain more information about the data deviation metric than other deviation metrics and therefore may provide better deviation metric estimates.

[0230] Note that the deviation measure does not necessarily imply that the generated distribution is known, i.e., that enough information has been collected about the expected value of the generated distribution in order to calculate the deviation from the deviation measure. However, if the underlying distribution later becomes known, then the deviation measure already contains the information necessary to estimate the deviation of the data. That is, if the identifiers are decorrelated, for example using a decorrelation module, then the resulting generated distribution will be trivial to estimate.

[0231] The most basic example of a deviation metric is to maintain a list of the original access group identities or noise-masked identities along with any associated additional data. This provides anonymity but can be inefficient in terms of storage space because they contain redundant information. However, in some cases, maintaining such original anonymized identities allows for better optional post-processing, such as removing outliers, and greater flexibility in changing the deviation metric specifically for various purposes.

[0232] Another example of a simple deviation metric is an access counter. This access counter counts the number of identities detected in each subject state for each hash group. For example, this number could be a vector with the numbers 5, 10, 8, and 7 representing the number of access identities assigned to each of the four group identities in a particular subject state.

[0233] More generally, the deviation metric may consist, for example, of two or more sums and / or integrals over the convolution of: some mapping from the anonymous identifier space to a scalar value; and the sum of Dirac or Kronecker increment functions of the anonymous identifiers that access the subject state. In other words, the identifier distribution is measured in two different ways. In the specific case where the anonymous identifiers are discrete, such as an enumeration, and the corresponding mapping is a Dirac increment d(i) for i=1:n, this is equivalent to a visit counter. In other words, the deviation metric is a generalization of the anonymous visit counter. In other words, the deviation metric is two or more counts of the number of anonymous identifiers detected from some defined subset of the set of possible anonymous identifiers, where the counts can be weighted by any function that depends on the anonymous identifier. In other words:

[0234] sum_i f(x_i)

[0235] where x_i is an anonymous identifier that accesses the subject's state, i is some index of all anonymous identifiers that access the subject's state and f(x) is some mapping from the anonymous identifier space to (not necessarily positive) scalar values.

[0236] The above sum can be viewed as a density estimate for the access subpopulation. Because we are estimating the distribution of actual access identifiers, a finite and known population rather than a properly unknown distribution, we also use the less common but more precise term "density metric" to describe this quantity. The simplest density metric is the total access count, corresponding to equal weighting between identifiers. This total access count can be used in conjunction with another density metric to derive a very simple deviation metric. In a preferred embodiment, one hundred or more density metrics are used as vector-valued deviation metrics.

[0237] Alternatively, the deviation metric may consist of information representing one or more differences between such density metrics. For example, given two counts, one may simply store the difference between the two counts as the deviation metric.

[0238] In other words, a deviation metric is typically vector-valued data consisting of information representing the deviation of an identifier from the expected distribution of all identifiers sampled from some larger population.

[0239] This information can be encoded in any manner. While this method could theoretically work with only a single difference between two density metrics, it is most often preferred to rely on as many density metrics as the desired level of anonymity allows, in order to reduce the variance of the population. In a preferred embodiment of the hash module, 10-1,000,000,000 density metrics are used, depending on the size of the group of potential access identities and the expected size of the dataset. From another perspective, achieving an average anonymity level roughly equivalent to k-anonymization with k=5 is almost always desirable, and a stricter k=50 or higher is recommended in most cases.

[0240] The key realization of the utility of the method is that, using a large number of density measures and / or other informative bias measures, the flow metric can achieve surprisingly low variance while still preserving the anonymity of the individuals. Very small numbers of density measures are impractical for the stated purpose due to prohibitive variance, but this disadvantage disappears as the bias information encoded in the bias metric (e.g., the number of density measures used) increases.

[0241] For example, a visit counter for two or more spatiotemporal locations (also referred to as space-time locations) can be used. This keeps track of the number of times a person from each of two or more hash groups was detected at a certain spatiotemporal location (e.g., a certain web page, a specific street, a certain store, etc.) at a certain time (repeated or unique).

[0242] As mentioned above, a more general deviation metric than access counters is a set of identifier density metrics, also referred to herein as density metrics. Density metrics indicate the density of identifiers in the data according to some weighting. For example, the deviation metric can be a set of Gaussian kernels in the space of possible identifiers. Specifically, the density metric associated with each kernel can include the sum of weighted distances, i.e., a Gaussian function of the distance from the center of the kernel to each anonymous identifier. Two or more such density metrics from different Gaussian kernels, or one or more comparisons between such density metrics, will then represent a deviation metric. The identifier density metric can measure the density of identifiers for identifying data and / or anonymous data.

[0243] This density metric can be correlated between two points, as in the case of the visit counters used in some specific examples described herein, to estimate group flow. This is true even if the density metric is different, such as when different density metrics are used at points A and B. For example, the same approach as for visit counters can be used, i.e., using Monte Carlo and / or analytical estimation to establish the minimum and maximum expected correlations based on the number of eligible visitors.

[0244] For the purpose of providing anonymity, it is important that the anonymization is performed effectively online (or in real time and / or near real time) to provide an anonymous deviation measure, i.e. there is a continuous but short delay between obtaining the identifier and generating and / or updating the deviation measure. In a preferred embodiment, the hashing is performed within a general-purpose computer located in the sensor system or within a general-purpose computer that immediately receives the value. The value should not be accessible from the outside with reasonable effort before being processed. The identifier should be deleted immediately after processing. However, in a preferred embodiment, if this extended type of online processing is necessary for reasonable technical requirements and if it is not considered to substantially weaken the anonymity of the subject provided, the data can be batched at different points and / or otherwise processed in small time intervals (e.g., batched transmission at night) (if necessary). In contrast, offline methods are usually applied after the entire data collection is completed. Due to the storage of personal data, such offline methods cannot be considered anonymous.

[0245] Subject Status and Access

[0246] Group identities, noise-masked identities, and other deviation metrics, such as visit counters and / or any data related to group identities and / or noise-masked identities can optionally be modified in any way, such as by removing outliers, filtering specific locations, filtering group identities that coincide with known individuals, or by performing further micro-aggregation of any data.

[0247] The spatial aspect of the aforementioned spatiotemporal location can also be an IP address, domain name, virtual range of a frame, or similar aspects that describe the connection between an individual and a portion of the state of an electronic device, and the state of the individual's interaction with the electronic device. The broader definition of subject state also covers these aspects.

[0248] A subject state is a state and / or other meaningful description of an individual's spatiotemporal location, health, actions, finances, behavior, physical attributes, clothing, orientation, categories assigned by a classifier, immediate environment, and / or interactions with computers, network services, and / or other services. In other words, a subject state is a category that describes an individual in relation to himself / herself or to other entities.

[0249] A visit is a connection between an identifier and a subject’s state. For example, a visit could identify an individual being tested in a specific area at a specific time, an IP address filling out a web form, or a subject being tested for a disease.

[0250] A spatiotemporal location is any range in space and / or time, not necessarily continuous. For example, it could be the number of visits to a certain subway station on any Friday morning. A count could be any information about the number of individuals. For example, the information could simply hold a Boolean value that keeps track of whether at least one individual has visited the spatiotemporal location. In another example, the Boolean value could keep track of how many additional individuals in a group have been visited compared to the average for all groups. The Boolean value could also keep track of more specific location data, such as specific geographic coordinates and timestamps, which are aggregated into a larger spatiotemporal location at a later point in time. This specific data is then considered to also implicitly keep track of visits to the larger location. Figure 4 An example of a possible access counter is shown in .

[0251] Spatiotemporal location and spatiotemporal place may generally be considered synonymous in the context of this document and may include any defined range of space, time, and / or spacetime.

[0252] Subject states can also be defined using fuzzy logic and similar partial member definitions. This will typically result in partial accesses rather than integer values and is generally compatible with the present invention.

[0253] Example - Anonymous Group Flow Estimation

[0254] The flow measurement uses data from the deviation metric to measure the flow of individuals from one subject state (A) to another subject state (B). Because each hash group and / or density metric represents a large number of individuals, it is impossible to know exactly how many people in a group or population that are in A are also in B. Instead, the present invention uses higher-order statistics to generate noise measurements.

[0255] A measure of flow is an estimate of the number of people who visit subject states A and B in some way. For example, the measure can be the number of people who transition from state A to state B and / or the percentage of the number of people who transition from A to B. The measure can also be, for example, a measure of the number of people who visit A, B, and a third subject state C (where the people who also visit C can then be considered a subpopulation for the purposes of the present invention). In another example, the measure can be the number of people who visit A and B, regardless of which subject state is visited first. There are many varieties of such measures available. Independent of any correlation between corresponding identities between subject states, the number of people who visit A and the number of people who visit B are not considered in this article to be group flow estimates, but rather two group estimates corresponding to two locations.

[0256] Because visiting individuals form a subset of all individuals in a hypothetical larger population, the identity of the individuals visiting the subject state is biased compared to the estimated visitation rate for all individuals from that larger population. If the same individual is visiting states A and B, this bias can be measured using a corresponding bias metric. This metric is complicated by the fact that the theoretical underlying distribution of visitors to states A and B is not necessarily known. For example, states A and B may display similar data biases due to phenotypes in their geographic regions. Such correlations would be difficult or impossible to isolate from overlapping visitors.

[0257] Some types of identifiers are truly and / or approximately, randomly and independently assigned to individuals in a population, for example when random numbers are chosen as pseudo-anonymous identifiers. Such identifiers will not show data bias between A and B for reasons other than coincidence of individuals between locations. In other words, the estimated distribution of the hypothetical larger population is known. In other words, the identity of each individual is then effectively independently sampled and the distribution of the assignments is known. This means that the exact expected distribution of identifiers in A and B is known. Since the expectation is known, deviations from that expectation can also be estimated without the need to collect data and without introducing bias. Furthermore, the independence of identifier assignment also means that deviation metrics such as the specific deviation metrics discussed above (i.e., linearly dependent on the weighted sum and integral of each detected identity) will become analytically derivable mappings to the number of coincident individuals.

[0258] For example, if the mapping is linear, then virtually any scalar value that is linearly dependent on the deviation metric can be used to construct a flow estimate. For the specific case of a certain maximum correlation between individuals in subject states A and B, and for the specific case where the individuals in the two subject states are different individuals, it is also straightforward to estimate this linear value, for example using Monte Carlo methods or analysis. Due to the independence of the identifiers, a flow estimate can be easily constructed using a linear interpolation between the two values. For simplicity, the preferred embodiment uses the correlation between two deviation metrics of the same type.

[0259] Note that the group flow metric may depend on the total or relative number of individuals in A and B, depending on its form (e.g., questions such as whether it is expressed as a percentage of visitors and / or as a total amount), in which case it may also be necessary to collect the group flow metric for each subject state.

[0260] Any nonlinear case requires more analytical steps in its design and may be computationally more expensive, but is otherwise simple and functionally equivalent.The preferred embodiment is linear due to its simplicity and efficiency.

[0261] However, many types of identifiers are not even approximately randomly assigned, such as home address geolocation data. For example, these identifiers may be a priori correlated with the frequency of access to the subject's state. In these cases, the present invention may optionally use a decorrelation hash module for group identifiers and a decorrelation module for noise-masked identifiers to remove any undesirable correlations present in the identifier distribution and make the identifiers approximately independently generated from one another and functionally equivalent to random and independent assignments. Once this has been accomplished, flow metrics such as linear transformations can be easily constructed without prior knowledge of the initial distribution as described above.

[0262] Specific examples and preferred embodiments of generating group flow estimates can be found in the various examples below.

[0263] In a preferred embodiment, a baseline is established by estimating the expected number of visits per group, for example by dividing the total number of visits for all groups in the visit counter by the number of groups. Such an expected baseline may also include a model of bias, for example in the case where the expected bias of a sensor system and / or similar system used directly or indirectly to generate anonymous identifiers can be calculated based on factors such as location, recording conditions and recording time. Additionally, the baseline may be designed taking into account group behavior models, for example: the tendency of each individual to repeatedly visit a location and / or the behavior of visitors that were not recorded for some reason. By subtracting this baseline, the preferred embodiment arrives at the deviation of each set of data. For example, the deviation of data may refer to how some specific data is distributed compared to the expectation from the generated distribution.

[0264] For example, the correlation between the variances of each group in A and B represents a deviation from the joint distribution. Careful consideration by the inventors revealed that the measure of the number of individuals can be achieved by exploiting the fact that the group identities and probabilities of individuals going from A to B can be effectively considered to be independent and identically distributed, which can be guaranteed by the design of the hash module and / or the decorrelation module. For example, by relying on the assumption of independence properties and by using: knowledge of the random aspects of the hash module distribution (which can include models of any sensor noise, transmission noise, and other relevant factors, if applicable); and a behavioral model describing the distribution of the number of visits for each individual, etc., a baseline deviation of the joint distribution (e.g., a Pearson correlation coefficient equal to 0) can be created, which would be expected if the two groups visiting A and B were independently generated from a random perspective. Similar behavioral models and / or knowledge of the random distribution in the hash module can also be used to estimate the deviation of the joint distribution (e.g., a Pearson correlation coefficient equal to 1) when the two groups are composed of exactly the same individuals. For example, such bias for completely overlapping populations can be adjusted based on a sensor noise model, where the sensor noise model can depend on other factors such as knowledge of the sensor noise model, location, group identity, identifier noise, and / or randomness in the hashing process. In a simple example with a homogeneous group, a hash module that includes a 50% chance of each individual having a consistent group assignment (which would otherwise be randomly assigned across all groups) can double the population estimate for the same bias compared to an estimate for a 100% accurate hash module.

[0265] A statistical measure of the number of individuals can then be generated by, for example, performing a linear interpolation between the two extremes based on the actual deviations measured by comparing the deviation metric. Note that these steps are merely examples, but the independence assumption will result in a population flow metric that can be expressed as a linear transformation, such as that indicated in one aspect described herein. A skilled person can derive various specific embodiments and ways to design such specific embodiments from this example and other examples and descriptions herein.

[0266] In some cases, the identifiers are already decorrelated from the outset. For example, this may be the case for unique identifiers assigned via biometric templates with random unique identifiers, where the unique identifier is a truly random or nearly random number generated for each biometric template.

[0267] Without the decorrelation assumption made possible by the inherent design of the hash module and in the case of noise masking of the identifier by the decorrelation module, the complexity of generating such a metric is in many cases prohibitive. Note that this simplification not only simplifies the precise design process of the embodiment, but also leads to cheaper, faster and / or more energy-efficient methods and systems due to the reduction and / or simplification of the number of processing operations in the required hardware architecture.

[0268] The groups in this example do not necessarily need to have the same distribution a priori (e.g., have the same estimated group size). For different expected group sizes, the group flow estimate will directly affect the estimated value and (normalized) correlation of each group of counters. Any related estimate of the variance of the group flow measure may become more complicated, for example, if the groups are very different, any Gaussian approximation of the correlation distribution may not be valid.

[0269] Likewise, the density metric and / or other deviation metric may differ in a variety of ways.

[0270] For example, more complex subject states can also be defined in order to compute accurate population flow estimates. Identifier deviation measures such as group identities can, for example, be stored together with subject states as described above (i.e. with "original" subject states) and access orders (i.e. ordinals), which then allows computing population flows from the original subject states to the original states before and / or after each specific access of the subject. From the perspective of the present invention, this can be viewed as aggregating many individual new subject states (i.e. one subject state per ordinal and original subject state) into larger subject states (i.e. states before and after a specific access) and aggregating population flow estimates into larger population flows (i.e. population flows from all subject states before a specific access x in state B, summed over all recorded accesses x in state B). This more complex computation allows computing population flows from A to B with lower variance, but the larger number of subject states results in a smaller number of anonymous identities in each subject state, which may weaken the anonymity provided by the present invention.

[0271] Example - Locality Sensitive Hashing

[0272] Correlation in anonymous identifiers can often be avoided by decorrelation, but this is not always the case. A special case where this is often unavoidable is certain noisy continuous identifiers. For example, continuous measurements of biometric data can be hashed using locality-sensitive hashing (LSH), which allows continuous measurements containing sensor noise to be used for micro-aggregation purposes. Such hash functions can be approximately and / or effectively (but not completely) decorrelated. Any choice of a particular LSH requires a balance between its decorrelation properties and its locality-preserving properties. Even if such a hashing method largely decorrelates the data, it may still retain some residual small bias in the hash distribution due to any correlation between the biometric measurements and the prior trend of visiting a certain location (if such correlation is present at all in the original continuous distribution). The term ("err") in the baseline, which will be further explained below, can then be used to compensate for this residual correlation. Note that decorrelation, such as decorrelation from the avalanche effect in this setting, is not strictly used, but rather it is assumed that small-scale correlations arising from local sensitivity have little impact on the resulting statistics (in other words, the correlations are effectively removed). Specifically, any significant correlation between the data and the prior trend of visiting a certain location is likely to be a large-scale pattern. The LSH-based hash module is not limited to continuous data and can also be used for other data, such as integer values.

[0273] As a specific example of LSH, a locality sensitive hash can be designed by dividing the space of continuous identifier values into 30,000 smaller regions. Cryptographic hashing, random tables, and / or other methods can then be used to effectively randomly assign the 30 regions to each of the 1,000 group identifiers. This means that two effectively independently sampled noisy continuous identifiers received from an individual have a high probability of being assigned to the same group. At the same time, since each group consists of 30 independently sampled regions of the feature space, the difference between two different groups may be negligible. Decorrelation is generally effective if the region is much smaller than the correlation pattern of interest. For many well-behaved continuous distributions, noise resistance (i.e., the robustness of the variance of the group flow estimate to the presence of noise such as identifier / sensor noise) and effective decorrelation of groups can be achieved simultaneously. Since individuals may be assigned to different regions simply due to noise in the identification data, it may be beneficial to compensate for the estimate of the resulting randomness in the group identity assignment.

[0274] As an example of the above concept about LSH, people with a height of more than 120 cm are much less likely to enter a toy store than people with a height of less than 120 cm, while the corresponding prior differences between people with a height of 119.5-120 cm and people with a height of 120.0-120.5 cm may be negligible and therefore approximately irrelevant.

[0275] Note that the decorrelation module may also use LSH as described above to produce locally preserved identification values that do not actually have the type of correlation described above. Compared to the anonymization module, the difference is that the number of possible decorrelated identifier values is large enough to uniquely identify an individual from the value. For example, the collision probability of the decorrelation hash may be very low. There may be a certain resulting probability that does not correctly identify an individual, but is not enough to be considered anonymous (i.e., the decorrelation module is decorrelated but not anonymous). Randomness then becomes a necessary additional anonymization step to LSH in order to protect the identity of the individual.

[0276] It can be noted that for a large number of samples and a large number of possible hashes, the correlation between the two independent populations is approximately normally distributed. This makes it easy to also present confidence intervals for the generated metrics, if desired.

[0277] Example - Behavioral Model

[0278] The group flow can optionally be modified by a behavioral model to derive derived statistics, such as the flow of unique individuals if each location is visited repeatedly. For example, such a behavioral model can estimate the expected number of revisits for each individual. Such a behavioral model can also be estimated iteratively along with the group flow, for example in an estimation maximization process, where the group flow and behavioral model are repeatedly updated to improve the joint probability of the observed identifier distribution.

[0279] Example Implementations

[0280] In an exemplary preferred embodiment, the server in the exemplary system applies a hashing module to the received identifier and stores an integer between 1 and 1000, which is effectively random due to the avalanche effect. Assuming the number of individuals at A and B is 10,000 each, and assuming that individuals move only once per day in one direction and that there is no other correlation between the corresponding groups at A and B, the expected average for the two points is 10,000 / 1000 = 10 individuals per group. The measured number of individuals in each group can be encoded in integer-valued vectors n_a and n_b, respectively. The unit-length relative variance vectors v_a and v_b can now be calculated as v_a = (n_a - 10) / norm(n_a - 10), etc. (where the function norm(x) is the norm of the vector and subtracting a scalar from a vector means removing the scalar value from each component). Assuming that every individual who passes through A also passes through B within a day, perfect correlation is obtained, E[v_a*v_b] = 1 (where * is the dot product (if used between vectors) and E[] is the expectation). In contrast, assuming that the populations in A and B always consist of different individuals, the baseline can be estimated as E[v_a*v_b]=0, using the uncorrelated assumption made possible by the use of the hash module. Now assume that the number of individuals c3 at B consists of two groups of individuals: c1 from A (with relative variance vector v_a1) and c2 not from A (with relative variance vector v_a2). In this case, the expected correlation becomes E[c3*v_b*v_a1]=E[(c1*v_a1+c2*va2)*v_a1]=c1. This means that the expected number of individuals that can be measured from A to B is nab=v_b*v_a1*10000. Assuming that the scalar product between v_b and v_a is measured to be 0.45 in this example, measurements are obtained for 4500 individuals from A, or 45% of the individuals in B. In other words, unbiased measurements are obtained using strictly anonymized micro-aggregated data, which can be implemented as a linear transformation using a decorrelating hash module. The data generated by the hash module in this example can be considered anonymous without storing personal data and can be uploaded to any database. The calculations described herein can then be performed on a cloud server / database, preferably using lambda functions or other suitable computing options for performing the low-cost computations required for the linear transformation.

[0281] As part of generating the estimate, the counters and / or correlations may be normalized or rescaled in any manner. The various calculations should be interpreted in a general sense and may be performed or approximated using any of a large number of possible variations in the order of operations and / or specific subroutines that implicitly and effectively perform the same mapping between input data and output data as the calculations referred to herein in their narrowest sense. Such variations are obvious to the skilled person and / or automatically designed, for example by a compiler and / or various other systems and methods. In the case of slightly imperfect hash functions, the resulting error in the above assumptions can be partially compensated by assuming E[v_a2*v_b]=err, where err is some correlation in the data that can be estimated, for example, empirically by comparing two different independent samples from a population (i.e., measuring the flow at two points that have no correlation with each other). It is then expected that the following equation will be followed: c1=E[(c1*v_a1+c2*va2)*v_b]-err. This err term can, for example, be used as a baseline or part of a baseline.

[0282] Note that this simple case becomes slightly more complex when there are more people in A than in B. Even if all people in B come from A, less than perfect alignment in the group distributions is expected. This maximum expected scalar product can be easily estimated from the total number of visits to A and B. In these cases, the linear transformation used to obtain the estimate becomes a function of the total number of visits to A and B, respectively.

[0283] If noise is used to mask identifiers, one can simply divide the identifier space into regions and compute a density estimate for each region. Calculations similar to the visit counters described above can be performed on these density measures.

[0284] Example - Anonymous Bias Metrics

[0285] A potential problem with using any biased metric is that the subject state is initially weakly populated through accesses, and if the identifier is known, an attacker can probabilistically link the identity to a large number of data points.

[0286] For example, a visit counter might have a group with a single visit to subject state A, then it is reasonable to assume that the individual is the only individual registered in that group in the dataset, or more specifically, it is reasonable to assume that he / she is the only individual in A.

[0287] Alternatively, it might be reasonable to infer a group identifier from sparsely populated data at a given location (e.g., a known home address). This could then be checked against the work address. In this case, it would be possible to infer that he / she was indeed present at location B with high probability. This particular case could be addressed by storing only the deviation measure in location A and generating a group estimate online without storing the deviation measure from location B, i.e., using the deviation measure from A to update it each time B is visited. However, this approach would be ineffective if a group flow estimate from B to A also needed to be calculated.

[0288] A solution to these weakly populated states, and a potential anonymization solution in its own right, is to use an anonymity bias metric.

[0289] Anonymous bias measurement works by adding a degree of noise to the stored bias measure. This can be done, for example, before data collection begins or at any point during the collection process. This noise can bias the estimates of group flows. The resulting bias can be compensated by calculating an estimate based on the noise. More problematically, this also increases the variance of the group flow estimates.

[0290] An alternative, improved mechanism can be devised. In this mechanism, a bias derived from the specific noise sample used and / or other information suitable for generating such a bias based on the specific noise sample can also be generated. For example, a random number of "virtual" visits for each group identifier can be generated and added to a visit counter. The total group flow from A to B, estimated by the spurious correlation of all such virtual visits in A and B, and the total number of virtual visits to each location are also stored as bias terms. Since the correlation from the actual virtual visits is precisely known at the time of their generation, this correlation can also be accurately calculated and removed by the bias term. This approach significantly reduces the variance of the data, although some cross-terms caused by spurious correlations between actual and virtual visits may still contribute to the variance. Instead of directly storing the bias term, any information required to generate the bias term can be stored instead. If too much information about the noise is stored, the data may be deanonymized. However, the necessary bias term is a single value, while noise is typically vector-valued, so there are many possible ways to store sufficient data without storing sufficient information about the noise to deanonymize the data.

[0291] In the specific illustrative example of access counters encoded in vectors v_a and v_b, we have:

[0292] v_a=f+a+n_a

[0293] v_b=f+b+n_b

[0294] where a and b are unique visits to agent states A and B, respectively, and f is the general population. n_a and n_b are noise terms.

[0295] In this example, the various measures of group mobility are related to the following values:

[0296] E[v_a'*v_b]=E[f'*f]+2E[(a+b)'*f]+2E[a'*b]-2E[(a+f)'*n_b]+2E[n_a'*(b+f)]-n_a'*n_b'

[0297] Where * is the dot product and ' is the transpose of the vector.

[0298] Note that if the noise level is large, computing the noise term directly rather than estimating it may reduce the variance significantly, and thus particularly if the variance of the noise is larger than the variance of the other terms, for example if the access counters are sparsely populated. Mixed noise / data terms such as a'*n_a can also be computed exactly if the noise is added after the data, or can be partially computed and partially estimated if the noise is added at some point during data collection.

[0299] As a final safety measure, a small amount of noise can be added to the compensating bias term generated from the virtual visits. Typically a very small random number (e.g. between 0 or 1) is sufficient to mask any individual contribution to the deviation metric, even in exceptional cases where that individual contribution can be isolated from the deviation metric. When a large number of subject states are used, such noise on the bias term may prevent reconstruction of the deviation metric noise. Optionally, the noise is high enough that the exact number of visits for any identity cannot be inferred with a probability higher than 0.5. For example, if the noise is generated based on a random integer number of visits for each group identifier, then the probability of any such specific number of visits for each group identifier should ideally be 0.5 or less.

[0300] Practical memory storage limitations often limit the range of noise that can be used. However, this is more of a theoretical problem if the probability of generating small values is high and the probability of adding larger noise is increasingly small. This lacks any effective maximum value unless the probabilities are negligible. For example, a probability density function that decays exponentially with the noise size can be used. This noise preferably has an expected value of 0 to avoid reaching high values with multiple additions of the noise. In other words,

[0301] p(x)=k1*exp(-k2 x)-k3

[0302] for some constants k1, k2, and k3 and x is greater than or equal to 0.

[0303] This can be removed by using the stored virtual visit counts for each subject state when calculating the population mobility percentage and total visit counts.

[0304] The above addition generally generates a new bias metric based on the bias metric and noise, but the actual addition is preferred because it is easy to isolate as a bias term for subsequent accurate correction.

[0305] Deviation metrics rendered anonymous by adding noise can be considered sufficient to provide anonymity without using an anonymization module. This is true even if the noise is only used once as an initialization before data collection. The weakness is that if the anonymized data can be accessed at two points in time, the number of visits to any particular individual between those points can be easily extracted.

[0306] Another alternative is to add this noise after each visit. The resulting approach is thus more or less equivalent to the noise masking anonymization module. Note that the approach described above for using instantaneous knowledge of the noise to generate accurate bias corrections in group flow estimates can also be applied to the noise masking anonymization module and / or the hashing module.

[0307] This approach can also be used in case of continuous deviation measures such as storing exact continuous identifiers.Such noise in the deviation measure can be generated, for example, based on a sufficient number of virtual visits that are indistinguishable from individual visits.

[0308] For most applications, a preferred embodiment is a combination of an approach with an initial anonymous noise bias metric with a stored bias correction term generated from a specific noise sample, combined with a bias metric generated by a hash module (e.g., a group identifier counter). If the accuracy of the group flow estimate is more important than anonymity, a random initialization that relies only on the identity bias metric may be more suitable for reducing variance.

[0309] A disadvantage of all noise-based methods is that there may be few true noise sources and many pseudo-random noise sources can be inverted, which significantly simplifies attacks on anonymization.

[0310] At a mechanical level, this measured anonymity deviation is generated by an anonymization module, typically online, partly from received identifiers and partly from identifier deviation metrics already stored in memory. Noise can be added by the anonymization module and / or a separate mechanism that adds noise to memory. If the noise level is high enough, each new identifier deviation metric generated based in part on this noisy identifier deviation metric can be rendered anonymous.

[0311] In the following, a non-exhaustive number of non-limiting examples will be outlined.

[0312] Example - anonymously tracking and / or analyzing visitor movement within a physical or online retail environment based on biometric data.

[0313] For example, a system, method, and computer program for anonymously tracking and / or analyzing visitor movement in a physical or online retail store are provided.

[0314] The system is configured to determine a group identifier for each retail store visitor in a set or group of multiple visitors based on a hash function using information representing the visitor's identity as input,

[0315] Each group identifier corresponds to a group of visitors, and identity information of the group of visitors generates the same group identifier, thereby effectively micro-aggregating the set or group of visitors into at least two groups.

[0316] The system is configured to keep track of visit data for each group, the visit data representing the number of visits to two or more spatiotemporal locations by visitors belonging to the group, and the system is also configured to determine at least one flow metric representing the number of retail store visitors passing from the first spatiotemporal location to the second spatiotemporal location based on the visit data for each group identifier.

[0317] Also provided is a method, system, and corresponding computer program for enabling estimation of a measure of flow or movement of retail store visitors in a collection or group of visitors between two or more spatiotemporal locations.

[0318] In an example, the method includes the following steps:

[0319] - receiving identification biometric data from two or more retail store visitors, wherein the identification data includes and / or is based on the biometric data;

[0320] - generating, online and by one or more processors, a group identity for each visitor that is not actually associated with group mobility (e.g. based on corresponding identifying biometric data); and

[0321] -Storing: a group identity for each visitor and data describing the spatiotemporal location; and / or a counter for each spatiotemporal location and group identity.

[0322] More generally, the method comprises the following steps:

[0323] - receiving identification data from two or more visitors, wherein the identification data includes and / or is based on biometric data;

[0324] - Generate an anonymous identifier for each visitor online and through one or more processors; and

[0325] -Store: an anonymous identifier for each visitor and data representing subject status; and / or deviation metrics for such anonymous identifiers.

[0326] Further, a method, system, and corresponding computer program are provided for generating a metric of the flow or movement of retail store visitors between spatiotemporal locations.

[0327] In this example, the method includes the following steps:

[0328] - configuring the one or more processors to receive a counter of anonymous and approximately independently distributed group identities derived from visits by retail store visitors to each of the two spatiotemporal locations, the counter being based on biometric data;

[0329] - generating, using the one or more processors, a group flow metric between two spatiotemporal locations using a linear correlation between counters of group identities for each of the two spatiotemporal locations;

[0330] - Storing said population flow metric in a memory.

[0331] More generally, the method comprises the following steps:

[0332] - configuring one or more processors to receive anonymous identifier deviation metrics generated based on biometric-based identifiers from visits by a visitor to each of two spatiotemporal locations or subject states and / or the visitor's presence in each of the two spatiotemporal locations or subject states, wherein each identifier represents an identity of an individual visitor and includes and / or is based on biometric data;

[0333] - generating, using the one or more processors, a population flow measure between two spatiotemporal locations or subject states by comparing the anonymous identifier deviation measures between the spatiotemporal locations or subject states;

[0334] - Storing said population flow metric in a memory.

[0335] Additional optional aspects as previously described may also be incorporated into the technical solution.

[0336] Similar systems and / or methods can also be used for purposes such as analyzing movement or flow in smart cities, public events, public transportation, security surveillance, buildings, airports, etc. For example, security cameras and / or specially installed cameras can be used to analyze the movement patterns of individuals. Such cameras can also use infrared, stereo vision, and other similar technologies to improve biometric measurements and / or more accurately locate individuals.

[0337] In another example, a camera is used in a retail environment to retrieve images containing facial image data. A face detector neural network is used to identify any facial locations. Faces are extracted from the images and a neural network-based hashing module is applied to create a group identifier for each face in the range of 1-1000 integers. The group identifier is stored along with an anonymous timestamp and location (e.g., area 3 in store 2). Optionally, additional data, such as activity, is stored along with the location, allowing for statistics not only on location (and time), but also on series of actions taken by customers or other similar events and / or situations. Correlations between normalized vectors of group counters at different locations and / or times can be used to measure how visitors move between or within stores, how many customers return to stores over various time spans, and how exposure to certain visual messages influences purchase propensity (e.g., by using proxies such as those seen on cameras near cash registers to estimate purchase propensity). Alternatively, facial images collected online from viewers of a digital marketing campaign (e.g., images retrieved from social media profiles) can be converted into anonymous group identifiers and associated with subsequent visits and / or actions in stores to anonymously measure the effectiveness of the digital marketing campaign.

[0338] Note that biometric data in this context refers to data that can theoretically be used to identify a person with a high probability. This is different from certain legal definitions, where image data, etc., is considered biometric data only when it is actually used or intended to be used for identification purposes. For example, facial images are considered biometric data in this context even if they are not intended to be used for identification.

[0339] For example, similar systems could be used to track people in smart cities, airports, security, and / or public transportation environments.

[0340] In another more complex example, wearable devices are used to automatically collect blood pressure data on a monthly basis. Blood pressure is divided into enumerable intervals, and self-reported dietary components are reported using a mobile app and categorized into multiple categories. The combination of blood level and diet serves as the subject's state. Upon self-reporting, the subject takes a photo, and a facial recognition neural network is used to generate an identifying facial recognition feature vector. The feature vector is hashed using a decorrelation module consisting of LSH, which enumerates multiple locations larger than the population size to produce a decorrelated hash with high probability of re-identification. An anonymization module is then used to anonymize the identifiers of those subjects who do not consent to the use of their personal data. The anonymization module then adds an integer drawn from an approximate Gaussian distribution of integer values to this enumeration. If the number is larger than the maximum population, a modulo operation is applied, generating a type of noise-masked identifier. The Gaussian distribution is chosen so that the distribution of each original integer is overlapping, making identification using the noise-masked identifier impossible. The noise-masked identifier is stored along with the subject's state and a description of the camera type and resolution used to take the photo. A vector counting the number of individuals for each noise-masked identifier and subject state serves as a bias metric. Randomly generated eigenvectors uniformly distributed across the feature space are then used to estimate the maximum and minimum correlations between the two states, depending on whether the states have independent or overlapping populations. These eigenvectors are fed into a decorrelation module, an anonymization module, a Monte Carlo estimate of the consent state, and a camera correlation model for eigenvector noise depending on the camera type and resolution. In other words, the Monte Carlo estimate is used to generate the parameters of a linear transformation that, when applied to the actual identifiers, generates population flow estimates. For those subjects who did not consent, these flow estimates are then used to anonymously investigate the impact of diet on blood pressure development by creating models of how subjects with each combination of diet and blood pressure flowed towards various blood pressure states over the following month, where diet was not used to distinguish between states in this second state.

[0341] It is also possible to divide the entire population into subpopulations of interest. For example, before applying the hash, patients can be divided into subpopulations such as male / female, age, region, etc. For the purposes of this article, each subpopulation is then considered a separate population under study, even though the same hash function may be shared between several subpopulations. This information can be stored as a separate counter, or additional information can be stored explicitly with the group identifier.

[0342] In each of these examples, multiple visits by the same individual would be indistinguishable from multiple visits from different individuals. Therefore, if an exact number of unique individuals is desired, then, as an example, a behavioral model could be combined with the generated metric. For example, one might look at correlations in time between a few different times for the same location and measure the average number of repeat visits per visitor. For example, as indicated in the more general description, such a behavioral model could then be used to compensate an advertising revenue model by dividing the total number of visits by the number of repeat visits and thus generating a metric for the number of unique visitors. Many other types of behavioral models could also be fit to the data using the general approach described herein and complex behavioral models could result from a combination of several such sub-models.

[0343] A specific example of a behavioral model used to derive unique visitors can be used to compensate for repeat visits that are more likely to occur within a short time interval. In these cases, visits from the same group within a certain time interval may be compensated or filtered. For example, two visits to the same location within 5 minutes may be considered a single visit or a fraction (such as 0.01 of a visit) based on some approximation of the probability that these visits are two separate identities.

[0344] It is also possible to divide the entire group into subgroups. For example, before applying the hash, visitors can be divided into subgroups such as male / female, age, region, etc. Each subgroup is then considered a separate group under investigation, even though the same hash function may be shared between several subgroups. This information can be stored as a separate counter, or additional information can be stored explicitly along with the group identity.

[0345] The above examples are not exhaustive of the possibilities.

[0346] Example - Implementation Details

[0347] It will be understood that the above-described methods and apparatus may be combined and rearranged in various ways, and that the methods may be performed by one or more appropriately programmed or configured digital signal processors and other known electronic circuits (e.g., discrete logic gates interconnected to perform specific functions, or application-specific integrated circuits).

[0348] Many aspects of the invention are described in terms of sequences of actions that may be performed by, for example, elements of a programmable computer system.

[0349] The above-described steps, functions, processes and / or blocks may be implemented in hardware using any conventional technology, such as discrete circuit or integrated circuit technology, including both general-purpose electronic circuitry and application-specific circuitry.

[0350] Alternatively, at least some of the above steps, functions, processes and / or boxes may be implemented in software to be performed by a suitable computer or processing device such as a microprocessor, a digital signal processor (DSP) and / or any suitable programmable logic device such as a field programmable gate array (FPGA) device and a programmable logic controller (PLC) device.

[0351] It should also be understood that the general processing power of any device implementing the present invention can be reused.Existing software can also be reused, for example by reprogramming the existing software or by adding new software components.

[0352] Solutions based on a combination of hardware and software can also be provided. The actual hardware-software partitioning can be determined by the system designer based on many factors including processing speed, implementation cost and other requirements.

[0353] Figure 11 1 is a schematic diagram illustrating an example of a computer implementation 100 according to an embodiment. In this particular example, at least some of the steps, functions, processes, modules, and / or blocks described herein are implemented in computer programs 125, 135, which are loaded into memory 120 for execution by a processing circuit system including one or more processors 110. The processor(s) 110 and memory 120 are interconnected to enable normal software execution. Optional input / output devices 140 may also be interconnected to the processor(s) 110 and / or memory 120 to enable input and / or output of relevant data, such as input parameter(s) and / or derived output parameter(s).

[0354] The term "processor" should be interpreted in a general sense as any system or device capable of executing program code or computer program instructions in order to perform a specific processing, determination or computing task.

[0355] The processing circuitry including the one or more processors 110 is thus configured to perform well-defined processing tasks such as those described herein when executing the computer program 125 .

[0356] In particular, the proposed technology provides a computer program comprising instructions that, when executed by at least one processor, cause the at least one processor to perform the computer-implemented method described herein.

[0357] The processing circuitry need not be dedicated to performing only the above-described steps, functions, processes, and / or blocks, but may also perform other tasks.

[0358] Furthermore, the present invention may be considered to be fully embodied in any form of computer-readable storage medium having stored therein an appropriate set of instructions for use by or in conjunction with an instruction execution system, apparatus or device (such as a computer-based system, a system containing a processor, or other system that can retrieve instructions from the medium and execute the instructions).

[0359] The software may be implemented as a computer program product, which is typically carried on a non-transitory computer-readable medium such as a CD, DVD, USB memory, hard drive, or any other conventional storage device. The software may thus be loaded into the operating memory of a computer or equivalent processing system for execution by a processor. The computer / processor need not be dedicated to performing only the steps, functions, processes, and / or blocks described above, but may also perform other software tasks.

[0360] When executed by one or more processors, one or more flowcharts presented herein can be viewed as one or more computer flowcharts. The corresponding apparatus can be defined as a set of functional modules, where each step performed by the processor corresponds to a functional module. In this case, the functional modules are implemented as computer programs running on the processor.

[0361] The computer program resident in the memory may thus be organized into suitable functional modules that are configured to perform at least a portion of the steps and / or tasks described herein when executed by the processor.

[0362] Alternatively, the module(s) may be implemented primarily through hardware modules or alternatively through hardware with appropriate interconnections between related modules. Specific examples include one or more appropriately configured digital signal processors and other known electronic circuits (e.g., discrete logic gates) interconnected to perform specific functions and / or application specific integrated circuits (ASICs) as previously mentioned. Other examples of usable hardware include input / output (I / O) circuitry and / or circuitry for receiving and / or transmitting signals. The scope of software versus hardware is purely an implementation choice.

[0363] Providing computing services (hardware and / or software) is becoming increasingly common, where resources are delivered as services over a network to a remote location. For example, this means that the functionality described herein can be distributed or relocated to one or more separate physical nodes or servers. The functionality can be relocated or distributed to one or more co-acting physical and / or virtual machines that can be located in a separate (multiple) physical nodes (i.e., the so-called cloud). Sometimes also referred to as cloud computing, this is a model for enabling ubiquitous, on-demand network access to a configurable pool of computing resources such as networks, servers, storage devices, applications, and conventional or customized services.

[0364] The above embodiments should be understood as several illustrative embodiments of the present invention. Those skilled in the art will appreciate that various modifications, combinations, and variations of the embodiments may be made without departing from the scope of the present invention. Specifically, different partial solutions in different embodiments may be combined in other configurations, where technically possible.

Claims

1. A system (10; 100) for anonymously estimating group mobility, comprising: - one or more processors (11; 110); an anonymization module (12) configured to perform the following operations via the one or more processors (11; 110): for each of a plurality of individuals comprising individual subjects and / or objects in a population of individuals, receive identification information representing the identity of the individual, wherein the identification information representing the identity of the individual comprises and / or is based on biometric data, and generate an anonymous identifier deviation measure based on the identification information of the one or more individuals, wherein the anonymization is performed in real time such that the identification information is deleted immediately after processing; - a memory (15; 120) configured to store at least one anonymous identifier deviation metric based on at least one of the generated identifier deviation metrics; - an estimator (13) configured to perform the following operations via the one or more processors (11; 110): receive a plurality of anonymous identifier deviation measures, at least one identifier deviation measure for each of at least two subject states of an individual from the memory and / or directly from the anonymization module, and generate one or more population flow measures related to individuals passing from one subject state to another subject state based on the received anonymous identifier deviation measures.

2. The system of claim 1, wherein: Each identifier deviation metric is generated based on two or more identifier density estimates and / or based on one or more values generated based on the identifier density estimates.

3. The system of claim 1 or 2, wherein: Each identifier deviation metric represents the deviation of the identification information of one or more individuals compared to the expected distribution of such identification information in the population.

4. The system of claim 3, wherein: The identifier bias metric of this anonymization module is based on group identifiers that represent a large number of individuals.

5. The system of claim 4, wherein: The identifier deviation metric is based on a visit counter.

6. The system of claim 5, wherein: The identifier deviation metric is generated based on the identification information using a hash function.

7. The system of claim 6, wherein: The one or more population flow metrics include the number and / or rate of visitors passing from one spatiotemporal location to another spatiotemporal location.

8. The system of claim 7, wherein: At least one of the one or more population flow metrics is generated based at least in part on a linear transformation of counter information of two or more access counters.

9. The system of claim 8, wherein: The anonymization module (12) and / or the identification information representing the identity of the individual are random, and the randomness of the identification information and / or the anonymization module (12) is taken into account when generating the linear transformation.

10. The system of claim 9, wherein: When generating this population flow metric, a baseline corresponding to the expected correlation from two independently generated populations was subtracted.

11. The system of claim 1, wherein: Each identifier bias metric is generated using a combination of that identifier and noise, such that contributions to that identifier bias metric are rendered anonymous due to a sufficient level of noise that access to the subject's state cannot be attributed to a specific identifier.

12. The system of claim 11, wherein: The identifier bias metric is based on two or more identifier density estimates.

13. The system of claim 12, wherein: - the anonymization module is configured to generate at least one identifier deviation metric based on the anonymous identifier deviation metric stored in the memory; and - Anonymity is provided by adding sufficient noise to the anonymous identifier deviation metric stored in memory at one or more time instants so that the total contribution from any single identifier cannot be determined.

14. The system of claim 13, wherein: Information about the generated noise samples is also stored and used to reduce the variance of the group flow metric.

15. The system of claim 14, wherein: The identification information representing the identity of the individual includes and / or is based on at least one of: an iris image, a facial image, a biometric vector and / or a body image, a fingerprint and / or a gait.

16. The system of claim 15, wherein: These subject states include spatiotemporal location, computer system state interacting with a user, and / or health state and health monitoring state of the subject.

17. The system of claim 16, wherein: These subject states are spatiotemporal positions or locations, and The anonymization module (12) is configured to generate a group identifier based on the identification information of the individual to micro-aggregate the group into corresponding groups; wherein the memory (15; 120) is configured to store a visit counter (16) for each of two or more group identifiers from each of two or more spatiotemporal locations or places associated with the corresponding individual; and The estimator (13) is configured to receive counter information from at least two visit counters and generate one or more group flow metrics associated with individuals passing from one spatiotemporal location to another spatiotemporal location.

18. The system of claim 17, wherein: The anonymization module (12) is configured to generate a group identifier based on identification information of the individual by using a hash function.

19. The system of claim 17 or 18, wherein: The system (10, 100) includes an input module (14; 140) configured to perform the following operations via the one or more processors (11; 110): for each of the plurality of individuals, receive location data representing a spatiotemporal location, and match the spatiotemporal location of the individual with a visit counter corresponding to the group identifier associated with the individual, wherein each visit counter of each group identifier also corresponds to a specific spatiotemporal location.

20. A system (10; 100) for anonymously tracking and / or analyzing the flow or movement of individual subjects, referred to as individuals, and / or objects between subject states based on biometric data, in, The system (10; 100) is configured to determine an anonymous identifier for each individual in a group of a plurality of individuals using as input identification information representing the identity of the individual, wherein the identification information representing the identity of the individual includes and / or is based on biometric data, wherein each anonymous identifier corresponds to any individual in a set of individuals whose identity information produces the same anonymous identifier with a probability such that no individual generates the anonymous identifier with a probability greater than the sum of the probabilities of all other individuals generating the identifier, and wherein anonymization is performed in real time such that the identifying information is deleted immediately after processing, wherein the system (10; 100) is configured to keep track of deviation metrics, one deviation metric for each of two or more subject states, wherein each deviation metric is generated based on anonymous identifiers associated with the corresponding individuals associated with a particular corresponding subject state; and Therein, the system (10; 100) is configured to determine at least one population flow metric representing the number of individuals passing from a first subject state to a second subject state based on the deviation metrics corresponding to the subject states.

21. The system of claim 20, wherein: These anonymous identifiers are group identifiers and / or noise masked identifiers.

22. The system of claim 20 or 21, wherein: The system (10; 100) is configured to determine a group identifier for each individual in a group of a plurality of individuals based on a hash function using information representing the identity of the individual as input, Each group identifier corresponds to a group of individuals, and the identity information of the group of individuals generates the same group identifier, thereby micro-aggregating the group into at least two groups. wherein the subject states are spatiotemporal locations or places and the deviation metrics correspond to visit data, and the system (10; 100) is configured to keep track of each set of visit data representing the number of visits to two or more spatiotemporal locations by individuals belonging to the set, and Therein, the system (10; 100) is configured to determine at least one group flow metric representing the number of individuals passing from a first spatiotemporal location to a second spatiotemporal location based on the access data for each group identifier.

23. The system of claim 22, wherein: The system (10; 100) comprises a processor (11; 110) and a memory (15; 120), wherein the memory comprises instructions that, when executed by the processor, cause the system to anonymously track and / or analyze the flow or movement of an individual.

24. A monitoring system (50) comprising the system (10) according to any one of claims 1 to 23.

25. A computer-implemented method for enabling anonymous estimation of the amount and / or flow of movement and / or coincidence of individual subjects and / or objects in a population, referred to as individuals, between two or more subject states based on biometric data, the method comprising the steps of: - receiving identification information from two or more individuals, wherein the identification information for each individual includes and / or is based on biometric data; - generating an anonymous identifier for each individual online and by one or more processors, wherein anonymization occurs in real time, thereby deleting the identifying information immediately after processing; and - Storing: an anonymous identifier for each individual and data representing the subject's status; and / or a measure of deviation of such anonymous identifiers.

26. The method of claim 25, wherein: The anonymous identifier is an anonymous identifier deviation measure or other anonymous identifier that is not actually related to group mobility.

27. The method of claim 26, wherein: The identification information is associated with the population flow, and wherein the deviation metric is decorrelated and / or the anonymous identifier is generated using a decorrelation module and / or a decorrelation hash module.

28. The method of claim 27, wherein: The anonymous identifier is an anonymous deviation metric and the anonymous deviation metric is generated based on stored anonymous identifier deviation metrics to which noise has been added at one or more time instances.

29. The method of claim 28, wherein: The anonymous identifier is generated by adding noise to the identification information.

30. The method of claim 29, wherein: A compensation term to be added to the group flow estimate and / or necessary information for generating such a group flow estimate is calculated based on the one or more generated noise samples used by the method.

31. The method of claim 30, wherein: Any two stored anonymous identifiers or identifier deviation measures are unlinkable to each other, i.e. there are no pseudo-anonymous identifiers that link states in the stored data.

32. The method of claim 31, wherein The anonymous identifier is a group identity, and the group identity for each individual is stored together with data representing the subject's state; and / or a counter for each subject's state and group identity.

33. The method of claim 32, wherein: The subject state is a spatiotemporal location, a state of a computer system interacting with a user, and / or a health state and / or health monitoring state of the subject.

34. The method of any one of claims 32 to 33, wherein Activity data representing one or more actions or activities of each individual is also stored along with the corresponding group identity and data describing the subject's state.

35. The method of claim 34, further comprising the step of generating a population flow metric between two subject states.

36. A computer-implemented method for generating a measure of the flow or movement of an individual subject and / or object between subject states, referred to as an individual, based on biometric data, the method comprising the steps of: - configuring one or more processors to receive anonymous identifier deviation metrics generated in real time based on biometric-based identifiers from the individual's visits to each of the two subject states and / or the individual's presence in each of the two subject states, Each identifier represents the identity of an individual and includes and / or is based on biometric data and is deleted immediately after processing; - generating, using the one or more processors, a population flow metric between two subject states by comparing the anonymous identifier deviation metrics between the subject states; - Storing said population flow metric in a memory.

37. The method of claim 36, wherein: These subject states are spatiotemporal location, computer system state interacting with a user, and / or health state and / or health monitoring state of the subject.

38. The method of any one of claims 36 to 37, wherein These anonymous identifier deviation metrics are counters of group identities.

39. The method of any one of claims 36 to 37, wherein A single visitor appearing in one subject state cannot be re-identified with high probability in another subject state using these anonymous identifier deviation measures.

40. The method of any one of claims 36 to 37, wherein The generation step is not based on data that already contains some measure of population mobility between locations at the individual level and / or micro-aggregate level.

41. The method of any one of claims 36 to 37, wherein These anonymous identifier bias measures are not actually correlated with this group mobility.

42. The method of any one of claims 36 to 37, wherein The group flow estimate is generated based on a linear mapping from these anonymous identifier deviation measures.

43. The method of any one of claims 36 to 37, wherein The group flow metric is further generated based on information of noise samples used to anonymize the data.

44. The method of any one of claims 36 to 37, wherein The configuring step includes configuring one or more processors to receive counters of anonymous and approximately independently distributed group identities derived from individual visits to each of two subject states; and the generating step includes using a linear correlation between the counters of group identities for each of the two subject states to generate a group flow metric between the two subject states.

45. The method of claim 44, wherein: The subject states are spatiotemporal locations, and the group flow metric between two spatiotemporal locations is generated using a linear correlation between the counters of the group identities of each of the two subject states.

46. The method of any one of claims 36 to 37, wherein The anonymous identifier or identifier bias measure for each subject state is based on two or more identifier density estimates.

47. A computer program product comprising instructions which, when executed by at least one processor (110), cause the at least one processor (110) to perform the computer-implemented method of any one of claims 25 to 46.

48. A system for performing the method of any one of claims 25 to 46.

Citation Information

Patent Citations

  • A system and method for storing and controlling access to behavioural data

    GB201607522D0

  • Anonymous biometric player tracking

    US20130137516A1