Obfuscation and Deletion of Personal Data in Loosely Coupled Distributed Systems

By replacing and managing personal identifiers with fuzzy values ​​and reversible mapping techniques in telemetry data, the complexity of personal data removal in telemetry data in the prior art is solved, and efficient analysis of telemetry data and secure deletion of personal data is achieved.

CN112771525BActive Publication Date: 2025-05-27MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980032545.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-12-10
Filing Date
2019-05-09
Publication Date
2025-05-27
Estimated Expiration
2039-05-09

AI Technical Summary

Technical Problem

The prior art is difficult to effectively find and remove personal data in telemetry data, especially when the data is stored in multiple different locations, which leads to the complexity of data access and deletion.

Method used

By replacing the personal identifier contained with the telemetry data using fuzzy values, and using reversible mapping techniques, the fuzzy personal identifier is mapped back to the original value when needed, while using password hash functions and deletion tables to manage the deletion requests of the personal identifier.

Benefits of technology

The ability to analyze telemetry data without citing personal identifiers and effectively remove personal identifiers from reversible maps through relatively cheap constant time operations, simplifying the removal process of personal data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112771525B_ABST
    Figure CN112771525B_ABST
Patent Text Reader

Abstract

A real-time event processing system receives event data that includes telemetry data and one or more personal identifiers. The personal identifiers in the event data are replaced with obfuscated values so that the telemetry data can be used without reference to the personal identifiers. A reversible mapping is used to reverse the obfuscated personal identifier back to its original value. In the case where a request to delete the mapped personal identifier is received, the link to the entry in the reversible mapping is broken by associating the personal identifier with a different obfuscated value.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Telemetry data generated during the use of a software product, web page, or service ("resource") is often collected and stored to study the performance of the resource and / or the behavior of users utilizing the resource. Telemetry data provides a deep understanding of the use and performance of a resource under varying conditions, some of which may not have been tested or considered in the design of the resource. Telemetry data is useful for identifying causes of failures, delays, or performance issues and identifying ways to improve customer engagement with the resource.

[0002] To facilitate real-time data processing of large amounts of telemetry data, the telemetry data may be stored in different storage locations in various devices. The telemetry data may include personal information of users of the resource. Personal information may include a personal identifier that uniquely identifies the user, such as name, phone number, email address, social security number, login name, account name, machine identifier, etc.

[0003] However, removal of personal data found in telemetry data can be complicated by the way that telemetry data is stored. Data may be stored such that there may be no way to track down all storage locations containing all telemetry data with a personal identifier without searching all possible storage locations. This may not be possible or practical when large amounts of telemetry data are collected. Summary of the invention

[0004] The present invention content introduces some concepts in a simplified form, which will be further described in the following detailed description. The present invention content is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0005] The real-time event processing system replaces personal identifiers included with telemetry data with obfuscated values ​​so that the telemetry data can be analyzed without reference to the personal identifiers. When the identity of the user is needed, the reversible map is used to map the obfuscated personal identifiers to the original personal identifiers. The personal identifiers in the reversible map are removed in response to a request from the associated user by unlinking the key accessing the corresponding entry in the reversible map.

[0006] In one aspect, a cryptographic hash function is used to obfuscate a personal identifier into an obfuscated value. The cryptographic hash function uses a different seed in order to distinguish an obfuscated value representing a personal identifier that has not been subject to a deletion request from an obfuscated value representing a personal identifier that has been subject to a deletion request. The deletion request is implemented by associating a different obfuscated value with the personal identifier in the deletion table, thereby eliminating access to its corresponding entry in the reversible map. In this way, the personal identifier is deleted from the reversible map using a relatively inexpensive constant time operation.

[0007] These and other features and advantages will be apparent from the following detailed description and from reading and examining the associated drawings.It should be understood that the foregoing general description and the following detailed description are merely illustrative and are not restrictive of the aspects claimed. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 An exemplary real-time system for obfuscation and removal of personal information from telemetry data is illustrated.

[0009] FIG. 2A to FIG. 2E The states of a Reversible Map ("RMAP") table and a Delete table at different points in event processing are illustrated.

[0010] Figure 3 is a flow chart illustrating an exemplary method for obfuscating personal identifiers in telemetry data.

[0011] Figure 4A The insertion of an entry into an identifier ("ID") map is illustrated.

[0012] Figure 4B is a flow chart illustrating an exemplary method for inserting an entry into an ID map.

[0013] Figure 5 is a flow chart illustrating an exemplary method for processing a delete request.

[0014] Figure 6 is a flow chart illustrating an exemplary method for generating a new version of a deletion table.

[0015] Figure 7 is a block diagram illustrating an exemplary operating environment DETAILED DESCRIPTION

[0016] Overview

[0017] The disclosed technical solution is about a real-time event processing system that receives events containing telemetry data and one or more personal identifiers. The real-time event processing system obfuscates the personal identifiers contained with the telemetry data so that the telemetry data can be analyzed without reference to the personal identifiers. When the identity of the user is needed, a reversible mapping is used to map the obfuscated personal identifiers to the original personal identifiers.

[0018] The user of the personal identifier can request that their personal identifier be deleted from the real-time event processing system. In this case, the link to the personal identifier in the reversible mapping is broken, and the obfuscated personal identifier is retained to track that the user has previously issued a deletion request. When telemetry data is received from a user who has issued a previous deletion request, the personal identifier in the new event is obfuscated with a different obfuscated value so that the user identity cannot be derived from the previous obfuscated value.

[0019] The personal identifier is obfuscated using a cryptographic hash function that maps the personal identifier to a hash value. The cryptographic hash function uses a different seed or salt so as to distinguish hash values ​​representing the deleted personal identifier from hash values ​​representing the obfuscated personal identifier not associated with a deletion request. In one aspect, the cryptographic hash function generates a 32-byte hash value to represent the obfuscated personal identifier not associated with a deletion request and generates a 33-byte hash value to represent the obfuscated personal identifier associated with a deletion request.

[0020] In order for the real-time event processing system to handle large amounts of data, the system is configured as a loosely coupled distributed system with multiple independent domains, each configured to process events for a specific resource. The use of a loosely coupled distributed system facilitates the rapid processing of large amounts of data. However, in order to maintain the high throughput achieved by the loosely coupled distributed system, the obfuscation and deletion of personal identifiers needs to be performed in a cost-effective manner. This process is complicated by the fact that a single user can be associated with multiple resources, with different personal identifiers appearing in different domains. To overcome this complexity, the deletion of all personal identifiers associated with a user is facilitated by an identifier mapping and a common deletion table. The identifier mapping is used to track additional personal identifiers and the domains in which they are located, so as to completely remove access to the user's personal data in the system. The common deletion table is used by each domain to track deleted obfuscated personal identifiers that may reside in any domain. The common deletion table is periodically updated by an offline process that simultaneously synchronizes the newly updated deletion table in each domain.

[0021] Attention now turns to the description of a system for obfuscating and removing personal information in telemetry data processed by a loosely coupled distributed system.

[0022] system

[0023] Figure 1 An exemplary system for real-time processing of telemetry data is illustrated. In one aspect, system 100 includes a loosely coupled distributed system 102 that processes events 110 from one or more users 104 via a network 106. Loosely coupled distributed system 102 is composed of multiple domains 108A-108N ("108") that operate independently of each other without central control. Domain 108 can be configured with one or more computing devices. In one aspect, domain 108 can be configured to process events associated with specific resources. For example, a domain can be configured to process events caused by the use of a specific software product.

[0024] A user 104 intervenes with a resource through the use of a computing device. The resource uses an agent that is coupled to the resource (i.e., an add-on, extension, plug-in, operating system) or embedded within the resource. The agent generates events 110 during operation of the system or during interaction between the user and the resource. The agent generates events 110 in response to the occurrence of a user action, in response to a user action, or measures attributes of the user's behavior at certain time intervals. Events 110 contain telemetry data, which may include performance metrics, timing data, memory dumps, stack traces, exception information, and / or measurements in addition to personal information. There may be multiple events 110 created during user intervention with a resource, and events 110 may be generated for multiple and different resources with which the user intervenes.

[0025] For example, events 110 may be generated from actions performed by the operating system, based on user interactions with the operating system, or caused by user interactions with applications executed under the operating system. Events 110 may be system generated logs generated by resources, such as software products, to fix problems and improve the quality of the product. In this case, events 110 may include data from crashes, hangs, unresponsive user interfaces, high CPU usage, high memory usage, and the like. The data may include memory dumps, stack traces, and exception information.

[0026] Events 110 may include personal information. Personal information may include one or more personal identifiers that uniquely represent a user, and may include name, phone number, email address, social security number, login name, etc. An agent on the user's computing device transmits the event to a specific domain 108.

[0027] Each domain 108 may include an ingestion node 116A-116N ("116"), a reversible mapping ("RMAP") table 124A-124N ("124"), and a delete table 126A-126N ("126"). An ingestion node 116 is associated with a dedicated endpoint that is registered with a computing device of a user 104. The endpoint is a port or uniform resource locator (URL) that a computing device uses to connect to a particular domain. An ingestion node 116 may be implemented as a computing device or a component within a computing device.

[0028] Each ingestion node 116 receives events 110, which may include personal information or input values ​​114A-114N ("114") that need to be obfuscated. Ingestion nodes 116 include ingestion modules 118A-118N ("118") that search for personal identifiers in events. The personal identifiers in events 110 are then obfuscated so that the remaining information in the events can be used by further event processing 122A-122N ("122"). The obfuscation of the personal identifiers is tracked to avoid access to storage locations associated with the obfuscated personal identifiers when a deletion request is received.

[0029] Each ingestion node 116 may include a reversible mapping ("RMAP") table 124A-124N ("124") and a deletion table 126A-126N ("126"). The RMAP table 124 provides a reversible mapping of obfuscated representations of personal identifiers to original personal identifiers. In some cases, it may be necessary to contact a user to discuss an issue raised by telemetry data. The user identity can be obtained by reverse mapping the obfuscated values ​​to the original values ​​and using the original values ​​to discover the user identity.

[0030] The delete table 126 contains a list of deleted obfuscated personal identifiers for the entire system. An entry in the delete table 126 indicates that the personal identifier has been previously deleted from the system. The delete table 126 is accessed by an obfuscated representation of a first personal identifier that maps to an obfuscated representation of a second personal identifier that is stored in the delete table 126. Each domain uses the same delete table 126.

[0031] In one aspect, the personal identifier is obfuscated using a hash function that converts the personal identifier into a hash value. The hash function can be any common hash function or cryptographic hash function, such as the Message Digest 5 (MD5) algorithm, the Secure Hash Algorithm (SHA), etc. In one aspect, the SHA-256 function is used, with the original value of the personal identifier and a seed as its parameters. The seed or salt is random data (i.e., a randomized value) used in the hash function to protect the stored hash value.

[0032] The hash function converts the personal identifier into a first-generation hash value H (input value, seed 1 ). The first seed 1 is used in the calculation of the first generation hash, and the second seed 2 Used in the calculation of the second generation hash.

[0033] The output of the hash function is a fixed-length hash value. In order to distinguish between personal identifiers that have been deleted from the system and personal identifiers that remain in the system as obfuscated values, two different seeds are used. The first generation of hash values ​​uses the first seed. 1 A 32-byte hash value is generated to represent an obfuscated personal identifier that has not been deleted from the system. A subsequent hash value obtained using a hash function with a subsequent seed generates a 33-byte hash value to identify a personal identifier that has been deleted from the system.

[0034] like Figure 1 As shown, RMAP table 124 includes keys 128A-128N ("128") and original input values ​​130A-130N ("130"), where key 128 is a first generation hash value. Delete table 126 includes keys 132A-132N ("132") and values ​​134A-134N ("134"). Key 132 is a first generation hash value, and value 132 is a second or subsequent generation hash value. Subsequent generations of hash values ​​use a different seed than the first generation hash values.

[0035] The event processing system also receives delete requests from users 104 to delete their personal identifiers. In some cases, a user may have several personal identifiers. For example, a user may have an account name for a cloud service, a login name for a software product, and an email address for a mail service. Each personal identifier may have been obfuscated in a different domain and associated with a different obfuscated value. In this scenario, the different domains may not be aware of the various personal identifiers associated with the user. To ensure that all personal identifiers belonging to the user are deleted from the event processing system, delete requests 112 are collected and processed offline in the update domain, separate from the ingestion process, to synchronize the delete tables used in each domain.

[0036] In one aspect, the delete request process is performed offline outside of domain 108 and in update domain 146. The delete request process utilizes global storage 136, update module 138, and ID map 140 to generate updated delete table 148. Global storage 136 is accessible to each domain in domain 108 and can be implemented as a database, cloud storage, etc.

[0037] Once the event has been updated with the obfuscated input values ​​120A-120N ("120"), the event is stored in the global storage 136. The update module 138 generates an ID map 140 from the event stored in the global storage 136. The ID map 140 contains a hash value for each individual identifier in each domain 142, based on a key 144 common to the individual identifiers in each domain 142. The update module 138 creates an updated version of the deletion table 148 for each domain.

[0038] Attention now turns to further detailing the use of RAMP and the Delete table. FIG. 2A to FIG. 2C .

[0039] FIG. 2A to FIG. 2E Several different states of the RMAP and deletion tables are illustrated. Figure 2A 2 shows the state of the RMAP table 206 after the personal identifier in the event has been obfuscated and before a deletion request from a user associated with the personal identifier is received. Figure 2A As shown, an event 202 with an identifier ID 204 is received. The RMAP table 206 contains an obfuscated value 212 as a key 208, and an original identifier 214 as a value 210. The hash value 212 is hashed with the identifier 204 and the first generation seed. 1 Generated. There is no entry in the deletion table because a deletion request has not yet been received for this personal identifier.

[0040] Figure 2B 220 and the RMAP table 230 after a delete request has been processed and the delete table 220 has been updated. A second hash value 228 is used to represent the personal identifier stored in the delete table 220 along with its first generation hash value 226. The RMAP table 230 contains a key 232 and value 234 pair for each entry. Access to an entry in the RMAP table 230 is accomplished by merging the personal identifier with a different seed or salt. z The generated second hash value 228 is associated to be removed. Thus, there is no longer any access to the entry for the personal identifier 238 in the RMAP table 230 based on the first generation hash value 236.

[0041] like Figure 2B As shown, the delete request 216 includes an identifier ID 218. The delete table 220 contains a key 222 and a value 224 for each entry. The hash value corresponding to the identifier 218 is used as a key 226 in the delete table 220. The key 226 uses the identifier 218 and the first generation seed seed. 1 (expressed as H(ID,seed 1)) is generated. The value 228 represents the second hash value H(ID,seed) for the identifier Z ), this second hash uses a different seed z The link corresponding to the identifier ID in the table 230 is broken, since the hash value 236 based on the first generation seed cannot access the entry in the RMAP table.

[0042] Figure 2C The diagram shows the state of the RMAP table 254 after another event is received from a user who has previously issued a delete request. In this case, there is an existing entry in the delete table 244 for the identifier 242. However, the entry contains a second hash value 252 that is different from the previous hash value 228, so there is no way to access the previous entry in the RMAP table. The second hash value 252 is then used to generate another entry in the RMAP table for the identifier.

[0043] like Figure 2C As shown, there is an event 240 with an identifier 242. A first generation hash value 250 for the identifier is used to search a delete table 244. Using the first generation hash value 250 as a key 246, a corresponding entry is found in the delete table 244. A corresponding value 248 represents a second hash value 252 for the identifier. The second hash value 252 is used as a key 260 in an RMAP table 254. The RMAP table 254 includes a pair of a key 256 and a value 258 for each entry. The key 256 is the second hash value 260, and the value 262 is the identifier.

[0044] Figure 2D 274. x Used to generate the hash value 276 placed in the value column 272. In this way, there is no longer any access to the entry in the RMAP table 279 for the personal identifier 266.

[0045] like Figure 2DAs shown, the delete request 264 includes an identifier ID 266. The delete table 268 contains a key 270 and a value 272 for each entry. A hash value 274 corresponding to the identifier ID is used as the key 270 in the delete table 268. The key 274 uses the identifier ID 266 and the first generation seed seed. 1 (expressed as H(ID,seed 1 )) is generated. The value 272 represents the second hash value H(ID,seed) for the identifier x ), this second hash uses a different seed x The link to the entry in the RMAP table 279 corresponding to the identifier is broken, and the entry in the RMAP table 279 cannot be accessed due to the hash value.

[0046] Figure 2E The diagram shows the state of the RMAP table 290 after another event is received from a user who has previously issued two delete requests, both for the same personal identifier. In this case, there is an existing entry for identifier 284 in the delete table 286. However, this entry contains a second hash value 288 that is different from the previous hash value 276, so there is no way to access the previous entry in the RMAP table. The second hash value 289 is then used to generate another entry for the identifier in the RMAP table.

[0047] like Figure 2E As shown, event 284 with identifier 284 arrives at the ingest node. The first generation hash value 286 for the identifier is used to search the delete table 286. Using the first generation hash value 285 as the key 287, the corresponding entry is found in the delete table 286. The corresponding value 288 represents the second hash value 289 for the identifier. The second hash value 289 is used as the key 260 in the RMAP table 290. The RMAP table 290 includes a pair of key 292 and value 293 for each entry. The key 291 is the second hash value 289, and the value 293 is the identifier 284.

[0048] A number of different aspects of the system 100 may be implemented using hardware elements, software elements, or a combination of both. Examples of hardware elements may include devices, components, processors, microprocessors, circuits, circuit elements, integrated circuits, application-specific integrated circuits, programmable logic devices, digital signal processors, field programmable gate arrays, memory cells, logic gates, and the like. Examples of software elements may include software components, programs, applications, computer programs, applications, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, processes, software interfaces, application program interfaces, instruction sets, computing codes, code snippets, and any combination of the above examples. Determining whether an aspect is implemented using hardware elements and / or software elements may vary according to any number of factors, such as desired computing rates, power levels, bandwidth, computing time, load balancing, memory resources, data bus speeds, and other design or performance constraints desired for a given implementation.

[0049] It should be noted Figure 1 The components of the system are shown in one aspect of an environment in which various aspects of the invention can be practiced. Figure 1 The exact configuration of components shown may not be required to practice the various aspects and may not depart from the spirit and scope of the invention. Figure 1 Variations of the configurations and component types shown in may be made. For example, other obfuscation techniques such as randomization and the like may be used instead of the hash encryption techniques described above.

[0050] method

[0051] Attention now turns to the description of various exemplary methods utilizing the systems and devices disclosed herein. Operations for various aspects may also be described with reference to various exemplary methods. It will be appreciated that, unless otherwise noted, representative methods are not necessarily performed in the order presented or in any particular order. In addition, various activities described with respect to the methods can be performed in a serial or parallel manner, or in any combination of serial and parallel operations. In one or more aspects, the method illustrates operations for the systems and devices disclosed herein.

[0052] Figure 33 is a flow chart illustrating an exemplary method 300 performed by an ingestion module in each domain for obfuscating one or more personal identifiers received in an event. An ingestion node in an event processing system receives a large number of events with one or more personal identifiers that need to be obfuscated (block 302). In this case, when there are several personal identifiers that need to be obfuscated, one personal identifier is selected as a primary identifier. A first obfuscated value is calculated for the primary identifier using a cryptographic hash function (block 304) that utilizes a first generation seed and the primary identifier. The first generation hash value is then used as a key to search a deletion table (block 306). If the first generation hash value is found in the deletion table (block 308-yes), the primary identifier has been previously deleted. In this case, a second obfuscated value is obtained from the deletion table (block 314) and is used to replace the primary identifier in the event (block 316). The event is then stored in a global storage device and passed to other downstream components for additional operations (block 312).

[0053] If the first obfuscated value is not found in the deletion table (block 308 - No), the first generation hash value and the original personal identifier are stored in the RMAP table (block 310). The original value of the primary identifier in the event is replaced with the first generation hash value (block 310). The event is then stored in global storage and passed to other downstream components for additional operations (block 312). This process is repeated for each event received by each ingest node (blocks 302-316).

[0054] FIG. 4A to FIG. 4B The insertion of a personal identifier into an ID map is illustrated. Figure 4A is a block diagram illustrating components used to insert a personal identifier into an ID map, and Figure 4B is an exemplary flow chart illustrating the insertion process 420. FIG. 4A to FIG. 4B , events are received by the event processing system and stored in a global storage device 400 accessible to each domain. Event 402 contains hash values ​​404, 406 for each individual identifier in event 402. ID map 408 contains a row for each event and columns 410A-410N ("410") for each domain. Update module 412 scans each event stored in global storage device 400 (box 422). Events have a primary identifier and may have additional dependent identifiers. Update module 412 inserts a row for the event using the hash value of the primary identifier as a key (box 424). Update module 412 places each hash value 402, 406 in event 402 in a column corresponding to the domain 410 where the event was received.

[0055] Attention now turns to discussion of techniques for updating the delete table and synchronizing each domain with the updated delete table. Figure 5 In one aspect, an offline process is used to update a deletion table to reflect a deletion request. The deletion request contains at least one personal identifier that can be associated with other dependent identifiers. In one aspect, the deletion request can be queued and processed offline at specified time intervals.

[0056] The offline process may use a copy of the current version of the offline table. The offline process will update the deletion table (box 502), remove the entry in the RMAP table corresponding to the personal identifier associated with the deletion request (box 504), and synchronize the update of each domain with the updated deletion table (box 506).

[0057] Steering Figure 6 , there is an exemplary method 600 of a process performed by an update module to process a delete request. The update module will generate a copy of the delete table and update its copy when the domain utilizes the existing delete table (box 602). A hash value for the primary identifier in the delete request is generated using the current generation seed (box 604). This hash value is used to calculate the primary search identifier for the delete request (box 606). If the hash value is found in the delete table, the primary search identifier is set to a hash value based on the current generation seed. Otherwise, the primary search ID is set to a hash value based on the first generation seed. In one aspect, the hash value can have different sizes. For example, the hash value based on the first generation seed can be a size of 32 bytes, and the hash value based on the current seed can be a size of 33 bytes. The update module can search the delete table to find the stored hash value based on the size of the hash value, and determine the generation of the hash value from the size (generally, 606).

[0058] The primary lookup identifier is then used to search the ID map to find the dependent identifier (box 608). The primary lookup identifier is used as a key to the ID map. If an entry based on the primary lookup identifier exists in the ID map, the update module processes each entry in each domain column (box 608).

[0059] If the hash value in the domain column of the ID map is not found in the key column of the delete table, a new entry for the hash value is added to the delete table (box 608). The new entry in the delete table adds the hash value as a key, and the value is the hash value calculated using the current generation seed (box 608). If the hash value in the domain column of the ID map is found in the key column of the delete table, the hash value in the value column of the delete table is updated with the new hash value calculated using the next generation seed (box 608). This process is repeated for each entry in the value column of the ID map (box 608).

[0060] Next, the delete table is updated for the primary search identifier (box 610). If the primary search identifier is found in the ID map, a new entry is added in the delete table for the primary search identifier (box 610). The new entry in the delete table adds the primary search identifier in the key column and places a new hash value using the first generation seed (box 610). The entry that exists in the delete table and matches the primary search identifier is updated with the new hash value using the next generation seed and placed in the value column of the delete table. (box 610).

[0061] Return to Figure 5 After the deletion table is updated, the entry in the RMAP table corresponding to the primary identifier in the deletion request is deleted (block 504), and the deletion table in each domain is replaced with the new version. In one aspect, the event processing system is temporarily suspended so that the RMAP table in each domain is updated to delete the entry matching the primary identifier in each deletion request, and the deletion table in each domain is replaced with the updated version of the deletion table (block 506).

[0062] Exemplary Operating Environment

[0063] Attention now turns to the discussion of exemplary operating embodiments. Figure 7 An exemplary operating environment 700 including computing devices 702, 706, 746 is illustrated. Computing devices 702, 706, 746 can be any type of electronic device, such as, but not limited to, a mobile device, a personal digital assistant, a mobile computing device, a smart phone, a cellular phone, a handheld computer, a server, a server array or server farm, a web server, a network server, a blade server, an Internet server, a workstation, a minicomputer, a mainframe computer, a supercomputer, a network appliance, a web appliance, a distributed computing system, a multiprocessor system, or a combination thereof. The operating environment 700 can be configured in a network environment, a distributed environment, a multiprocessor environment, or a stand-alone computing device with access to a remote or local storage device.

[0064] The computing devices 702, 706, 746 may include one or more processors 708, 726, 748, communication interfaces 710, 728, 750, one or more storage devices 712, 730, 752, one or more input and output devices 714, 732, 754, and memories 716, 734, 756. The processors 708, 726, 748 may be commercially available or custom processors, and may include dual microprocessor and multi-processor architectures. The communication interfaces 710, 728, 750 facilitate wired or wireless communications between the computing devices 702, 706, 746 and other devices. The storage devices 712, 730, 752 may be computer-readable media that do not contain propagation signals (such as modulated data signals transmitted by carrier waves). Examples of storage devices 712, 730, 752 include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD), or other optical storage, cassettes, magnetic tape, disk storage, all of which do not contain propagating signals (such as modulated data signals transmitted via carrier waves). There may be multiple storage devices 712, 730, 752 in computing devices 702, 706, 746. Input / output devices 714, 732, 754 may include a keyboard, a mouse, a pen, a voice input device, a touch input device, a display, a microphone, a printer, etc., and any combination of the above.

[0065] The memory 716, 734, 756 may be any non-transitory computer-readable storage medium that can store executable processes, applications, and data. The computer-readable storage medium is not suitable for propagating signals (such as modulated data signals transmitted by carrier waves). It may be any type of non-transitory memory device (e.g., random access memory, read-only memory, etc.), magnetic storage device, volatile storage device, non-volatile storage device, optical storage device, DVD, CD, floppy disk drive, etc. that is not suitable for propagating signals (such as modulated data signals transmitted by carrier waves). The memory 716, 734, 756 may also include one or more external storage devices or remotely located storage devices that are not suitable for propagating signals (such as modulated data signals transmitted by carrier waves).

[0066] In one aspect, computing device 702 is used by a user, computing device 706 is used by an ingestion node, and computing device 746 is used by an offline process in an update domain. Memories 716, 734, 756 may contain instructions, components, and data. A component is a software program that performs a specific function and is also referred to as a module, program, engine, and / or application. The memory 716 of computing device 702 may contain an operating system 718, one or more resources 720 (such as one or more software products), an event generation agent 722 that is communicatively coupled to a resource that generates an event 723, and other applications and data 724. The memory 734 of computing device 706 may include an operating system 736, an ingestion module 738, an RMAP table 740, a deletion table 742, an event 743, and other applications and data 744. The memory 756 of computing device 746 may contain an operating system 758, an update module 760, a global storage device 762, an ID map 764, an updated deletion table 766, and other applications and data 768.

[0067] in conclusion

[0068] Various aspects of the technical solutions disclosed herein are directed to the technical problem of obfuscating and deleting personal identifiers contained in a large number of events, where the events contain telemetry data and are processed in a loosely coupled distributed system. Technical features associated with solving this problem use a technique of obfuscated values ​​that replace personal identifiers so that telemetry data can be used without referencing the personal identifiers. The obfuscated values ​​are constructed so that they can distinguish between personal identifiers that have been subject to a deletion request and personal identifiers that have not been subject to a deletion request.

[0069] A system is disclosed, which includes at least one domain, wherein the at least one domain includes one or more processors and memory. The at least one domain performs actions, the actions of: maintaining a reversible mapping of obfuscated values ​​to at least one personal identifier; maintaining a deletion table, the deletion table including a plurality of entries for deleted personal identifiers, the entries including a first obfuscated value and a second obfuscated value, the first obfuscated value being associated with a first generation of randomized values, and the second obfuscated value being associated with a current generation of randomized values; obtaining at least one event including the at least one personal identifier; replacing the at least one personal identifier in the at least one event with the first obfuscated value when the at least one personal identifier is not included in the deletion table; replacing the at least one personal identifier in the at least one event with the second obfuscated value when the at least one personal identifier is included in the deletion table; and utilizing event data without referencing the at least one personal identifier.

[0070] When the at least one personal identifier is not found in the deletion table, access to the reversible mapping is based on the first obfuscated value. When the at least one personal identifier is found in the deletion table, access to the reversible mapping is based on the second obfuscated value. The at least one domain performs further actions, which are: in response to receiving a deletion request for the at least one personal identifier, searching the deletion table for the first obfuscated value associated with the at least one personal identifier; and when the first obfuscated value is present in the deletion table, unlinking access to the reversible mapping by changing the second obfuscated value associated with the at least one personal identifier in the deletion table. The at least one domain performs further actions, which are: when the first obfuscated value is not present in the deletion table, generating an entry in the deletion table for the at least one personal identifier, the entry including the first obfuscated value and the second obfuscated value, the second obfuscated value being associated with the next generation of randomized values.

[0071] The system includes an update domain having at least one processor and a memory. The update domain has at least one module configured to perform actions, the actions of: generating an identifier table to store dependent personal identifiers associated with the at least one personal identifier in the at least one event; and constructing an updated deletion table to reflect personal identifiers received in events from any domain that have been subject to deletion requests.

[0072] At least one module in the update domain performs further actions that: create an entry in an updated deletion table for the personal identifier that is subject to the deletion request that is not present in the deletion table; and update an obfuscated value associated with the personal identifier that is already present in the deletion table. Further actions include removing the entry for the personal identifier that is subject to the deletion request in the reversible mapping, and updating the current generation randomized value. The magnitude of the first obfuscated value distinguishes the first obfuscated value from the second obfuscated value.

[0073] A method is disclosed, comprising the following actions: obtaining a first delete request to remove at least one personal identifier at an ingestion node of a computing system, the ingestion node comprising at least one processor and a memory, the ingestion node storing the at least one personal identifier in a first table, the at least one personal identifier in the first table being accessed by a current fuzzy value; searching for the at least one personal identifier in a delete table using the first fuzzy value, the first fuzzy value being based on the at least one personal identifier, the delete table being accessed by the first fuzzy value, the delete table storing a second fuzzy value, the first fuzzy value being different from the second fuzzy value; when the first fuzzy value is not found in the delete table, inserting an entry for the at least one personal identifier in the delete table; when the at least one personal identifier is found in the delete table, unlinking access to the at least one personal identifier in the first table by replacing the second fuzzy value in the delete table with a fuzzy value different from the current fuzzy value; and utilizing telemetry data associated with the at least one personal identifier without referencing the at least one personal identifier.

[0074] Searching the deletion table for the at least one personal identifier also includes generating a first obfuscated value using the at least one personal identifier and the first randomized value. The method performs the further actions of: obtaining an event including the at least one personal identifier; and storing the at least one personal identifier in the first table using a second obfuscated value from the deletion table when the at least one personal identifier has a corresponding entry in the deletion table. The method also includes: obtaining a second deletion request for the at least one personal identifier; and unlinking access to the at least one personal identifier in the first table by replacing the second obfuscated value in the deletion table with an obfuscated value that is different from the second obfuscated value.

[0075] The method also includes: obtaining an event associated with the first personal identifier; determining that the first personal identifier is not associated with an entry in the deletion table; storing an obfuscated value representing the first personal identifier in the first table; replacing the first personal identifier with the obfuscated value in the event; and utilizing data in the event without referencing the first personal identifier.

[0076] A device having at least one processor and a memory is disclosed. The at least one processor is configured to: generate a mapping that maps a current fuzzy value to a personal identifier; generate a second mapping that maps a first fuzzy value to a current fuzzy value; obtain first event data including the personal identifier; when the first fuzzy value is not mapped to the current fuzzy value, replace the personal identifier in the event data with the first fuzzy value; when the first fuzzy value is mapped to the current fuzzy value, replace the personal identifier in the event data with the current fuzzy value; retain the first event data without the personal identifier.

[0077] The first obfuscated value is based on the personal identifier and the first randomized value, and the second obfuscated value is based on the personal identifier and the second randomized value, the first randomized value being different from the second randomized value. The device is also configured to: unlink access to the first mapping to the personal identifier by changing the second obfuscated value; and generate a mapping from a subsequent generation of the obfuscated value to the personal identifier in response to receiving second event data containing the personal identifier. The current obfuscated value and the first obfuscated value were generated using different randomized values. The event data includes telemetry data.

[0078] Although the technical solution has been described with specific structural features and / or methodological actions, it should be understood that the technical solution defined in the claims is not necessarily limited to the specific features or actions described above. The specific features and actions described above are more disclosed as example forms of implementing the claims.

Claims

1. A system for processing events, comprising: at least one domain, wherein the at least one domain includes one or more processors and a memory; wherein the at least one domain includes at least one module that performs actions, the actions: obtain an event including a personal identifier; generate a first obfuscated value to replace the personal identifier; search for the first obfuscated value in a deletion table, wherein the deletion table includes a plurality of entries, the entries including the first obfuscated value and a second obfuscated value, the first obfuscated value being associated with a first-generation randomized value, the second obfuscated value being based on a current-generation randomized value; when the first obfuscated value is not found in the deletion table, replace the personal identifier in the event with the first obfuscated value and store it in a reversible mapping table; when the first obfuscated value is found in the deletion table, obtain the second obfuscated value of the personal identifier from the deletion table and replace the personal identifier in the event with the second obfuscated value; and utilize the event without referring to the personal identifier.

2. The system according to claim 1, wherein the at least one module performs actions, the actions: receive a deletion request for the personal identifier; generate the first obfuscated value of the personal identifier; search for the first obfuscated value in the deletion table ; and when the first obfuscated value exists in the deletion table: generate a new obfuscated value; and store the new obfuscated value as the second obfuscated value in the entry in the deletion table associated with the personal identifier, wherein the new obfuscated value breaks the link to the reversible mapping table for access to the personal identifier.

3. The system according to claim 2, wherein the at least one module performs further actions, the further actions: when the first obfuscated value does not exist in the deletion table, generate an entry for the personal identifier in the deletion table, the entry including the first obfuscated value.

4. The system according to claim 1, further comprising: an update domain including at least one processor and a memory; at least one module in the update domain, the at least one module being configured to perform actions, the actions: generate an identifier table to store a dependent personal identifier associated with the personal identifier in the event; and construct an updated deletion table to reflect personal identifiers that have received deletion requests in events from any domain.

5. The system according to claim 4, wherein the at least one module in the update domain performs further actions, the further actions: create an entry in the updated deletion table for a personal identifier that has received a deletion request and does not exist in the deletion table; and update the first obfuscated value associated with a personal identifier that exists in the deletion table.

6. The system according to claim 5, wherein at least one module in the update domain performs a further action, the further action removing an entry for the personal identifier subject to the deletion request in the reversible mapping.

7. The system according to claim 6, wherein at least one module in the update domain performs a further action, the further action updating the value of the current generation randomization.

8. The system according to claim 2, wherein the magnitude of the first obfuscated value distinguishes the first obfuscated value from the second obfuscated value.

9. A method for processing personal identifiers, comprising: at an ingestion node of a computing system, obtaining a first deletion request to remove a personal identifier, the ingestion node including at least one processor and a memory, the personal identifier being stored in a reversible mapping table of the ingestion node, the personal identifier being accessed in the reversible mapping table by a first obfuscated value; generating the first obfuscated value to replace the personal identifier; searching for the personal identifier in a deletion table using the first obfuscated value, wherein the deletion table includes a plurality of entries, the entries including the first obfuscated value and a second obfuscated value, the first obfuscated value being associated with a value of a first generation randomization, the second obfuscated value being based on a value of a current generation randomization; when the personal identifier is not found in the deletion table, inserting an entry for the personal identifier in the deletion table, the entry containing the first obfuscated value; when the personal identifier is found in the deletion table, unlinking access to the personal identifier in the reversible mapping table by replacing the second obfuscated value in the deletion table with a new obfuscated value; and utilizing telemetry data associated with the personal identifier without referencing the personal identifier.

10. The method according to claim 9, wherein searching for the at least one personal identifier in the deletion table further includes generating the first obfuscated value using the personal identifier and the value of the first generation randomization.

11. The method according to claim 9, further comprising: obtaining an event including the personal identifier; searching for the personal identifier in the deletion table; when the personal identifier has a corresponding entry in the deletion table, storing the personal identifier in the reversible mapping table using the second obfuscated value from the deletion table.

12. The method according to claim 11, further comprising: obtaining a second deletion request for the personal identifier; and unlinking access to the personal identifier in the reversible mapping table by replacing the second obfuscated value in the deletion table with a new obfuscated value.

13. The method according to claim 9, further comprising: obtaining an event associated with a first personal identifier; determining that the first personal identifier is not associated with an entry in the deletion table; generating a first obfuscated value for the first personal identifier; Store the first blurred value for the first personal identifier in the reversible mapping table; Replace the first personal identifier with the first blurred value for the first personal identifier in the event; And Utilize the data in the event without referring to the first personal identifier.

14. A device for processing event data, Comprising: At least one processor and a memory; Wherein the at least one processor is configured to: Generate a deletion table with multiple entries that map a first blurred value to a second blurred value; Generate a reversible mapping table with multiple entries that map the first blurred value to a personal identifier; Obtain event data including the personal identifier; Generate the first blurred value to replace the personal identifier; When the first blurred value is not found in the deletion table, replace the personal identifier in the event data with the first blurred value; When the first blurred value is found in the deletion table, replace the personal identifier in the event data with the second blurred value; Retain the event data without the personal identifier.

15. The device according to claim 14, wherein the first blurred value is based on the personal identifier and a first randomized value, and the second blurred value is based on the personal identifier and a second randomized value, and the first randomized value is different from the second randomized value.

16. The device according to claim 14, wherein the at least one processor is further configured to: Receive a first deletion request for the personal identifier; By changing the second blurred value in the deletion table, unlink the access to the entry of the reversible mapping table for the personal identifier.

17. The device according to claim 16, wherein the at least one processor is further configured to: Receive a second event including the personal identifier; and Generate an entry in the reversible mapping table for the personal identifier using the second blurred value for the personal identifier from the deletion table.

18. The device according to claim 17, wherein the at least one processor is further configured to: Receive a second deletion request for the personal identifier; Generate a new blurred value for the personal identifier; and By storing the new blurred value as the second blurred value in the deletion table, unlink the access to the entry of the reversible mapping table for the personal identifier.

19. The device according to claim 14, wherein the event data includes telemetry data.

Citation Information

Patent Citations

  • Methods and systems for deleting requested information

    CN105940412A

  • System and method for obfuscating an identifier to protect the identifier from impermissible appropriation

    EP3040898A1