System and method for normalizing personally identifiable information and identity verification

The system addresses the challenge of standardizing and securely processing diverse PII formats by converting them into encoded objects for privacy-preserving identity verification and fraud prevention, improving accuracy and security in data exchange.

WO2025146571A1PCT designated stage expired Publication Date: 2025-07-10WINR DATA PTY LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2024/060423
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-03
Filing Date
2024-10-23
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Current systems face challenges in standardizing and securely processing personally identifiable information (PII) such as postal addresses, names, and dates of birth due to diverse formatting, leading to inefficiencies in identity verification and fraud prevention, particularly in the absence of standardized entry formats and secure data transfer.

Method used

A system and method that converts PII into encoded objects using encryption techniques, enabling privacy-preserving searching and linking across networks, while ensuring secure data exchange and compliance with privacy regulations, utilizing country-specific parsers, phonetic encoders, and hash encoders to normalize and encrypt data.

Benefits of technology

Enhances the accuracy of identity verification and fraud prevention by securely processing and linking diverse PII formats, reducing the risk of data breaches, and ensuring compliance with privacy regulations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024060423_10072025_PF_FP_ABST
    Figure IB2024060423_10072025_PF_FP_ABST
Patent Text Reader

Abstract

A system and method for normalizing personal identifiable information (PII) and for identity verification in a computing environment is disclosed. The systems and methods describe receiving, at a server, a data request from at least one third party in the computing environment, wherein the data request comprises a data set of clear text PII of a customer, normalizing the clear text PII, wherein normalizing the clear text PII comprises parsing the clear text PII into a series encoded objects, transferring the encoded objects to an internal search engine, at the internal search engine, querying an internal database of normalized data objects for matches to the encoded objects, receiving, from the internal search engine, a first result set based on the encoded objects that are matched within the internal database, and forming an amalgamated response, wherein the amalgamated response comprises an identity verification confirmation, a confidence score, or both.
Need to check novelty before this filing date? Find Prior Art

Description

TITLESystem and Method for Normalizing Personally Identifiable Information and Identity VerificationTECHNICAL FIELD

[0001] The present disclosure generally relates to data normalization Personally Identifiable information (PH) and identify verification in a computing environment or networked environment. More specifically, the present disclosure relates to a system and method for efficiently processing PH including but not limited to postal addresses, names, and dates of birth (DOB) from diverse geographic regions, and creating encoded objects such as keys that facilitate privacy-preserving searching, sorting, and linking of records within databases and across networks for identity verification, resolution and fraud mitigation.BACKGROUND

[0002] Anonymity online can create a liability for businesses and consumers in everyday online transactions. In fact, global losses resulting from online fraud is expected to reach US$55B by 2024. This is in large part due to identify theft based on the lack of safe care of (PH), which itself falls into a large subset of online fraud generally. As a result, governments worldwide have instituted regulations for identity verification (IDV) one of which is called KYC or Know Your Customer (KYC) to protect consumers and businesses from the risks of fraud and other criminal activities. However, KYC verification has a myriad of drawbacks such as high amounts of manual labor, false positives, and incongruent data schema, to name a few. Consequently, in the realm of identity verification and fraud prevention, there is a growing demand for robust systems capable of linking disparate address points and personal identifiers such as names and dates of birth (DOB) associated with specific individuals.

[0003] While certain solutions have been proposed to enhance consumer identification procedures, there are still multiple pitfalls. For example, the format of personally identifiable information (PH) has proven extremely problematic for companies and their providers. For example, postal addresses, names, and DOBs, while mostly essential toverification, are inherently sensitive, and consumers often craft them according to personal conventions. Variations in formatting, including diverse abbreviations for elements such as apartments and street types, have historically posed challenges to address handling systems. In other words, there is a lack of standardized entry format, and the absence of standardized address data introduces variability and complexity into real-time parsing algorithms, making it challenging to accurately dissect and categorize address components. Furthermore, stringent latency requirements for address parsing and database searching are compounded by hardware limitations, such as processor type and speed, as well as the computational expense of the parsing algorithm itself.

[0004] In some systems, users may enter their address in many forms (street name, apartment name, apartment number, city, etc.). However, sometimes users input incorrect information. For example, sometimes a user may misspell a street name or include an incorrect numeral as part of a street address. Other times, a user will input correct information, but that information will be non-standard and thus difficult to properly map or locate for sending to other systems (e.g. to a delivery worker's device). For example, a user may input a postal code but may only include the first digits of the postal code (and not the remaining digits of a full postal code number), or a user may enter an abbreviation for a street name that is not a standard abbreviation. As a result, this address information may be difficult to map or locate for sending to other systems.

[0005] Current address correction systems may provide for correction of misspellings and may attempt to standardize non-standard address input. Current address correction systems may also provide a user interface for entry and storage of multiple addresses and may allow for display of a corrected address for the user. For example, a user may enter a shipping address for an online purchase of a product, and current corrected address systems may fix or correct the shipping address for display to the user before the user approves and subsequently completes purchase of the online product. In this instance, a user may be able to review and approve the corrected address before proceeding to the next step of the online purchase.

[0006] However, current addresses correction systems are limited in the user interfaces that may be displayed to a user and are limited in the type of errors in input address information that they may be able to detect and correct. Furthermore, current address correction systems are unable to analyze address information over timeto determine patterns for automatic correction of new address input. Moreover, current systems do not take into account the global nature of businesses and the myriads of ways addresses (and any type of PI I) may be input by a user based on those jurisdictions formalities.

[0007] Furthermore, current systems are also reliant on the presence of an authoritative database comprising all delivery points typically issued under license from the postal authority and government mapping agencies. In emerging markets these databases may not be readily available.

[0008] In parallel, the secure transfer of Pll between organizations to verify the identity of individuals is a massive enterprise and is critical to prevent fraud and identify theft.Currently, organizations that need to secure inbound Pll use data-centric encryption and decryption as it is shared between organizations or with a single organization. However, when sharing between two organizations or more, the data becomes less secure and susceptible to intercept or targeted hacking and phishing scams and assumes equal (or at the very least, a certain level of trust between the sender and receiver).

[0009] In light of all the above-mentioned drawbacks, there is a need for a system and method that can normalize, efficiently process and securely transfer data such as Pll, and verify identification.SUMMARY

[0010] The present disclosure describes systems and methods for efficiently processing Pll including but not limited to postal addresses, names, and dates of birth (DOB) from diverse geographic regions, and creating encoded objects such as keys that facilitateprivacy-preserving searching, sorting, and linking of records within databases and across networks or in a computing environment for identity verification, resolution and fraud mitigation and may be used as a supplement or in place of KYC verification

[0011] In operation, the systems and methods described herein provide the practical application of allowing for two or more parties to employ encryption techniques without exposing the information to each other, thus removing the need to expose clear text data, such as PI I . Further, this limits the possibility of data breach due to the parties not having a cache of re-identifiable PI I.

[0012] The disclosure provides the practical application of providing a universal address key generation system and method for data security and privacy compliance while enabling efficient data management, particularly in the context of identity verification and fraud prevention. Furthermore, it enables parties to safely exchange addresses, names, and date of birth (DOB) information for data matching without exposing the actual PH, thus safeguarding individual privacy and increasing data security.

[0013] The disclosure describes a system and method that takes postal addresses, names, and DOBs as inputs from customer's various geographic regions and converts them into encoded objects, which in embodiments, are privacy-preserving keys. These keys support identity verification and fraud prevention because, in one embodiment, they enable the secure aggregation of related address points, names, and DOBs and other PI I . This not only enhances the accuracy of identity verification but also significantly bolsters fraud prevention capabilities. Moreover, it facilitates secure data exchange for data matching to third party databases or databases that comprise standardized information.

[0014] Furthermore, because the keys contain the normalized data of customers, it speeds up the verification process over the networks because there is no need to hand verify the data set due to the normalization of the customer addresses, in particular.

[0015] In embodiments, the systems and methods perform data conversion on the identify verification (IDV) application program interface API server on the system side. It exclusively transfers the encrypted keys, minimizing the exposure of Pll during the identity verification process. This approach ensures that sensitive information remains confidential while maintaining data security and privacy compliance. Furthermore, in embodiments, the system and method allow collaborating organizations to enhancesecurity and reduce data re-identification by incorporating an additional "salt" into the encryption algorithm, such as password hashing or string lengthening.

[0016] In embodiments, the systems and methods securely link various address points, names, and DOBs (and other PH) related to streets or buildings, even in the presence of diverse formats. Moreover, it enables parties to safely exchange linkage identifiers without revealing sensitive personal information, in some embodiments, using artificial intelligence.

[0017] In embodiments, a method for normalizing personal identifiable information (PI I) and for identity verification in a computing environment is disclosed. The method comprises receiving, at a server, a data request from at least one third party in the computing environment, wherein the data request comprises a data set of clear text Pll of a customer, normalizing the clear text Pll, wherein normalizing the clear text Pll comprises parsing the clear text Pll into a series encoded objects, transferring the encoded objects to an internal search engine, at the internal search engine, querying an internal database of normalized data objects for matches to the encoded objects, receiving, from the internal search engine, a first result set based on the encoded objects that are matched within the internal database, and forming an amalgamated response, wherein the amalgamated response comprises an identity verification confirmation, a confidence score, or both.

[0018] In embodiments a system for normalizing personal identifiable information (Pll) and for identity verification in a computing environment is disclosed. The system comprises one or more memories storing instructions, one or more processors executing the instructions to receive, at a server, a data request from at least one third party in the computing environment, wherein the data request comprises a data set of clear text Pll of a customer, normalize the clear text Pll, wherein normalizing the clear text Pll comprises parsing the clear text Pll into a series encoded objects, transfer the encoded objects to an internal search engine, at the internal search engine, query an internal database of normalized data objects for matches to the encoded objects, receive, from the internal search engine, a first result set based on the encoded objects that are matched within the internal database, form an amalgamated response, wherein the amalgamated response comprises an identity verification confirmation, a confidence score, or both.

[0019] In embodiments, a non-tangible, computer-readable recording medium storing instructions is provided, that, when executed by a processor, cause the processor to receive, at a server, a data request from at least one third party in the computing environment, wherein the data request comprises a data set of clear text PH of a customer, normalize the clear text PH, wherein normalizing the clear text PH comprises parsing the clear text Pll into a series encoded objects, transfer the encoded objects to an internal search engine, at the internal search engine, query an internal database of normalized data objects for matches to the encoded objects, receive, from the internal search engine, a first result set based on the encoded objects that are matched within the internal database and form an amalgamated response, wherein the amalgamated response comprises an identity verification confirmation, a confidence score, or both.

[0020] In embodiments, the system and methods may reside on a distributed network architecture to enhance scalability, fault tolerance, and data redundancy. This design enables seamless data collaboration across various entities while maintaining robust security.

[0021] In embodiments, the systems and methods incorporate data normalization and matching capabilities, streamlining the process of reconciling disparate data formats, resulting in a comprehensive and cohesive dataset for improved record management. Furthermore, the systems and methods enable secure connections to data sources, both internally and externally, fostering a safe environment for data exchange. This secure data sourcing ensures compliance with privacy regulations and protects individual identities, all while preventing exposure of personal data.

[0022] The systems and methods described herein may be used for identity verification and can be integrated into identity verification systems used in various sectors, including financial services, healthcare, and online authentication, providing a robust solution for complex addresses, names, and DOB, all while ensuring minimal Pll transfer and enhanced security through salt incorporation. The systems and methods described herein may be used in fraud detection to enhance capabilities in industries vulnerable to identity-related fraud, such as banking, insurance, and e-commerce, accommodating diverse data formats while safeguarding privacy and data security. The systems and methods described herein may be used in public safety and law enforcement agencies, which can use the system to connect addresses, names, and DOBs associated withcriminal activities, aiding in investigations, and securely exchanging information for data matching.

[0023] In embodiments, the disclosure describes a computing system comprising a processor, a memory and cryptographic circuit coupled to the memory. The cryptographic circuit is configured to generate a data pipeline between partner networks who need to verify a person's identity, and platform clients that verify identities using, in embodiments, KYC standards. The platform clients have a database including PH and processes and extract customer or client data need to verify identity for the partner network. The application parses the PH and loads only the encrypted keys into a database which is made available to the platform clients (verifying organization) API.

[0024] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter. The claimed subject matter is not limited to implementations that solve any or all disadvantages noted in the Background.BRIEF DESCRIPTION OF FIGURES

[0025] Aspects of the present disclosure are illustrated by way of example and are not limited by the accompanying figures for which like references indicate like elements.

[0026] FIG. 1 illustrates a system architecture block diagram for processing Pll and securely linking records within databases across networks for identity verification and fraud mitigation in accordance with one embodiment.

[0027] FIG. 2 illustrates a block of memory for parsing and normalizing Pll using artificial intelligence (Al) in accordance with one embodiment.

[0028] FIG. 3 illustrates a system architecture block diagram for processing and normalizing Pll within databases across networks for identity verification and fraud prevention in accordance with one embodiment.

[0029] FIG. 4 illustrates a combination architecture and flowchart executed by software code for processing an Identity Verification (IDV) in accordance with one embodiment.

[0030] FIG. 5 illustrates a system architecture block diagram for processing Pll and securely linking records within databases across networks for identity verification and fraud prevention in accordance with one embodiment.

[0031] FIG. 6 illustrates a stepwise diagram for processing Pll and securely importing and verifying records within databases across networks for identity verification and fraud prevention in accordance with one embodiment.

[0032] FIG. 7 illustrates a flowchart for the secure, privacy-compliant data-sharing process within a Clean Room environment, highlighting the use of secure keys and encoding to protect sensitive information shown generally at 700 in accordance with one embodiment.

[0033] FIG. 8 illustrates a stepwise diagram for encoding Pll and securely processing customer transactions in accordance with one embodiment.

[0034] FIG. 9 illustrates another stepwise diagram for and encoding Pll and securely processing customer transactions in accordance with one embodiment in accordance with one embodiment.

[0035] FIG. 10 illustrates a stepwise diagram for fraud detection in accordance with one embodiment in accordance with one embodiment.

[0036] FIG. 11 illustrates process map for the parser of a postal address shown in accordance with one embodiment; and

[0037] FIG. 12 illustrates a stepwise diagram for processing Pll and securely linking records within databases for identity verification in accordance with one embodiment.DETAILED DESCRIPTION

[0038] Exemplary embodiments are discussed below with reference to the Figures.

[0039] In the descriptions above and in the claims, phrases such as "at least one of or "one or more of' may occur followed by a conjunctive list of elements or features. The term "and / or" may also occur in a list of two or more elements or features. Unless otherwise implicitly or explicitly contradicted by the context in which it used, such a phrase is intended to mean any of the listed elements or features individually or any of the recited elements or features in combination with any of the other recited elements or features. For example, the phrases "at least one of A and B;" "one or more of A and B;" and "A and / or B" are each intended to mean "A alone, B alone, or A and B together." A similar interpretation is also intended for lists including three or more items. For example, the phrases "at least one of A, B, and C;" "one or more of A, B, and C;" and "A, B, and / or C" are each intended to mean "A alone, B alone, C alone, A and B together, A and C together, B and C together, or A and B and C together." Use of the term "based on," above and in the claims is intended to mean, "based at least in part on," such that an unrecited feature or element is also permissible.

[0040] The subject matter described herein can be embodied in systems, apparatus, methods, and / or articles depending on the desired configuration. The implementations set forth in the foregoing description do not represent all implementations consistent with the subject matter described herein. Instead, they are merely some examples consistent with aspects related to the described subject matter. Although a few variations have been described in detail above, other modifications or additions are possible. In particular, further features and / or variations can be provided in addition to those set forth herein. For example, the implementations described above can be directed to various combinations and sub combinations of the disclosed features and / or combinations and sub combinations of several further features disclosed above. In addition, the logic flows depicted in the accompanying figures and / or described herein do not necessarily require the particular order shown, or sequential order, to achieve desirable results. Other implementations may be within the scope of the following claims.

[0041] Those skilled in the art will recognize that this example is illustrative and not limiting and is provided purely for explanatory purposes. An example of a computing system environment is disclosed. The computing system environment is not intended to suggest any limitation as to the scope of use or functionality of the system and method described herein. Neither should the computing environment be interpreted as having any dependency or requirement relating to any one or combination of components illustrated in the exemplary operating environment.

[0042] Embodiments of the disclosure are operational with numerous other general purposes or special purpose computing system environments or configurations. The embodiments of the disclosure may be described in the general context of computerexecutable instructions, such as program modules, being executed by a computer or smart device. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular data types. The systems and methods described herein may also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through or overlayed by a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media including memory unit or storage devices. Tasks performed by the programs and modules are described below and with the aid of figures. Those skilled in the art can implement the exemplary embodiments as processor executable instructions, which can be written on any form of a computer readable media in a corresponding computing environment according to this disclosure.

[0043] Computers and smart devices may comprise a variety of computer readable media. Computer readable media can be any available media that can be accessed by computer and comprises both volatile and non-volatile media, removable and non-removable media. By way of example, and not limitation, computer readable media may include computer storage media and communication media. Computer storage media comprises both volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data.

[0044] The non-transitory computer-readable includes all types of computer-readable media or mediums, including magnetic storage media, optical storage media, and solid-state storage media and specifically excludes signals. It should be understood that the software can be installed in and sold with the device. Alternatively, the software can be obtained and loaded into the device, including obtaining the software via a disc medium or from any manner of network or distribution system, including, for example, from a server owned by the software creator or from a server not owned but used by the software creator. The software can be stored on a server for distribution over the Internet, for example.

[0045] Computer-readable storage media (medium) exclude (excludes) propagated signals per se, can be accessed by a computer and / or processor(s), and include volatile and nonvolatile internal and / or external media that is removable and / or non-removable. For the computer, the various types of storage media accommodate the storage of data in any suitable digital format. It should be appreciated by those skilled in the art that other types of computer readable medium can be employed such as zip drives, solid state drives, magnetic tape, flash memory cards, flash drives, cartridges, and the like, for storing computer executable instructions for performing the novel methods (acts) of the disclosed architecture.

[0046] As used herein, Pll means any representation of information that permits the identity of an individual to whom the information applies to be reasonably inferred by either direct or indirect means. Further, Pll is defined as information: (i) that directly identifies an individual (e.g., name, address, social security number or other identifying number or code, telephone number, email address, etc.) or (ii) by which an agency intends to identify specific individuals in conjunction with other data elements, i.e., indirect identification. (These data elements may include a combination of gender, race, birth date, geographic indicator, and other descriptors). Additionally, information permitting the physical or online contacting of a specific individual is the same as personally identifiable information. This information can be maintained in either paper, electronic or other media.

[0047] As used herein, the terms "elements" and "components" may be used interchangeably.

[0048] As used herein, "customer" may mean any type of consumer or potential user, and all may be used interchangeably.

[0049] The described embodiments are to be considered in all respects only as illustrative and not restrictive. The scope is, therefore, indicated by the appended claims rather than by the foregoing description. All changes that come within the meaning and range of equivalency of the claims are to be embraced within their scope. Referring now to FIG. 1, which illustrates a system architecture block diagram for processing, normalizing, and encoding Personal Identifiable Information ( PI I) data 104 and linking records within a secure, distributed network for identity verification and fraud prevention is shown generally at 100.

[0050] The system 100 may be implemented on a distributed network but may also be implemented on other architectures as well.

[0051] This system is a secure, in operation, is a non-PII intermediary 110 between client 106 Identity Verification (IDV) requests and a distributed data partner network 134. Identity Verification can be made using any of the Pll data; including but not limited to, Know-Your-Customer (KYC) queries that are typically Name and DOB, Name and Address, or Name, DOB and Address. In operation, customers or consumers using smart phones or other devices 102 utilizing the client platforms 106a-c enter their clear-text Pll 104 on their user devices 102a-120c (n+1). In exemplary embodiments, the data may comprise addresses, date of birth, email addresses, phone numbers, and names, but other categories of Pll may be ingested from the user based on the needs of the client platform network or verifying organization 106a-106c (n+1). Examples of verifying organizations include but are not limited to obliged or regulated industries that includes Financial Services, Insurance, Real Estate, Gambling, Digital Payments, Legal & Accounting firms involved in financial transactions, Precious Metals & Gems. Other examples include e-commerce marketplaces and ticketing systems.

[0052] A first firewall 108a may exist on the client platform 106a-c side and the secure encoding platform side 110. A second firewall 108b may exist on the data partner network 134 side and the secure encoding platform side 110. Both firewall 108a and 108b sets are configured to block certain traffic based on a set of predetermined security rules.

[0053] The plain-text Pll for users 102 is securely transmitted from the client 106 via a secure sockets layer to the IDV Application Programming Interface (API) module or server 112. The IDV API 112 is configured to pass the plain-text Pll to a series ofencoders and normalizers 114 which include, but are not limited to, country-specific address parsers 116, phonetic name encoders 118, an email normalizer 120, a date of birth normalizer 122, and telephone number normalizers 124. Each of these is described in greater detail in the subsequent paragraphs.

[0054] The address parsers 116 are configured to utilize country-specific rules to tokenize, analyze, normalize, and decompose, or any combination thereof, the customer address into individual address components. In exemplary embodiments, the components may comprise unit / apartment number, street number, street name, street type, suburb / city, state / region, and postcode. This granular decomposition enables a more precise alignment during the matching process due to non-standard or irregular data format inputs from users - for example when users or consumers integrate supplementary geographic elements such as suburb, state, and postcode into the primary street address, exemplified by entries like '1 / 23 Elm St Auburn NSW 2034.' The inclusion of additional elements within the street address complicates parsing algorithms by introducing data variability and increasing the risk of misclassification. This is accomplished by tokenizing the address string into subcomponents; flagging subcomponents that are numeric / alphanumeric / single alpha; perform a country-based match for street and unit types and ranking each match for all non- numeric / alphanumeric / single alpha subcomponents; check each alphanumeric subcomponent for matches with a country-base street and unit types, splitting into two subcomponents and ranking if matched; discard subcomponents that match city / state / postcode; identify and parse country-specific address patterns; identify all non-address subcomponents and remove; parse remaining subcomponents into address elements (unit number, street number, street name, and street type).

[0055] The phonetic encoders 118 are configured to utilize language-specific phonetic encoding methods to normalize the customer's name into a phonetic key. This enables more accurate matching of names that having minor spelling variations - for example 'Lawrence, 'Laurence', and 'Lawrance'. Phonetic encoder methods may include, but are not limited to, Soundex, Metaphone, Double-Metaphone, and the New York State Identification and Intelligence System phonetic code (NYSIIS).

[0056] The email normalizer 120 are configured to normalize a customer email address using methods that include, but are not limited to, text case standardization, removal ofplus addressing, and adherence to the RFC 5322 Internet Message Format standard. This will be accomplished using regular expression pattern matching / parsing and text transformation functions. The resultant output may be an RFC 5322 Internet Message Format standardized email address or a metadata representation of the address (e.g. "USERNAME=fred / DOMAIN=farkus.com").

[0057] The date of birth normalizer 122 is configured to translate the customer date of birth format into a standardized format, which could include epoch, Julian, ordinal, or a metadata representation of the date (e.g. "YEAR=YYYY / MONTH=MM / DAY=DD").

[0058] The phone number normalizers 124 is configured to utilize country-specific formatting rules to translate the customer phone number into an International Telecommunication Union Telecommunication Standardization Sector (ITU-T) E.164 standardized format or by means of a novel International Numbering Plan (INP). This INP could accommodate meta data tags about the number and use a combination of alpha numeric representations to support better encryption. For example, "CC+TN-AC212- SN5551234", where CC is Country Code, TN is the Trunk Number, AC is Area Code, and SN is Subscriber Number. Optionally, an International Numbering Plan (INP) may be used. This INP may accommodate meta data tags about the number and use a combination of alpha numeric representations to support encryption; EG Tags may be used for Each Segment: For example, "CC+1-AC212-SN5551234-EXT23", where CC is Country Code, AC is Area Code, SN is Subscriber Number, and EXT is Extension.

[0059] The query distributor 128 is configured to take the normalized, hashed data elements and - based on pre-defined client rules 129 - securely transmits the data via secure sockets layer (SSL) asynchronously to zero to many partner network platforms 134. The query distributor 128 is also configured to pass the data and the client rules129 asynchronously to the internal search engine 130. The query distributor 128 starts a response timer (the length of which is set in the client rules 129 based on pre-defined Service Level Agreements).

[0060] The internal search engine 130 is configured to receive the normalized, hashed data elements, and using the client rules, determine which internal data sources to utilize in the query. A query is made to the key store database 132, which is an internal database of hashed keys, created and refreshed on a periodic basis. The internal search engine130 is configured to return the match status for each of the normalized, hashed dataelements based on the database query. For example, if the query set contain normalized, hashed values for name, date of birth, and address, the returned response may indicate a successful match for name and address, but an unsuccessful match for date of birth.

[0061] The normalized, hashed data elements that have been sent to the partner network platform 134 servers are used to query the data partners' 136a-136x pre-encoded key store databases 137a-137x for matching data elements. Each data partner network platform 136a-136x is configured to return the match status for each of the normalized, hashed data elements based on the database query, using the same response format logic as the internal search engine 130.

[0062] The query distributor 128 is configured to collect all match responses that have been received within the time limit set from the client rules 129 and create a single, amalgamated response object, which is return to the IDV Application Programming Interface (API) module 112. The amalgamated response is then passed back to the original calling client platform 106.The dedicated servers (also referred to as "always on" or headless server) is available to both the organization and partner side and, in embodiments, may be ring fenced from the core business server. The servers 110 and 134 has a processor and memory, wherein the plurality of application modules reside on. The server 110 may also be in communication with an international database from disparate sources, which may comprise a plurality of countries each having different schemes for entering or inputting the data ( PI I). The amalgamated response may comprise sending an identity verification confirmation, a confidence score, or both. An identify verification may be a binary response such as a yes / no, whereas the confidence score may be a score of 1-10 or 1-100 based on the data result sets. As an example, for a user, the system may find that the address provided is correct street number, city, state and country, but not be able to verify the apartment number if that address is an apartment building. As such, a confidence score of 8 out of 10 may be given for that particular portion of data. If the system is only able to verify the user's city, state and country, a confidence score of 3 out of 10 may be provided. In this way, the client may verify or not verify a user's identify based on a set of rules on the client side.

[0063] Referring now to FIG. 2, an exemplary schematic of an Al training module (e.g., neural network) for parsing and normalizing PH such as global addresses is shown generally at 200.

[0064] In operation, data from both client and from client organization 218 is received at the server 102 is segmented 202a-202n+l.Then, once segmentation occurs, the segments are parsed 204a-204n+l. The PH is then normalized at 206a-206n+l and reconfigured into data 208a-208n+l in a way that is usable for fraud prevention by client organization 218. In exemplary embodiments, to maximize the probability of achieving accurate address data matching, both consumer-submitted and client-query address data are segmented into seven distinct components: unit / apartment number, street number, street name, street type, suburb / city, state / region, and postcode. This granular decomposition enables more precise alignment during the matching process. However, due to the vast array of formats that exist machine learning module may be used to increase accuracy. Potential machine learning algorithms would include - but are not limited to - Logistic Regression, K-Nearest Neighbor, Support Vector Machine, and / or K Means Clustering.

[0065] In embodiments, Supervised Learning may be used such as Classification and Regression Trees (CART). This is useful for decision-making processes like determining which address component a piece of data belongs to. Support Vector Machines (SVM) may be used and are effective for classification tasks that can distinguish between different types of address data and Neural Networks may be used (e.g., deep learning models), which can model complex patterns in address data, useful for parsing and standardizing addresses from diverse formats.

[0066] In embodiments, Unsupervised Learning may be used such as Clustering (e.g., K- means, DBSCAN) can be used to identify natural groupings of address data that may not be explicitly labeled, helping to uncover underlying patterns or commonalities in address formatting; Principal Component Analysis (PCA)which the dimensionality of address data, which can be useful in simplifying the data before processing.

[0067] In embodiments, Semi-Supervised Learning may be used which combines a small amount of labeled address data with a large amount of unlabeled data, which can be practical in situations where comprehensive labeling is costly or impractical.

[0068] In embodiments, Reinforcement Learning may be used in a system where the model learns to make better parsing decisions over time based on feedback from the accuracy of its address matching.

[0069] In embodiments, Deep Learning such as: Convolutional Neural Networks (CNNs) can be adapted for sequence data like addresses, such as Recurrent Neural Networks (RNNs) and Long Short-Term Memory Networks (LSTMs) and ideal for dealing with sequences, such as the sequence of components in an address.

[0070] In embodiments, Natural Language Processing (NLP) may be used. Tokenization and Part-of-Speech Tagging techniques can be used to identify and classify components of addresses correctly; or Named Entity Recognition (NER): Could be specifically trained to recognize and extract components of addresses such as street names, city names, etc.

[0071] In embodiments, machine learning models associated algorithms, augmented with specialized processing techniques is provided to provide a high parsing accuracy for a range of international address types, and utilize international data to initially training the nodes on the exemplary Al model 130 (e.g., neural network).

[0072] In the training method, a neural network comprises at least one input layer 210, intermediate layers 212 and output layers 214. In a training step, the input layer 210 is configured to receive the training data from server 130 (and database 126), the confidence output layer 214 is configured to output a predicted confidence map which represents the confidence, predicted by employing the neural network, that the address (or other PI I ) is corrected segmented, parsed, and normalized. In this way, nodes 216 are continuously trained to ensure the Pll is accurate.

[0073] With reference now to FIG. 3, a system architecture block diagram for processing, normalizing, and encoding existing data Pll data for subsequent use in the secure, distributed network for fraud prevention shown is shown generally at 300. The system 300 may be implemented on the data partner network using code-secure technologies.

[0074] To initiate the data normalization and encoding process, the data partner exports their database 302 of existing Personal Identifiable Information (Pll) into a data file 304, which may include, but is not limited, to a comma-separated values (CSV) file or other file formats such as tab-delimited text file, Parquet, JSON, or XML. This is adaptable to new and emerging file formats that may be developed in the future.

[0075] The data file 304 is read by the file reader 308 and each plain-text Pll record is passed to a series of encoders and normalizers 310 which include, but are not limited to, country-specific address parsers 312, phonetic name encoders 314, an email normalizer 316, a date of birth normalizer 318, and telephone number normalizers 320. Each of these is described in greater detail herein.

[0076] The address parsers 312 is configured to utilize country-specific rules to tokenize, analyze, normalize, and decompose the customer address into individual address components. In exemplary embodiments, the components may comprise unit / apartment number, street number, street name, street type, suburb / city, state / region, and postcode. This granular decomposition is configured to enable more precise alignment during the matching process due to non-standard or irregular data format inputs - for example when consumers integrate supplementary geographic elements such as suburb, state, and postcode into the primary street address, exemplified by entries like '1 / 23 Elm St Auburn NSW 2034.' The inclusion of additional elements within the street address complicates parsing algorithms by introducing data variability and increasing the risk of misclassification.

[0077] The phonetic encoders 314 are configured to utilize language-specific phonetic encoding methods to normalize the customer's name into a phonetic key. This enables more accurate matching of names that having minor spelling variations - for example 'Lawrence, 'Laurence', and 'Lawrance'. Phonetic encoder methods may include, but are not limited to, Soundex®, Metaphone®, Double-Metaphone®, and the New York State Identification and Intelligence System phonetic code (NYSIIS).

[0078] The email normalizer 316 is configured to normalize a customer email address using methods that include, but are not limited to, text case standardization, removal of plus addressing, and adherence to the RFC 5322 Internet Message Format standard.

[0079] The date of birth normalizer 318 is configured to translate the customer date of birth format into a standardized epoch format or other format described herein.

[0080] The phone number normalizers 320 is configured to utilize country-specific formatting rules to translate the customer phone number into an International Telecommunication Union Telecommunication Standardization Sector (ITU-T) E.164 standardized format or other formats described herein.

[0081] The now-normalized customer data from the encoders and normalizers 310 is passed to the hash encoders 322. The hash encoders 322 are configured to translate the clear-text normalized PH into a series of keys utilizing hashing methods that include, but are not limited to, methods like the secure hashing algorithm with a 256-bit digest (SHA- 256), message digest algorithms (MD5), and hash salting.

[0082] After each customer record completes the hash encoding 322, it is passed to the file writer 324, which writes each normalized, hashed record to the encoded data file 326, which includes but is not limited to a comma-separated values (CSV) file or other file formats such as tab-delimited text file, Parquet, JSON, or XML. The process is designed to be adaptable to new and emerging file formats that may be developed in the future.

[0083] After the complete plain-text input text file 304 has been encoded and written to the output text file 326, the file is closed and imported into the key store database 137a using the data partners internal Key Store Database import process.

[0084] Once the key store database has been loaded, the partner network platform 134 is available to accept Identity Verification (IDV) queries from the centralized query distributor 128. The partner network platform 134 is securely protected behind firewall 108b which restricts data traffic from / to the centralized query distributor 128.

[0085] When the encoded IDV request is received by the IDV API server 136a, it passes the normalized, hashed data elements to the search engine 328, which queries the key store database 137a. The search engine 328 is configured to return the match status for each of the normalized, hashed data elements based on the database query to the IDV API server 136a, which in returns the match status to the query distributor 128.

[0086] Referring now to FIG. 4, a combination flowchart and architecture executed by software code to process an Identity Verification (IDV) query is shown. At step 402, receiving, at a server, an Identity Verification Request from at least one third party on the network, wherein the IDV request comprises a data set of clear text PI I, and wherein the Pll comprises at least a client address. In operation, the client 106 makes a IDV API request containing a set of customer Personal Identifiable Information (Pll) or protected persona information Pll data in clear text.

[0087] At step 404, normalizing the clear text Pll, wherein normalizing the clear text Pll comprises parsing the client address into a series encoded objects occurs. The IDV API server 112 sends the request data to the standardization / encoding module 114 whichwill take the individual customer Pll elements and employ methods of data standardization and hashing techniques to translate the data into non-PII elements. The methods of standardization include, but are not limited to, parsing the customer address data into a set of unique keys representing the street, premise and household location of the customer, translating the customer's name into a phonetic code, formatting the customer telephone number into a standard international format, and converting the customer date of birth into an epoch date. The methods of hashing techniques include, but are not limited to, methods such as the secure hashing algorithm with a 256-bit digest (SHA-256), message digest algorithms (MD5), and hash salting.

[0088] At step 406, transferring the encoded object to an internal search engine and to the at least one third party provider occurs. An encoded object containing the non-PII data is transferred to the query distributor 128 for subsequent distribution to the internal search engine 130 and zero to many third-party data providers 134, based on the preconfigured rules for each client 106.

[0089] At step 408, at the internal search engine, querying the internal database of normalized data objects for matches to the encoded objects occurs. In operation, the encoded object containing the non-PII data is sent to the internal search engine 130, which will query its internal databases of standardized, hashed data objects for matches to the query data.

[0090] At step 410, receiving, from the internal search engine, a first result set based on the encoded objects that are matched with in the internal database occurs. The encoded object containing the non-PII data is sent to zero to many third-party data providers 134, which will query their internal databases of standardized, hashed data objects for matches to the query data.

[0091] At step 412, receiving, from the internal search engine, a first result set based on the encoded objects that are matched with in the internal database occurs. The internal search engine 130 is configured to return the result set based on the encoded data objects that were matched within the internal database 132. In operations, the use of indexed keys enables the processor / servers / computers to handle massive data volumes within an acceptable time scale. As the encoded element keys are hash values, the uniform distribution of the hexadecimal values enables queries of extremely largevolumes (tens of billions) of records utilizing database-specific techniques such as sharding, data partitioning and key indexing.

[0092] At step 414, within a predetermined time, receiving, from the at least one third party provider, a second result set based on the encoded objects that have been matched in a third-party provider database occurs. The third-party data providers 134 may return their result sets based on the encoded data objects that were matched within their databases, at step 414.

[0093] At step 416, when the internal search engine receives the encoded elements, using third party rules, determining which internal data sources to utilize in a query, and querying a key store database, wherein the key store database is an internal database of hashed keys, and wherein the database of hashed keys are created and refreshed on a predetermined basis. The sets that are returned from third-party data providers 134 outside of the fixed time window may be ignored by the query distributor 128 may not be included in the amalgamated response.

[0094] At step 418, creating a composite IDV response using the first result set, and the second result set occurs. The query distributor takes the individual result sets from the internal search engine 130 and any third-party result sets that were returned with the predetermined time window to create a composite IDV response to the IDV API module 112.

[0095] At step 420, sending an amalgamated response to the client, wherein the amalgamated response comprises an identity verification confirmation, a confidence score, or both. The IDV API module 112 returns the amalgamated response back to the client 106.

[0096] Referring now to FIG. 5, a system architecture block diagram for an implementation of processing, normalizing, and encoding PH data 104 and linking records within a secure, distributed network for fraud prevention is shown generally at 500. The system 500 is a variation of system 100, but with the normalization and encoding modules residing within the client platform 502.

[0097] In this embodiment, Customers or consumers 102 utilizing the client 106 applications enter their clear-text Pll 104 on their user devices 102a-120c. In exemplary embodiments, the data may comprise addresses, dates of birth, email addresses, phonenumbers, and names, but other categories of PH may be ingested from the user based on the needs of the platform network or verifying organization.

[0098] The customer plain-text PH is passed to a series of encoders and normalizers 114 which include, but are not limited to, country-specific address parsers 116, phonetic name encoders 118, an email normalizer 120, a date of birth normalizer 122, and telephone number normalizers 124. Each of these is described in greater detail in herein.

[0099] The address parsers 116 are configured to utilize country-specific rules to tokenize, analyze, normalize, and decompose the customer address into individual address components. In exemplary embodiments, the components may comprise unit / apartment number, street number, street name, street type, suburb / city, state / region, and postcode. This granular decomposition should enable more precise alignment during the matching process due to non-standard or irregular data format inputs - for example when consumers integrate supplementary geographic elements such as suburb, state, and postcode into the primary street address, exemplified by entries like '1 / 23 Elm St Auburn NSW 2034.' The inclusion of additional elements within the street address complicates parsing algorithms by introducing data variability and increasing the risk of misclassification. This is accomplished by tokenizing the address string into subcomponents; flagging subcomponents that are numeric / alphanumeric / single alpha; perform a country-based match for street and unit types and ranking each match for all non- numeric / alphanumeric / single alpha subcomponents; check each alphanumeric subcomponent for matches with a countrybase street and unit types, splitting into two subcomponents and ranking if matched; discard subcomponents that match city / state / postcode; identify and parse countryspecific address patterns; identify all non-address subcomponents and remove; parse remaining subcomponents into address elements (unit number, street number, street name, and street type).

[0100] The phonetic encoders 118 is configured to utilize language-specific phonetic encoding methods to normalize the customer's name into a phonetic key. This enables more accurate matching of names that having minor spelling variations - for example 'Lawrence, 'Laurence', and 'Lawrance'. Phonetic encoder methods may include, but are not limited to, Soundex, Metaphone, Double-Metaphone, and the New York State Identification and Intelligence System phonetic code (NYSIIS).

[0101] The email normalizer 120 is configured to normalize a customer email address using method that include, but are not limited to, text case standardization, removal of plus addressing, and adherence to the RFC 5322 Internet Message Format standard. This will be accomplished using regular expression pattern matching / parsing and text transformation functions. The resultant output may be an RFC 5322 Internet Message Format standardized email address or a metadata representation of the address (e.g. "USERNAME=fred / DOMAIN=farkus.com").

[0102] The date of birth normalizer 122 is configured to translate the customer date of birth format into a standardized format, which could include epoch, Julian, ordinal, or a metadata representation of the date (e.g. "YEAR=YYYY / MONTH=MM / DAY=DD").

[0103] The phone number normalizers 124 is configured to utilize country-specific formatting rules to translate the customer phone number into an International Telecommunication Union Telecommunication Standardization Sector (ITU-T) E.164 standardized format or by means of a novel International Numbering Plan (INP). This INP could accommodate meta data tags about the number and use a combination of alpha numeric representations to support better encryption. For example, "CC+TN-AC212- SN5551234", where CC is Country Code, TN is the Trunk Number, AC is Area Code, and SN is Subscriber Number.

[0104] The now-normalized customer data from the encoders and normalizers 114 is passed to the hash encoders 126. The hash encoders translate the clear-text normalized PH into a series of keys utilizing hashing methods that include, but are not limited to, methods like the secure hashing algorithm with a 256-bit digest (SHA-256), message digest algorithms (MD5), and hash salting.

[0105] The normalized, hashed data elements are passed to a communications module 504 which establishes a secure sockets layer connection the IDV API 112. A firewall 108a may exist on the client platform 502 side and the secure encoding platform side 110. The firewall is configured to block certain traffic based on a set of predetermined security rules, for example to only allow ingress and egress of HTTPS protocol data between the encoding platforms.

[0106] The query distributor 128 is configured to take the normalized, hashed data elements and - based on pre-defined client rules 129 - securely transmits the data via secure sockets layer asynchronously to zero to many partner network platforms 134.The query distributor 128 also is configured to pass the data and the client rules 129 asynchronously to the internal search engine 130. The query distributor 128 starts a response timer (e.g., the length of which is set in the client rules 129 based on predefined Service Level Agreements).

[0107] The internal search engine 130 is configured to receive the normalized, hashed data elements and using the client rules, determines which internal data sources to utilize in the query. A query is made to the key store database 132, which is an internal database of hashed keys, created and refreshed on a periodic basis. The internal search engine 130 is configured to return the match status for each of the normalized, hashed data elements based on the database query. For example, if the query set contain normalized, hashed values for name, date of birth, and address, the returned response might indicate a successful match for name and address, but an unsuccessful match for date of birth.

[0108] The normalized, hashed data elements that have been sent to the partner network platform 134 servers are used to query the data partners' 136a-136x pre-encoded key store databases 137a-137x for matching data elements. Each data partner network platform 136a-136x will return the match status for each of the normalized, hashed data elements based on the database query, using the same response format logic as the internal search engine 130.

[0109] The query distributor 128 is configured to collect all match responses that have been received within the time limit set from the client rules 129 and create a single, amalgamated response object, which is return to the IDV Application Programming Interface (API) module 112.

[0110] The amalgamated response is then passed back to the original calling client platform 502.

[0111] With reference now to FIG. 6, a flowchart for the onboarding of a data partner into the secure Identity Verification network is shown generally at 600. This system enables a data partner to securely normalize and encode their Personally Identifiable Information ( PI I ) on their own server platform and participate in the IDV network without exposing or receiving clear-text PH, which limits the partner's ability to infer or inherit information about a data subject as no PH is re-identifiable. In operation, this limits the partner'sability to infer or inherit information about a data subject as no Pll is re-identifiable because the database may not contain a client ID.

[0112] At step 602, a Data Partner organization initiates the process by requesting access to join the secure Identity Verification (IDV) network. This is the entry point for partners seeking to utilize the platform's capabilities for data verification.

[0113] At step 604, upon approval, the Data Partner installs the encoding and IDV API application onto their computing platform Additionally, they set up a dedicated database to interact with the IDV API. For enhanced security, the Data Partner has the option to configure a firewall for controlled data access.

[0114] At step 606, once the installation process is completed, the Data Partner uses the encoding application to securely encode their existing Pll. This encoding process ensures that sensitive information is protected when integrated with the platform.

[0115] At step 608, the encoded Pll records are then imported into the IDV API database. This allows the data to be processed and verified through the secure network, utilizing the distributed capabilities of the IDV platform.

[0116] At step 610, the Data Partner's IDV API is integrated into the broader network by configuring it within the query distributor. This configuration ensures that queries can be distributed and handled efficiently across the platform.

[0117] At step 612, after successful configuration, the secure partner IDV verification process begins. This allows the Data Partner to verify the identity of individuals or data sets in a secure, encoded manner, leveraging the full capabilities of the platform's distributed network.

[0118] With reference now to FIG. 7, a flowchart for the secure, privacy-compliant data- sharing process within a Clean Room environment, highlighting the use of secure keys and encoding to protect sensitive information shown generally at 700. This system provides the ability for two or more data owners to generate common encoded keys, based on their private Pll data, that can be used in a clean room environment to match data.

[0119] At step 702, two parties (e.g., companies) come to a mutual agreement to share data in a manner that complies with privacy regulations. As part of this agreement, they establish a secure Clean Room environment where data sharing can take place without compromising sensitive information.

[0120] At step 704, both companies install the standardization and encoding application within their respective computing environments. This application is essential for preparing the data for secure sharing.

[0121] At step 706, each company processes their Personally Identifiable Information (PI I ) through the installed application. During this process, the Pll is standardized and encoded, generating unique secure keys for each data set. These secure keys ensure that the data is protected and can be linked between the two companies without directly sharing the raw data.

[0122] At step 708, once the data has been encoded, the secure keys are appended to the Pll data. The keys, not the actual Pll, are then uploaded to the Clean Room system, ensuring that only secure, anonymized data is shared.

[0123] At step 710, inside the Clean Room, the secure keys from both companies are matched. The results of these matches are then downloaded by both companies. This process allows the companies to securely compare and link their data sets without ever sharing unprotected Pll, ensuring both privacy and compliance.

[0124] With reference now to FIG. 8, a flowchart for the secure Identity Verification (IDV) process using encoded data, focusing on the standardization, encoding, and verification steps for protecting sensitive customer shown generally at 800. At step 802, a client who needs to perform secure, encoded Identity Verification (IDV) queries begins by installing the standardization and encoding application within their compute environment. This application is key for handling the sensitive data required for KYC processes.

[0125] At step 804, on the client's platform, a customer enters their Personally Identifiable Information (Pll). This Pll is then sent to the locally installed standardization and encoding application for further processing.

[0126] At step 806, the locally installed application processes the Pll by standardizing and encoding it the resulting encoded elements (referred to as keys) are then sent to the IDV API for secure processing.

[0127] At step 808, the IDV API takes the encoded keys and matches them against its internal key store. Additionally, it can perform optional matching with several external partner network key stores to enhance the verification process.

[0128] At step 810, once the matching process is complete, the IDV API consolidates the results and sends them back to the client. Based on these results, the client can then make an informed decision on whether to proceed with the customer transaction.

[0129] With reference now to FIG. 9, a flowchart for the secure Identity Verification (IDV) process using plain-text PH focusing on standardization, encoding, and matching to ensure a secure and reliable Identity Verification shown generally at 900.

[0130] At step 902, a customer visiting the client's platform inputs their Personally Identifiable Information (PH). This data is then sent to the locally installed application for further processing.

[0131] At step 904, the locally installed application transmits the plain-text PH directly to the Identity Verification (IDV) API for analysis and verification.

[0132] At step 906, the IDV API standardizes and encodes the plain-text PH into secure elements, referred to as keys. This ensures that the data is anonymized and securely processed during the verification.

[0133] At step 908, the encoded keys are matched against the internal key store of the IDV system. Optionally, the system may also match these keys with several external partner network key stores to enhance the accuracy of identity verification.

[0134] At step 910, once the matching process is complete, the IDV API consolidates the match results and sends them back to the client. Based on the verification results, the client can decide whether to proceed with the customer's transaction.

[0135] With reference now to FIG. 10, a flowchart for the fraud detection process using secure encoding and matching, allowing clients to detect and block transactions involving fraudulent Pll shown generally at 1000.

[0136] At step 1002, a client, seeking to perform secure and encoded fraud detection, starts by installing the standardization and encoding application within their compute environment. This application is essential for handling the sensitive data needed for detecting fraudulent activities.

[0137] At step 1004, the client receives a file containing fraudulent Personally Identifiable Information (Pll) from a third-party source. This file is processed through the standardization and encoding application, where the fraudulent Pll is encoded into secure elements, known as keys. The encoded keys are then stored in the client's local database for future reference.

[0138] At step 1004, when a customer enters their Pll into the client's application (e.g., during a transaction), the application triggers the locally installed standardization and encoding application. This application encodes the customer's Pll into secure keys, similar to the process done with the fraudulent Pll.

[0139] At step 1006, the encoded keys generated from the customer's Pll are then matched against the internal fraudulent key store, which contains the previously encoded fraudulent Pll.

[0140] At step 1008, if a match is found between the customer's encoded Pll and the fraudulent key store, the system flags the transaction as potentially fraudulent. The client is then prompted to reject the customer's transaction to prevent fraud.

[0141] With reference to FIG. 11 illustrates process map for the parser of a postal address shown in accordance with one embodiment. This system enables the free-form address to by normalized into a standard set of address elements and address keys, which facilitates improved matching speed and accuracy.

[0142] At step 1102, receiving an input of the address from a client, wherein the address comprises address elements and removing or replacing punctuation dependent upon a type of punctuation and a type of characters surrounding the punctuation to form an address string occurs. The free-form postal address is received as input to the address parser 116 and undergoes a series of transformations to remove and / or replace punctuation marks dependent on the type of punctuation mark and the characters surrounding the punctuation mark.

[0143] At step 1104, matching the address string against a country specific table to identify if the address has a numeric street name, and if the address string does not contain one of the numeric street names, assigning a street name metadata label to a portion of the address string and checking each of the address elements for the presence of the metadata labels occurs. This transformed address string is matched against countryspecific tables to identify if the address string contains a numeric street name (for example: "10 Septembre 1943" or "25 de Diciembre"). If the address string does contain one of the numeric street name patterns, a Street Name metadata label is assigned to that portion of the address string.

[0144] At step 1106, the address string is split into multiple address elements using the space delimiter. If the address string has a portion identified with a Street Namemetadata label, this portion of the address string is preserved as a single element (e.g. "10 Septembre 1943" remains as a single element and is not split into the three elements "10", "Septembre", and "1943").

[0145] At step 1108, each address element is checked for the presence of a metadata label. The elements that are not labelled are passed to step 1110, where the element's text is examined to match certain patterns, such as all numeric, alphanumeric, or a single alpha character. If the element's text matches any of these patterns, a metadata label is assigned to the address element indicating the corresponding text pattern type.

[0146] At step 1112, each address element is checked for the presence of a metadata label. The elements that are not labelled are passed to step 1114, where the element's text is examined to match against a country-specific set of street and unit types. If the element's text matches any of these patterns, a metadata label is assigned to the address element indicating the corresponding text pattern type and a ranking of priority (which has been pre-assigned to each country-specific set of street and unit types).

[0147] At step 1116, each address element is checked for the presence of a metadata label. The elements that are not labelled are passed to step 1118, where the element's text is examined to match against the city, state, and postcode that were supplied as part of the full postal address. If the element's text matches any of these patterns, an 'ignore' metadata label is assigned to the address element.

[0148] At step 1120, the full array of address elements is parsed in sequence, examining the country-specific relationships between the address elements that have been assigned metadata labels & priority ranking and their adjacent address elements to extract the unit number, street number, street name, and street type (known as 'Address Line 1 components').

[0149] At step 1122, the Address Line 1 components, other postal address components, and the country code are used to create three hashed address keys: the Street Key, the Premise Key, and the Household Key.

[0150] FIG. 12 illustrates a stepwise diagram for processing Pll and securely linking records within databases for identity verification in accordance with one embodiment.

[0151] The steps comprise receiving, at a server, a data request from at least one third party in the computing environment, wherein the data request comprises a data set of clear text Pll of a customer, step 1202, normalizing the clear text Pll, whereinnormalizing the clear text PH comprises parsing the clear text Pll into a series encoded objects, step 1204, transferring the encoded objects to an internal search engine step 1206, at the internal search engine, querying an internal database of normalized data objects for matches to the encoded objects, step 1208, receiving, from the internal search engine, a first result set based on the encoded objects that are matched within the internal database step 1210 and forming an amalgamated response, wherein the amalgamated response comprises an identity verification confirmation, a confidence score, or both step 1212.

[0152] Preferred embodiments of this invention are described herein, including the best mode known to the inventors for carrying out the invention. It should be understood that the illustrated embodiments are exemplary only and should not be taken as limiting the scope of the invention.

[0153] The foregoing description comprise illustrative embodiments of the present invention. Having thus described exemplary embodiments of the present invention, it should be noted by those skilled in the art that the within disclosures are exemplary only, and that various other alternatives, adaptations, and modifications may be made within the scope of the present invention. Merely listing or numbering the steps of a method in a certain order does not constitute any limitation on the order of the steps of that method. Many modifications and other embodiments of the invention will come to mind to one skilled in the art to which this invention pertains having the benefit of the teachings in the foregoing descriptions. Although specific terms may be employed herein, they are used only in generic and descriptive sense and not for purposes of limitation. Accordingly, the present invention is not limited to the specific embodiments illustrated herein.

Claims

CLAIMSI claim:

1. A method for normalizing personal identifiable information (PI I ) and for identity verification in a computing environment, the method comprising: receiving, at a server, a data request from at least one third party in the computing environment, wherein the data request comprises a data set of clear text Pll of a customer; normalizing the clear text Pll, wherein normalizing the clear text Pll comprises parsing the clear text Pll into a series encoded objects; transferring the encoded objects to an internal search engine; at the internal search engine, querying an internal database of normalized data objects for matches to the encoded objects; receiving, from the internal search engine, a first result set based on the encoded objects that are matched within the internal database; forming an amalgamated response, wherein the amalgamated response comprises an identity verification confirmation, a confidence score, or both.

2. The method of claim 1, wherein: the series of encoded objects are a series of keys; the computing environment is a communications network; and the Pll comprises a customer address.

3. The method of claim 1, wherein the data request comprises an Identity Verification Request an (IDV) request, a Know Your Customer (KYC) query, a clean room environment request, or a combination thereof.

4. The method of claim 3, further comprising: within a predetermined time, receiving, from the at least one third party, a second result set based on the encoded objects that have been matched in a third- party database; andcreating a composite IDV response using the first result set and the second result set.

5. The method of claim 1, wherein parsing the customer address further comprises utilizing country-specific rule sets to decompose the customer address into individual address components to normalize the individual address components.

6. The method of claim 5, wherein parsing the customer address further comprises: receiving an input of the address from the customer, wherein the address comprises address components; removing or replacing punctuation dependent upon a type of punctuation and a type of characters surrounding the punctuation to form an address string; matching the address string against a country specific table to identify if the address string has a numeric street name, and if the address string does not contain the numeric street name, assigning a street name metadata label to a portion of the address string; checking each of the address components for the presence of the metadata labels.

7. The method of claim 6, further comprising: if the address components are not metadata labeled, evaluating the components to match at least a predetermined pattern, wherein the pattern comprises numeric, alphanumeric or single alpha characters; if the address components match the patterns, assigning an additional metadata label the address component indicating the patterns; checking again for the presence of the metadata labels, and if the address component is not labeled, examining the address components text to match against a country-specific set of street and unit types;if the address components text matches any of the patterns, assigning an additional metadata label to the address element indicating the corresponding text pattern type; ranking a priority to each country-specific set of street and unit types; checking again for the presence of the metadata labels, and if the address element is not labeled, examining the address components to match the address component to a city, state, postcode, or any combination thereof that were part of the customer address; if the address element's text matches any of these patterns, assigning an ignore metadata label the address element; parsing a full array of address elements in sequence; examining the country-specific relationships between the address elements that have been assigned metadata labels, the priority ranking and their adjacent address components to extract a unit number, street number, street name, street type, or any combination thereof, wherein the unit number, street number, street name, street type is labeled as address line one components; using the address line one components, and a country code to create three hashed address keys comprising a street key, a premise key, and a household key.

8. The method of claim 2, further comprising passing the clear text Pll to a hash encoder, wherein the hash encoder is configured to translate the clear text Pll into the series of keys using hashing algorithms comprising a 256-bit digest, message digest algorithms (MD5), hash salting, or any combination thereof.

9. The method of claim 1, further comprising sending the encoded objects to a third party network platform server, wherein the third party platform network server is configured to query a pre-encoded key store databases on the third party network to match the encoded objects and return a match status for each of the encoded based on the database query.

10. The method of claim 1, further comprising normalizing customer names, dates of birth, emails, telephone numbers, email address, or any combination thereof.

11. A system for normalizing personal identifiable information ( PI I) and for identity verification in a computing environment, the system comprising: one or more memories storing instructions; one or more processors executing the instructions to: receive, at a server, a data request from at least one third party in the computing environment, wherein the data request comprises a data set of clear text Pll of a customer; normalize the clear text Pll, wherein normalizing the clear text Pll comprises parsing the clear text Pll into a series encoded objects; transfer the encoded objects to an internal search engine; at the internal search engine, query an internal database of normalized data objects for matches to the encoded objects; receive, from the internal search engine, a first result set based on the encoded objects that are matched within the internal database; form an amalgamated response, wherein the amalgamated response comprises an identity verification confirmation, a confidence score, or both.

12. The system of claim 11, wherein: the series of encoded objects are a series of keys; the computing environment is a communications network; and the Pll comprises a customer address.

13. The system of claim 11, wherein the data request comprises an Identity Verification Request an (IDV) request, a Know Your Customer (KYC) query, a clean room environment request, or a combination thereof.

14. The system of claim 13, further comprising causing the processor to: within a predetermined time, receive, from the at least one third party, a second result set based on the encoded objects that have been matched in a third- party database; and create a composite IDV response using the first result set and the second result set.

15. The system of claim 12, wherein parsing the customer address further comprises causing the processor to utilize country-specific rule sets to decompose the customer address into individual address components to normalize the individual address components.

16. The system of claim 15, wherein parsing the customer address further comprises causing the processor to: receive an input of the address from the customer, wherein the address comprises address components; remove or replace punctuation dependent upon a type of punctuation and a type of characters surrounding the punctuation to form an address string; match the address string against a country specific table to identify if the address string has a numeric street name, and if the address string does not contain the numeric street name, assigning a street name metadata label to a portion of the address string; check each of the address components for the presence of the metadata labels.

17. The system of claim 16, further comprising causing the processor to: if the address components are not metadata labeled, evaluate the components to match at least a predetermined pattern, wherein the pattern comprises numeric, alphanumeric or single alpha characters; if the address components match the patterns, assign an additional metadata label the address component indicating the patterns; check again for the presence of the metadata labels, and if the address component is not labeled, examine the address components text to match against a country-specific set of street and unit types; if the address components text matches any of the patterns, assign an additional metadata label to the address element indicating the corresponding text pattern type; rank a priority to each country-specific set of street and unit types;check again for the presence of the metadata labels, and if the address element is not labeled, examining the address components to match the address component to a city, state, postcode, or any combination thereof that were part of the customer address; if the address element's text matches any of these patterns, assign an ignore metadata label the address element; parse a full array of address elements in sequence; examine the country-specific relationships between the address elements that have been assigned metadata labels, the priority ranking and their adjacent address components to extract a unit number, street number, street name, street type, or any combination thereof, wherein the unit number, street number, street name, street type is labeled as address line one components; use the address line one components, and a country code to create three hashed address keys comprising a street key, a premise key, and a household key.

18. The system of claim 12, further comprising causing the processor to pass the clear text PH to a hash encoder, wherein the hash encoder is configured to translate the clear text Pll into the series of keys using hashing algorithms comprising a 256-bit digest, message digest algorithms (MD5), hash salting, or any combination thereof.

19. The system of claim 11, further comprising sending the encoded objects to a third party network platform server, wherein the third party platform network server is configured to query a pre-encoded key store databases on the third party network to match the encoded objects and return a match status for each of the encoded based on the database query.

0. A non-tangible, computer-readable recording medium storing instructions, that, when executed by a processor, cause the processor to: receive, at a server, a data request from at least one third party in the computing environment, wherein the data request comprises a data set of clear text Pll of a customer; normalize the clear text Pll, wherein normalizing the clear text Pll comprises parsing the clear text Pll into a series encoded objects; transfer the encoded objects to an internal search engine; at the internal search engine, query an internal database of normalized data objects for matches to the encoded objects; receive, from the internal search engine, a first result set based on the encoded objects that are matched within the internal database; form an amalgamated response, wherein the amalgamated response comprises an identity verification confirmation, a confidence score, or both.

Citation Information

Patent Citations

  • Online data decomposition

    US11729074B1

  • Establishing a link between identifiers without disclosing specific identifying information

    US20180218168A1

  • Contact discovery service with privacy aspect

    US20190340385A1

  • Massive scale heterogeneous data ingestion and user resolution

    US20220138238A1

  • Method, apparatus, and computer-readable medium for postal address indentification

    US20220327403A1