Systems and methods for language-agnostic name comparison

The language-agnostic name screening system addresses the challenge of comparing names across languages by standardizing and phonetically representing names in a vector space, achieving efficient and accurate matching and improving regulatory compliance.

WO2025125874A1PCT designated stage expired Publication Date: 2025-06-19MOZN INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2023/062645
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-14
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing name screening and threat detection systems face challenges in comparing names across different languages and transliterations, leading to inefficiencies and inaccuracies, particularly when dealing with non-Latin script names.

Method used

A language-agnostic name screening system that standardizes names, generates phonetic phrases, and represents them in a vector space for comparison with prescreened entities, allowing for accurate matching across different languages and transliterations.

Benefits of technology

The system enables efficient and accurate comparison of names from various languages and transliterations, reducing false positives and negatives, and improving compliance with regulations such as anti-money laundering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2023062645_19062025_PF_FP_ABST
    Figure IB2023062645_19062025_PF_FP_ABST
Patent Text Reader

Abstract

A system for language-agnostic name matching includes a processor configured to phonetically screen a name of interest by comparison against prescreened names or entities. The name of interest is transliterated to a predetermined language, and the name is parsed by the processor using predetermined parsing logic to generate a plurality of name components. For each name component, the processor generates a phonetic embedding and computes a measure of similarity against phonetic embeddings associated with the prescreened names to identify possible matches and / or a degree of similarity between the name and one or more of the known prescreened names.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR LANGUAGE-AGNOSTIC NAME COMPARISONFIELD

[0001] The present disclosure generally relates to cyber technologies such as fraud and threat detection; and in particular, to systems and one or more methods for phonetically comparing names for screening purposes or otherwise.BACKGROUND

[0002] A prominent challenge in the domain of name screening and general threat detection is the abundant use of English language in adverse media, sanctions lists, or even social media. That means non-English names are transliterated into English. This transliteration process is often performed by nonEnglish speakers, resulting in variations of spellings. For example, when the name in Arabic script is converted to English, it might take several forms “Hashem”, “Hashim”, “Hachem”, or “Hachim.” Naturally, a challenge that follows is which one of those possible transliterations should be used to search the adverse media, social media, sanctions lists or any other sources of interest. Another challenge is that most algorithms in the name screening and threat detection domains are developed and implemented based on the English language, and do not consider or sufficiently address non-Latin script languages. A typical or conventional solution would likely require keeping a list of known transliterations for each name. Nevertheless, maintaining, updating, and validating such a list is technically impractical.

[0003] It is with these observations in mind, among others, that various aspects of the present disclosure were conceived and developed.SUMMARY

[0004] The present disclosure provides a number of examples that describe language-agnostic name screening such as matching and associated operations. In the context of the disclosed methods, devices, techniques, apparatus, systems, and so on, the terms “operable to,” “configured to,” and “capable of’ used herein are interchangeable.

[0005] In a first set of illustrative examples, the language-agnostic name screening techniques are embodied by a system including a processor. The processor accesses instructions from a memory that configures the processor to: access a name for screening; create a standardized name in a language based on the name, the language having a first alphabet of characters and a direction of pronunciation; generate at least one phonetic phrase based on respective groupings of characters in the direction of pronunciation for the standardized name; generate an embedded phrase to represent the at least one phonetic phrase in a vector space; access a plurality of embedded candidate phrases in the vector space based on a relative proximity to the at least one phonetic phrase, the plurality of embedded candidate phrases corresponding to prescreened entities; and compute a degree of similarity between the embedded phrase and an embedded candidate phrase associated with at least one prescreened name.

[0006] In a second set of illustrative examples, the language-agnostic name screening techniques are embodied by a method. The method can comprise steps of: standardizing a name into a script defining a standardized representation of the name associated with a language; parsing the script of the name into a plurality of name components in the language; accessing a plurality of prescreened names associated with entities including one or more prescreened components defining prescreened phonetic embeddings; and computing, for each component of the plurality of name components, a phonetic similarity with the one or more prescreened components using one or more measures of similarity. Computing the phonetic similarity with the one or more prescreened components using one or more measures of similarity can include generating a phonetic embedding from phonetic features extracted from each component, the phonetic embedding defining a vector, and measuring a similarity in a vector space between the phonetic embedding of the component and the prescreened phonetic embeddings.

[0007] The foregoing examples broadly outline various aspects, features, and technical advantages of examples according to the disclosure in order that the detailed description that follows may be better understood. It is further appreciated that the above operations described in the context of the illustrative example method, device, and computer-readable medium are not required and that one or more operations may be excluded and / or other additional operations discussed herein may be included. Additional features and advantages will be described hereinafter. The conception and specific examples illustrated and described herein may be readily utilized as a basis for modifying or designing other structures for carrying out the same purposes of the present disclosure. Such equivalent constructions do not depart from the spirit and scope of the appended claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0008] FIG. 1A is a simplified block diagram of a general computer- implemented system for leveraging artificial intelligence such as machine learning for language-agnostic name screening such as matching.

[0009] FIG. 1 B is a simplified block diagram demonstrating example general services the processor of FIG. 1A can be configured to execute.

[0010] FIG. 2 is a general flow diagram illustrating example functionality associated with the system of FIGS. 1A-1 B.

[0011] FIG. 3 is a simplified diagram of an example implementation of the system of FIG. 1A for transliterating names of any language into a common predetermined format such as common Latin script.

[0012] FIG. 4 is an illustration of phrases vs. entities vector space demonstrating that searching in the ‘phrase space’ is much faster than searching in the entities space.

[0013] FIG. 5 is a simplified flow diagram illustrating example functionality for creating a cache to provide two-way mapping between entities and phrases (tokens).

[0014] FIG. 6 is a simplified process flow of an example method that can be executed by the system of FIGS. 1A-1B and can include functionality of FIG. 2 and other features described herein.

[0015] FIG. 7 is an exemplary simplified block diagram of a computing device that may be configured to implement various methodologies described herein.

[0016] Corresponding reference characters indicate corresponding elements among the view of the drawings. The headings used in the figures do not limit the scope of the claims.DETAILED DESCRIPTION

[0017] Aspects of the present disclosure relate to computer- implemented systems and associated methods for indexing and phonetically comparing names of a subject (entities, organizations, individuals, or any combinations thereof) with other names, for, e.g., screening purposes and general threat detection. In some examples, the systems and methods can be used to screen names of customers of financial institutions against watchlists, which helps for complying with anti-money laundering regulations. In other examples, the systems and methods leverage advanced natural language processing (NLP) and artificial intelligence (Al) techniques to phonetically match names to verify payees of financial transactions. The systems and methods described herein can match, or at least compute a degree of similarity between names from different languages and names with different variations of spelling based on transliterations and phonetic representation as further described herein.

[0018] Each of the foregoing examples of the system may leverage information from one another. In addition, other examples, features, and subfeatures are contemplated; and examples of the system are not mutually exclusive, i.e., features of one example can be shared and implemented by other examples. For instance, one example is contemplated that includes all of the features of the main examples described herein. Other variations are contemplated and supported.

[0019] Introduction (& Technical Challenqes / Problems)

[0020] Name Parts Recognition

[0021] Several cultures use different naming conventions and there is no size that fits all cultures. Spanish naming conventions usually require having a surname that is composed of two separate names (e.g., “RODRIGUEZ OLIVERA, Luis”). Arabic naming conventions usually require having the name of the father as a middle name. A typical solution would be to consider the last part of the name to be a last name when a comma is not present. Otherwise, the last name is the part preceding a comma. Those conventions will eventually result in conflicts in identifying which part of the name belongs to the given name, middle name, and last name. Based on our observations, for instance, people who are not familiar with the Arabic culture will consider the last part of the following names “Abd Al-Bari”, “Abd Al-Basit”, and “Abd Al-Baki” as a last name, “AL-BARI, 'Abd”, “AL-BASIT, Abd”, and“AL-BAKI, 'Abd.” The problem is that people who are committing such cultural mistakes are usually those who build and collect media and sanctions databases.

[0022] Matching Techniques

[0023] Edit-Distance: Many name-matching solutions are promoting their products using variants of the well-known edit-distance algorithm. The variants include but are not limited to “Levenshtein”, “Jaro”, or “Jaro-Winkler”. This algorithm and all its variants are sensitive to short names. A basic example would be the names and “^U” which can be transliterated into “Said” or “Saeed” and “Raid” or “Raeed”. All variants of this algorithm will fail because they count the number of modifications to turn one string into another. For example, let us apply the edit distance algorithm on “Said” and “Saeed.” First, turn “i” into an “e,” and second add an extra “e” to turn “Said” into “Saeed,” with a total of 2 modifications. On the other hand, “Said” and “Raid,” require 1 modification to turn S into R. Indeed, this algorithm and all its variants will consider "Said" more similar to “Raid” than to “Saeed.” Consequently, when searching for “Said” it will be missing “Saeed” as a potential hit (i.e., false negative).

[0024] Phonetic Index: More sophisticated solutions to the name matching problem suggest transforming text into a compressed representation that captures the phonetic features to speed up the searching process. There are several well-known algorithms achieving this goal, such as Soundex, Metaphone, and Double-Metaphone. Table 1 below shows the output of these algorithms for a set of example names.Table 1: Example output

[0025] From the first glance, one might conclude that these algorithms solve the problem of “Said” and “Raid” we discussed above. However, these algorithms have a radical bias towards American pronunciation, and they also introduce insensitivity to vowels. Table 2 below shows how these algorithms are failing:Table 2: Drawbacks of known algorithms

[0026] An observant reader will conclude that “Said” and “Saad” will have the same phonetic codes. The same is true for “Hasan” and “Hussein” and the list goes on. As a result, this approach will result in more false positives leading to hours of manual verification.

[0027] Hybrid Approach: Current name-matching solutions have recognized the weaknesses of the prior approaches and challenges addressed above. It has been realized that edit-distance will cause false negatives and phonetic-indexing will cause false positives, so combining several approaches would seem to alleviate the shortcomings of relying on a single approach. A hybrid approach seems reliable and convenient; however, there are major drawbacks that will be discussed in the following subsection.

[0028] Drawbacks to Prior Approaches: A major drawback of all previous approaches is that they are sensitive to unusual spellings. For instance, the name can be spelled commonly by “Haitham” or “Haytham.” However, there is a less common spelling for the exact name “Haisam.” Another drawback is insensitivity to similar sounding names. For example, the namescan be spelled commonly as “Awad,” “Aiedh,” “Awaadh,” “Eidha.” While the editdistance approach will lead to a very low similarity between the common and less common spellings, the phonetic indexing not only will yield irrelevant names such asbut also different phonetic codes.Table 3: Spelling issues associated with previous approaches

[0029] Given the above major drawbacks, it is imperative to develop solutions that capture finer phonetic details. To illustrate that, consider the lettersand which can be mapped to different English letters such as “d,” “dh,” or “z.” A similar sounding letter is “j” and can be written as “th” or “z.” While the lettercan be written as “th” or “t.” Given these mappings, one might conclude that the letteris more similar to “j” than toIndeed, those finer details can be captured in the inventive approach as further described herein.

[0030] Exemplary Technical Solutions

[0031] Referring to FIGS. 1A-1 B, an inventive concept responsive to the aforementioned technical challenges may take the form of a computer- implemented system, designated system 100, comprising any number of computing devices or processing elements. In general, the system 100 leverages artificialintelligence for language-agnostic name comparison including screening or matching, by comparing aspects of a name of interest with candidate entities. While the present inventive concept is described primarily as an implementation of the system, it should be appreciated that the inventive concept may also take the form of tangible, non-transitory, computer-readable media having instructions encoded thereon and executable by a processor, and any number of methods related to examples of the system described herein.

[0032] In some examples, the system 100 includes at least one processor 102 or processing element, and at least one of a memory 103 or storage device 103 storing instructions 104 accessible by the processor 102 to perform various functions and operations described herein. The system 100 can further include a network interface 106 (or multiple network interfaces), and a bus (or wireless medium) for interconnecting the aforementioned components. The network interface 106 includes the mechanical, electrical, and signaling circuitry for communicating data over links (e.g., wires or wireless links) within a network (e.g., the Internet). The network interface 106 may be configured to transmit and / or receive data using a variety of different communication protocols, as will be understood by those skilled in the art.

[0033] In general, the processor 102 is configured (via the instructions 104) to execute services (150 in FIG. 1B) and perform operations including implementation of one or more models (e.g., Al models) to phonetically compare an input name with known or prescreened names (or entities). In various examples, the input name can be accessed from input data 114A provided to an external computing device 108 (e.g., by a user engaging a user interface 112 rendered along a display 110), and the processor 102 can return an output (e.g., a decision) associated with the input name in the form of output data 114B for access by the computing device 108. In addition, the processor 102 can leverage datasets and other information to train, tune, and / or update one or more artificial intelligence models which can be leveraged to screen and / or to perform language-agnostic name screening / matching functions as described herein. For instance, the processor 102 can access data from one more data source devices 120 shown by example as device 120A, device 120B, and device 120C. Devices 120 can provide access to, e.g., large datasets of entities and names and English transliterations thereof and other information for use by the processor 102 or otherwise. Datasets and otherinformation can be preprocessed and stored within a database 118 as shown. Al models as referenced herein can include classification models, supervised or unsupervised learning models such as K Nearest Neighbors, linear regression models, neural networks, deep learning models, etc.

[0034] As indicated in FIG. 1 B, the instructions 104 executed by the processor 102 may define various services 150 associated with the languageagnostic name comparison techniques described herein. In some examples, the services 150 can include standardization and preprocessing 150, which can include functionality for generating a standardized representation of a name of interest including, e.g., reordering elements of the name to a predetermined format, removing stop words, and / or removing elements from the name such as honorifics. Services 150 can include transliteration 156 for transliterating names written in any language into a predetermined common language, such as a common Latin script or English. Services 150 can include parsing 154 which defines parsing logic for parsing the name of interest into individual components, also referred to herein and interchangeable with tokens, character grams, n-grams, or simply portions of the name of interest. Services 150 can include phonetic embedding 158 for generating a phonetic embedding from phonetic features extracted from the name components derived from the parsing 154 service. Services 150 can further include screening & comparison 158 including functionality for executing one or more measures of similarity (such as phonetic comparison) to derive a degree of similarity between the name and one or more entities or prescreened names (of, e.g., a predetermined list such as a watchlist). In some examples, “names” or entities being compared can be supplemented with other attributes (e.g., date of birth, nationality, etc.) and the attributes can supplement the comparison. The services 150 shown are not intended to be exhaustive but merely illustrated to demonstrate example concepts and operations associated with the system 100. Implementation of the services 150 and further details are elaborated upon in FIG. 2 and associated description below.

[0035] The instructions 104 and / or services 150 defined thereof can be implemented as code and / or machine-executable instructions executable by the processor 102 that may represent one or more of a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, an object, a software package, a class, or any combination of instructions, data structures, or program statements, and the like. In other words, the instructions 104 or any operationsperformed by the processor 102 described herein may be implemented by hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware or microcode, the program code or code segments to perform the necessary tasks (e.g., a computer-program product) may be stored in a computer-readable or machine-readable medium (e.g., the memory 103), and the processor 102 performs the tasks defined by the code.

[0036] Referring to FIG. 2, a general example process 200 associated with language-agnostic name comparison is shown that can be implemented by the processor 102 or otherwise. As indicated, the processor 102 of FIGS. 1A-1B may access or receive a process request 201, such as a query including input elements 202 defining a name 204 of interest. The input elements 202 can include other parameters 203 associated with the name 204, such as a nationality 203A and a date of birth (DOB) 203B. As indicated, the preprocessing 152 can be applied to the name 204 and other aspects of the input elements 202. For example, the name 204 can be preprocessed to generate a standardized representation of the name 204 including, e.g., reordering elements of the name 204 to a predetermined format and / or removing elements from the name such as honorifics, stop words, and the like.

[0037] As indicated in FIG. 2, the processor 102 can apply the transliteration 156 service to derive a transliterated version of the name 204, or transliterated name 206. In this manner, the transliterated name 206 is expressed in a predetermined common language, such as common Latin script, or English text. As indicated in FIG. 3 and further elaborated upon herein, transliteration may be omitted where the name 204 of the input elements 202 is already expressed in the predetermined common language. Transliteration can include transliterating from Arabic to English orthography using a deep learning model, but it also can include normalizing names, removing non-English letters and diacritics. As indicated in FIG. 2, the name 204 as transliterated (transliterated name 206) is represented by the elements “alhasan” and “ahmed” after parsing and transliteration.

[0038] As further indicated, parsing logic can be executed by the processor 102 in view of the name 204 as preprocessed and / or standardized to derive one or more name components 208. Parsing logic, as further elaborated upon below, can be configured in different ways where the name 204 represents aperson vs. a company or organization. In some examples, applying the parsing logic (via the parsing 154 service) gives results of name components 208 where persons’ names, include components such as a first-name, middle-name, and last-name; and organizations’ names, gives results such as an organization-type, organizationname, and organization-commercial-activity.

[0039] As further indicated, phonetic embedding 158 can be applied to the name components 208 to derive one or more phonetic embeddings 210 associated with the name 204. In general, the phonetic embeddings 210 are numerical representations, or vectors, defining pronunciation or phonetic properties of the name 204. In other words, the processor 102 generates phonetic embeddings 210 from phonetic features extracted from the transliterated name components 208. In some examples, the name components 208 can define phrases 211 , e.g., abstractions of each component, and the phonetic embedding 158 service extracts phonetic features from a phrase 211 and generates the embedding vectors (phonetic embeddings 210) from the phrases 211. In some examples, phrases 211 can be associated with a class containing a string token, phonetic features, and an index.

[0040] Neighbors 212 represent neighboring phrases, or phrases of prescreened entity names closest to the phrases 211 of the name 204 in a phrase vector space. In some examples, the phonetic embeddings 210 associated with the phrases 211 can be fed to a K-nearest neighbors’ (KNN) algorithm to return the nearest (e.g., 100) phrases neighboring or closest to the phrases 211 in a phrase vector space. Examples of the K-nearest neighbors’ algorithm implementations include Hierarchical Navigable Small Worlds (HNSW) Approximate KNN, and / or Greedy KNN. In general, neighbors 212 include phonetically similar phrases compared with the phrases 211 associated with the subject name 204.

[0041] Candidate matches 214 define possible matches with the name 204 of interest. In the example provided, candidate matches 214 include “Hasan Alahmed,” “Ahmad Alhasan,” and “Ahmed Alhassaan.” Candidate matches 214 may be retrieved from a cache 216. At this stage of the process 200, the names of the candidate matches 214 are at least phonetically similar with the name 204 as transliterated, including components of “alhasan” and “ahmed.”

[0042] As further shown in the process 200 of FIG. 2, candidate entities 218, or entities associated with the candidate matches 214 can be retrieved from the database 118. The candidate entities 218 include, for example, other parametersassociated with the names of the candidate matches 214, such as the DOB and nationality associated with each name of the candidate matches 214.

[0043] Finally, represented by “score” 220, the processor 102 can compute one or more similarity calculations or measures of similarity between phonetic embeddings associated with the name 204 and phonetic embeddings associated with the candidate entities 218 to output a degree of similarity or other measure of comparison for each name of the candidate entities 218. Alternatively, the “score” 220 can represent one or more similarity measures associated with the name, as well as other attributes associated with the candidate entities, such as date of birth or nationality. Examples of scoring are provided below and illustrated in Table 4. In some examples, as shown in FIG. 2, the processor 102 can return “Final Entities” or a list of possible matches 222 with the name of interest 204. The possible matches 222 are ordered such that the name of the possible match in the first row of the possible matches 222, “Ahmad Alhasan,” is determined to be most phonetically similar to the name 204 derived from the original process request 201.

[0044] Further detail associated with aspects of the process of FIG. 2 is set forth below.

[0045] Transliteration:

[0046] Referring to FIG. 3, examples of the system 100 can be configured for implementing a transliteration model to transliterate names written in any language into a predetermined format, such as a common Latin script. For instance, Arabic names can be transliterated to a common Latin script, or the English language. The transliteration model can transliterate names across multiple languages, such as transliterating Arabic and Greek names to a common a Latin script. Multiple transliteration models may also be used, such as a transliteration model for Arabic names to a Latin script and a transliteration model for Greek names to a Latin script.

[0047] In one example, a transliteration model can include a large dataset (stored, e.g., in database 118 of FIG. 1A) of hundreds of thousands of Arabic names and their possible English transliterations compiled from various sources. The following tuples are examples from the compiled dataset: (“^ , “Mosa”),“Raghad”). A preprocessing method can be applied to the input to make a standard representation of the text. For example, using small casing in English and turning “LS” into “ I ”. After that, these tuples can then be broken into small elements (known inthe literature of NLP as character-grams), resulting into tuples of this nature (“ OJ j* “Mo os sa”). The subject transformation makes the transliteration model generalizable on the gram level. To clarify, let us assumehas not been observed by our model,both have the gram “OJ,” so knowing how to transliterate “u-T forwill help the model to transliteratefor the name After transforming the dataset into grams as described, the dataset can then be fed to a deep learning (DL) model, e.g., Seq2Seq LSTM with Attention. In the preceding example, the inputs to the model are elements in Arabic, and the outputs of the model are elements in English. This example DL architecture was chosen carefully after evaluating the results thereof. FIG. 3 shows an example of the transliteration model and functionality thereof. In contrast to the described inventive approach of using components such as character-grams and a transliteration model, a standard approach of looking up a requested name (e.g., in Arabic) in a database to find a corresponding match (e.g., in English) will fail if the database does not contain an exact match for the requested Arabic name.

[0048] In general, the aforementioned transliteration steps accommodate non-Latin script being transliterated into English. The objective of this transliteration is to standardize the inputs into a common Latin script, which yields accurate measures of similarity.

[0049] Parsing names for identifying and reordering name strings into predetermined (regular) formats):

[0050] It was previously described that relying on data sources for determining the parts of the name is not a valid approach because the data sources might not be familiar with cultural conventions of naming. This challenge can be addressed by (i) reordering the name string associated with the name 204 of interest into a regular format, such as given-name middle-name and surname, (ii) Next, a name parser can be applied that is designed to preserve different naming conventions. For example, an input name, “Khalil Abd ur Rahman” can be shown on the United States Office of Foreign Assets Control (OFAC) list as “RAHMAN, Khalil Abd ur.” Any person familiar with the Arabic language would know “RAHMAN” is not the correct last name. A predetermined reordering scheme can be applied so that the input name is modified to “Khalil Abd ur Rahman.” After reordering, a parser can be applied to the modified input name to output “Khalil Abdurrahman.” In another step (iii) “ Khalil Abdurrahman” can be broken into components; so “Khalil” becomesa given name while “Abdurrahman ” is the surname. Furthermore, names exist that have multi-surnames and the system 100 can be configured to be capable of recognizing them. For instance, the name “AL-ZAHRANI, Ahmed Abdullah Saleh al- Khazmari” has two potential surnames “AL-ZAHRANI” and “al-Khazmari.” The parsing technique (ii) can be applied to both the data sources (prescreened names) and user input (e.g., input data 114A). As such, a user engaging with the processor 102 (e.g., via the computing device 108) does not need to worry about using a certain format to write a query to the processor 102 to analyze an input name. The reordering scheme can include a set of predefined rules that can be updated as requested.

[0051] Parsing Logic Examples (for parsing 154)

[0052] Names of a subject, named subjects, or subject names as referenced herein generally define names of interest (e.g., a name to be compared against prescreened names of a financial institution watchlist for security name screening or for whatever reason) and can include entity names. As such, names can include names of one or more persons or individuals, organizations or company names, and the like and combinations thereof. Parsing logic can be configured for execution by the processor 102 in different ways to accommodate different types of names (e.g., name of an individual vs. name of a company). Parsing logic can be embodied as predefined functions, operations, and / or one or more algorithms.

[0053] In one example, the steps of parsing include accessing an input (e.g., from a user) including a request; by example a “PersonRequest” parameter that can include a first name, middle name, last name, and a full name depending upon how the user wishes to submit the PersonRequest instance, and include one or more initial operations, by example:• Step 1 - extraction: identifying / extracting the first names, last names, middle names and full names of one or more of a name of interest from the incoming requests (PersonRequest).• Step 2 - preprocessing: transliterating the name into a predetermined language.• Step 3 - parsing logic: applying one or more operations (title removal, composite name detection, etc.,) data parsing logics, etc.

[0054] Parsing logic is applied on the transliterated version of the incoming name of interest. Various other specific examples of parsing logic to accommodate different types of names follows below.

[0055] Individual name parsing logic:

[0056] There are several cases that fall under this type of matching where parsing logic can be configured for names of individual persons. For example, med Ahmed^l”. Regardless of the complexity of these cases, the first step in the matching process described herein is to parse the subject name into its standardized representation. This step can allow for matching the standard representation accurately in the matching step. It is also possible to configure each step based on a given users’ needs.

[0057] In one example, the parser starts by the omission or removal of titles, suchfrom the subject name. Then, any honorifics, such as or “Prince” can further be removed from the name. Next, common stop words, such as “LH” or “bin” can also be removed from the subject name. Once this step is completed, the processor 102 applies the parsing logic to start detecting composite names. In some examples, the parsing logic takes the form of a predefined algorithm (composite name detection algorithm) executed by the processor 102 to start detecting any composite names such asA composite name detection algorithm can include set of rules in the source code; i.e. , predetermined linguistic rules related to composite names that have been collected and can be applied.

[0058] This detection of composite names accommodates identification of parts of the names accurately; such as first names, middle names, and surnames. The last step of the parser logic is to generate name variants. It is not always the case that the inputs are ordered in the standard format of first-name, middle-name, then a last-name. The parser can also generate all permutations of name variants for transliteration, where each permutation can be used later in the matching process. Another variant that is common is the use of the last-name before the first-name with the omission of the middle name. These variants provide flexibility for different deployments of the system 100 based on an end user’s preferences or behavior.

[0059] Organization name parsing logic:

[0060] There are several cases that fall under this type of matching. For example, matching “Regardless of the complexity of these cases, parsing logic for an organization or company can define a two-step approach similar to the person / individual parsing logic approach above.

[0061] Before delving into the specifics of the parsing logic for organization names, it should be noted that certain cases cannot be modeled using conventional methods. For example, establishments may have popular but not official names, and non-conventional acronyms. Therefore, examples of the system 100 can permit users to add arbitrary aliases to be taken into consideration at the time of the matching in addition to the official name.

[0062] Similar to the person / individual name parsing logic, parsing logic for organization names starts by the omission of stop words, such as “company” orThese stop words are not limited to a fixed set and other stop words can be defined and flagged for removal. Once the stop words are removed, the input organization name is run through the composite name detection algorithm, which allows for grouping of phrases of multiple words such as “dw^l j ." In general, the composite name detection algorithm finds and concatenates words that belong to the same name component. So for example, the three words “Abd el Rahman” are concatenated into one: ’’Abdelrahman.”

[0063] If the inputs have any non-Latin script, it is then transliterated into English. Then, the standardized input can be split into the following components: 1) organization type, such as company or establishment. 2) organization commercial activity, such as hospital, or medical center. 3) organization explicit name. In some examples, with these three components, the parsing of the organization name or its alias is complete.

[0064] Phonetic Embedding of Transliterated and Parsed Names:

[0065] Once the parsing logic is applied to the name of interest, output of any parsing logic to the name results into parsing the name into one or more components, referenced as name components 206 in FIG. 2. For a person’s name, the components can include a first-name, middle-name, and last-name. For an organization name, the components can include organization-type, organizationname, and organization-commercial-activity. Each component of a requested name can be matched and / or compared with its corresponding counterpart as we matchthe name associated with the transaction against the name associated with the beneficiary. For example, the requested first name can be matched against the first name associated with the beneficiary, and the requested last name can be matched against the last name associated with the beneficiary.

[0066] Before comparison of input or subject names against watchlists (defining, e.g., prescreened names associated with fraud or a security threat), names can be represented in a numerical or vectorized format to later be queried or processed by computational models. This process is commonly referred to as text representation. In some examples, the processor 102 of the system can be configured to process the outputs of the transliteration and parsing processes above to construct phonetic embeddings. Specifically, the processor 102 can be configured to transliterate a name to a common predetermined language such as Latin script or English, apply parsing logic to the name as transliterated to generate one or more components of the name, and generate or construct phonetic embeddings from the one or more components of the name.

[0067] It is important to define the word embedding first. Although the naming is not intuitive, embeddings are mathematical representations of points in high dimensional spaces. Phonetic embedding is a technique that transforms names into numerical representations (vectors) capturing the pronunciation or phonetic properties of words; i.e., words that are pronounced similarly would have highly similar vectors regardless of their spelling. Many techniques may be applied to produce, e.g., neural network-based models that generate phonetic embeddings, for example:• Feature extraction, phonetic indexing techniques may be used to extract features based on the pronunciation of words, for example: o Soundex and its variations o Metaphone and Double Metaphone and their variations o Any combination of feature extraction based on phonetic indexing and their variations for non-English sounds• Embedding techniques, e.g., Word2Vec, Phrase2Vec and their variations

[0068] Name Look-Up Examples

[0069] A variety of different name look-up or query operations can be implemented to search, e.g., variations of a name of interest, XXXX

[0070] Example method for indexing vectors for efficient searching

[0071] As indicated herein, a database such as database 118 can be developed to include one or more datasets of prescreened names for comparison with a name of interest or subject name that needs to be compared, screened or possibly matched with a name in the database for whatever reason. Those who are familiar with databases, or old filing systems, would recognize that having an index separate from the data itself is an efficient setting for searching.

[0072] In one approach, referencing FIG. 5, an efficient method to index any type of name in the database 118 can be implemented. For example, let us assume that the name “Mohammed” is present in 1 ,000 entities in the database 118. In one example, an indexing method first constructs a vector for the name “Mohammed " and keeps references to the 1 ,0000 entities that are linked to it. Therefore, with a single scan through the database, redundant information is removed so as to only keep track of essential information. This design is motivated by the fact that pages on Wikipedia for humans are roughly about 7 million different pages. Yet, there are less than 700,000 unique names. Furthermore, this indexing scheme allows building search-trees on top of the vector representation resulting in a logarithmic time complexity for searching. Using a vector for the index also provides a means to efficiently find other similar vectors in a high dimensional vector space. For instance, given a vectorized name, the system can quickly find both the vectorized name and additional similar vectors for names nearby based on a mathematical representation of similarity, including the distance between vectors, inverse Euclidean distance, cosine similarity, or a neural network that outputs a value of similarity. For name matching, finding nearby or similar names is critical because, as discussed previously, there may be several variations in how the name is spelled and there is no need for an exact match. Furthermore, a vectorized index is numerical and can generalize across many languages. For example, a name in French and a name in English can both be embedded into the same vector space, thereby accommodating cross language matching and similarity searching. In contrast, simply using a name, such as “Mohammed”, for the index does not provide the advantage of being able to efficiently search for similar names or similar names in other languages. In some examples, one or more strings associated with a given name are embedded into a vector space to get a representation of their phonetic properties; and names represented in the vector space can be considered similar orin close proximity if they are phonetically similar (as opposed to. e.g., semantically similar).

[0073] Example method for retrieving results related to a search query including a name of interest

[0074] In some examples, as shown in FIG. 4, whenever the processor 102 of the system 100 receives or accesses a query defining a name of interest, a search can be executed on the index yielding names that are phonetically similar to the name of interest. Regardless of the position of the name (i.e., first, middle, or last) in the entity, it can be considered as a relevant entity to these names and retrieved for further evaluation.

[0075] Comparison of a Name of Interest with Prescreened Names (e.g., candidate entities):

[0076] For each component, the matching process can start by using multiple measures of similarity. In some examples each measure has its weight or in other words importance contributing to the final similarity score. However, this weighting is configurable to the desired behavior of our users. The measures can include the following:• The first measure is phonetic embedding (e.g., Phrase2Vec). The objective of this measure is to execute a model configured to transform the name components 206 (in Latin-scripts) into an abstract mathematical representation capturing the phonetic features of the script. This representation allows measurement of the similarity in terms of how things sound.• The second measure is to perform phonetic encodings of the name components 206 and measure how similar these encodings are to each other.• The third measure is widely used and known as “Edit Distance” where the similarity in terms of how characters are like each other is measured.• The fourth measure is an advancement over “Edit Distance” where the similarity using substrings instead of individual characters can be measured.

[0077] All these measures of similarity can then be fed into an aggregation function to yield a final similarity score for each measure output. In some examples, the aggregation function has no restriction on its internal design; however, a logistic function can be implemented to aggregate the output of the measures. The output of the aggregation function may be used as is or corrected in cases of high similarity values, which eventually can be interpreted as the likelihood of the inputted names to be the same.

[0078] Comparing Entities of Interest with Candidate Entities

[0079] After finding candidate name matches for the name of interest, candidate entities can be retrieved from a database using the candidate name matches. The candidate entities can include attributes in addition to name components, such as date of birth or nationality. The entity of interest may also have additional attributes. Additional measures of similarity for the entity of interest and each candidate entity can be generated based on the entity name or any of the entity attributes included in the entity of interest and the candidate entity. The function to yield a similarity score for each attribute has no restrictions and may be configured by a user of the system. For each candidate entity, the similarity scores for name and one or more attributes can be combined into an overall similarity score by an aggregation function to represent the overall similarity between the candidate entity and the entity of interest. For instance, the aggregation function may be configured to avoid false positives because incorrectly matching names is costly. The overall similarity scores for candidate entities may be used to rank order candidate entities. Similarity scores for the entity name, other entity attributes, or any combination thereof may also be used to rank order candidate entities. The method of rank ordering the candidate entities may be configured by the user.

[0080] Presentation of Results:

[0081] One notable feature is how results are shown to the user. As discussed above, there are several metrics or measures that can be used for scoring the matching score. However, it was observed that common names might affect the quality of the results remarkably. For example, consider the query “ JA'("Mohammed Anwar”) the nameknown to be the most common name, so giving it the same attention as “JA*” ("Anwar”) is not reasonable. Therefore, we can give the nameless weight than “JA*” because it distinguishes true positives from false positives.

[0082] Example:

[0083] This computation becomes purely probabilistic in nature:P Entity = True | fname = Mohammed" ,mname = "Anwar")

[0084] Assuming independence of the random variables fname and name, the equation can be refactored as:P (Entity = True | fname = "Mohammed”) P (Entity = True | mname = "Anwar")

[0085] The quantity P (Entity = True | fname = "Mohammed") will be less than the quantity P(Entity = True | mname = "Anwar").

[0086] Now those quantities can be used to determine which parts of the queries need more emphasis. However, we might face the problem of giving 0 attention to the first name. To solve this problem, we provide a regularization term, which guarantees no 0 weights. This regularization term can also be used to put more emphasis on the surname or given name. Those parameters aid refining the results further and can be customizable.

[0087] Example method for scoring a phonetic relevance of a retrieved entity as compared with a name of interest

[0088] A retrieved entity in the subject example includes an entity of the database 118 of prescreened entities that needs to be compared, phonetically, with a name of interest (e.g., name 204). One or more metrics, or combinations thereof, may be used to score the relevance of a retrieved entity including the use of plurality of:• Distance-based metrics, e.g., inverse Euclidean distance (similarity is the inverse of distance. When things are close, their similarity is high, but their distance is low), Manhattan distance, etc.• Similarity-based metrics, e.g., Pearson’s correlation, Spearmen’s correlation, Kendall’s Tau, Cosine similarity, Jaccard similarity, etc.• Any combination of these, e.g., weighted combination of these or rule-based scoring techniques

[0089] Example: To illustrate the subject scoring process, consider the following query exampleFirst name = “Mohammed”Middle name = “Abdullah”Last name = “Alahmad”

[0090] Also, consider this entity with the following aliases to be scored for comparison with the name of the querying example above.Table 4: Output scoring metrics including score defining a degree of similarity

[0091] Alias 1 and alias 3 match on two fields out of the three fields provided in the query, and those aliases can be used in the next stage of scoring. After determining the best matching alias, we can score the initials phonetically. Each initial is converted into a phonetic code allowing more than one character to map to the same initial. For example, the letters “k” and “q” are different, but they have the same phonetic features. After matching the initials, we score the similarity between the alias and the query using the edit-distance method. There is also a phonetic similarity computed between the query and the matched alias. In this case, we apply the edit distance algorithm to compare the phonetic keys.

[0092] General Screening Example

[0093] Given and incorporating features of the above description, a general non-limiting example of screening a name of interest is illustrated in process 1000 of FIG. 6. The steps illustrated in the blocks 1001-1005 of FIG. 6 can be implemented by the processor 102 or otherwise implemented.

[0094] In block 1001 of the process 1000, a name of interest can be accessed (by the processor 102) from a query and preprocessed. For example, a user associated with an end-user device may submit a request including a query to screen the name of interest to determine comparison between the name and prescreened names associated with known bad actors of a database.

[0095] In block 1002, the name can be transliterated to a predetermined language such as English. As indicated in the decision block 304 ofFIG. 3 for example, the processor 102 can be configured by instructions in memory accessible to the processor 102 to first determine whether the input name of interest is Arabic (AR) or not; and can thereafter apply appropriate preprocessing.

[0096] In block 1003, the name can be parsed into one or more name components. In general, the name components represent portions of the name in the predetermined language.In blocks 1004-1005, multiple measures of similarity can be computed to compare the name components of the name of interest with components associated with prescreened entities (other names). Similarity measures examples are described herein. In particular, strings representing the name components of the name of interest can be embedded into a vector space along with name components of the prescreened entities to define a representation of phonetic properties associated with all name components. In this manner, the components can be phonetically compared in the vector space to assess similarity or matching. For example, the processor 102 can compute, for each component of the plurality of name components, a phonetic similarity with the one or more prescreened components (associated with prescreened entities) using one or more measures of similarity, by: generating a phonetic embedding from phonetic features extracted from each component, the phonetic embedding defining a vector, and measuring a similarity in a vector space between the phonetic embedding of the component and the prescreened phonetic embeddings.

[0097] Application Examples:

[0098] Various example applications of the language-agnostic name screening functionality are contemplated. It should be understood that such examples are non-limiting and various other examples are included within the scope of the present inventive concept.

[0099] Application 1: Anti-Money Laundering (AML)

[0100] AML refers to the set of regulations, policies, and procedures businesses and financial institutions implement to prevent and detect fraudulent activities related to money laundering. All financial institutions that are under regulatory supervision must make complying with AML their top priority.

[0101] Financial penalties and legal action may be taken against institutions that do not comply with AML regulations. There has been a rise in theamount of fines imposed for non-compliance, with reports indicating that over $6 billion in fines and restitution payments were issued by the SEC and FCA to various trading and brokerage firms in 2022.

[0102] Maintaining compliance with AML regulations is a comprehensive process that spans various aspects, including people, technology, and policies and procedures. Technological advancements have vastly facilitated AML compliance.

[0103] One of the core technologies for AML compliance is sanction screening. Sanction screening is a process widely recognized as verifying the identities of subjects (entities, individuals, transactions, or any combination of these). This is done by comparing their names against watchlists using name matching techniques such as what is being proposed in this invention. Watchlists are made up of individuals and entities who are suspected of engaging in activities such as terrorism, money laundering, or breaking international laws. It is important to ensure that these subjects are identified and prevented from engaging in such activities.

[0104] Application 2: Confirmation of Payee (CoP)

[0105] CoP is a process that is most used by banks and financial institutions to verify the recipient's identity before making a payment, aiming to prevent misdirected payments and authorized push payment fraud. CoP addresses the risks of human error, typos, and incorrect account details, reducing financial losses and disputes. By adding an extra layer of security, CoP instills confidence in customers, promotes trust in digital payments, and demonstrates the industry's commitment to protecting customer interests.

[0106] All these applications share the need for accurate techniques for matching and / or screening names across different languages and with spelling variations. However, false positive rates continue to pose a significant challenge. Despite proposed matching algorithms such as exact string matching, fuzzy matching, or phonetic matching, existing literature is limited due to inherent challenges in transliteration, name parts recognition, and other matching techniques. The novel concepts described herein directly address these technical challenges.

[0107] Exemplary Computing Device: Referring to FIG. 7, a computing device 1200 is illustrated which may can be configured, via one or more of an application 1211 or computer-executable instructions, to execute functionality described herein. More particularly, in some embodiments, aspects of the methodsherein may be translated to software or machine-level code, which may be installed to and / or executed by the computing device 1200 such that the computing device 1200 is configured to execute functionality described herein. It is contemplated that the computing device 1200 may include any number of devices, such as personal computers, server computers, hand-held or laptop devices, tablet devices, multiprocessor systems, microprocessor-based systems, set top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, digital signal processors, state machines, logic circuitries, distributed computing environments, and the like.

[0108] The computing device 1200 may include various hardware components, such as a processor 1202, a main memory 1204 (e.g., a system memory), and a system bus 1201 that couples various components of the computing device 1200 to the processor 1202. The system bus 1201 may be any of several types of bus structures including a memory bus or memory controller, a peripheral bus, and a local bus using any of a variety of bus architectures. For example, such architectures may include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus also known as Mezzanine bus.

[0109] The computing device 1200 may further include a variety of memory devices and computer-readable media 1207 that includes removable / non- removable media and volatile / nonvolatile media and / or tangible media but excludes transitory propagated signals. Computer-readable media 1207 may also include computer storage media and communication media. Computer storage media includes removable / non-removable media and volatile / nonvolatile media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules or other data, such as RAM, ROM, EEPROM, flash memory or other memory technology, CD- ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that may be used to store the desired information / data and which may be accessed by the computing device 1200. Communication media includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism andincludes any information delivery media. The term "modulated data signal" means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. For example, communication media may include wired media such as a wired network or direct-wired connection and wireless media such as acoustic, RF, infrared, and / or other wireless media, or some combination thereof. Computer-readable media may be embodied as a computer program product, such as software stored on computer storage media.

[0110] The main memory 1204 includes computer storage media in the form of volatile / nonvolatile memory such as read only memory (ROM) and random access memory (RAM). A basic input / output system (BIOS), containing the basic routines that help to transfer information between elements within the computing device 1200 (e.g., during start-up) is typically stored in ROM. RAM typically contains data and / or program modules that are immediately accessible to and / or presently being operated on by processor 1202. Further, data storage 1206 in the form of Read-Only Memory (ROM) or otherwise may store an operating system, application programs, and other program modules and program data.

[0111] The data storage 1206 may also include other removable / non- removable, volatile / nonvolatile computer storage media. For example, the data storage 1206 may be: a hard disk drive that reads from or writes to non-removable, nonvolatile magnetic media; a magnetic disk drive that reads from or writes to a removable, nonvolatile magnetic disk; a solid state drive; and / or an optical disk drive that reads from or writes to a removable, nonvolatile optical disk such as a CD-ROM or other optical media. Other removable / non-removable, volatile / nonvolatile computer storage media may include magnetic tape cassettes, flash memory cards, digital versatile disks, digital video tape, solid state RAM, solid state ROM, and the like. The drives and their associated computer storage media provide storage of computer-readable instructions, data structures, program modules, and other data for the computing device 1200.

[0112] A user may enter commands and information through a user interface 1240 (displayed via a monitor 1260) by engaging input devices 1245 such as a tablet, electronic digitizer, a microphone, keyboard, and / or pointing device, commonly referred to as mouse, trackball or touch pad. Other input devices 1245 may include a joystick, game pad, satellite dish, scanner, or the like. Additionally, voice inputs, gesture inputs (e.g., via hands or fingers), or other natural user inputmethods may also be used with the appropriate input devices, such as a microphone, camera, tablet, touch pad, glove, or other sensor. These and other input devices 1245 are in operative connection to the processor 1202 and may be coupled to the system bus 1201 but may be connected by other interface and bus structures, such as a parallel port, game port or a universal serial bus (USB). The monitor 1260 or other type of display device may also be connected to the system bus 1201. The monitor 1260 may also be integrated with a touch-screen panel or the like.

[0113] The computing device 1200 may be implemented in a networked or cloud-computing environment using logical connections of a network interface 1203 to one or more remote devices, such as a remote computer. The remote computer may be a personal computer, a server, a router, a network PC, a peer device or other common network node, and typically includes many or all of the elements described above relative to the computing device 1200. The logical connection may include one or more local area networks (LAN) and one or more wide area networks (WAN) but may also include other networks. Such networking environments are commonplace in offices, enterprise-wide computer networks, intranets and the Internet.

[0114] When used in a networked or cloud-computing environment, the computing device 1200 may be connected to a public and / or private network through the network interface 1203. In such embodiments, a modem or other means for establishing communications over the network is connected to the system bus 1201 via the network interface 1203 or other appropriate mechanism. A wireless networking component including an interface and antenna may be coupled through a suitable device such as an access point or peer computer to a network. In a networked environment, program modules depicted relative to the computing device 1200, or portions thereof, may be stored in the remote memory storage device.

[0115] Certain embodiments are described herein as including one or more modules. Such modules are hardware-implemented, and thus include at least one tangible unit capable of performing certain operations and may be configured or arranged in a certain manner. For example, a hardware-implemented module may comprise dedicated circuitry that is permanently configured (e.g., as a specialpurpose processor, such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC)) to perform certain operations. Ahardware-implemented module may also comprise programmable circuitry (e.g., as encompassed within a general-purpose processor or other programmable processor) that is temporarily configured by software or firmware to perform certain operations. In some example embodiments, one or more computer systems (e.g., a standalone system, a client and / or server computer system, or a peer-to-peer computer system) or one or more processors may be configured by software (e.g., an application or application portion) as a hardware-implemented module that operates to perform certain operations as described herein.

[0116] Accordingly, the term “hardware-implemented module” encompasses a tangible entity, be that an entity that is physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a certain manner and / or to perform certain operations described herein. Considering embodiments in which hardware-implemented modules are temporarily configured (e.g., programmed), each of the hardware- implemented modules need not be configured or instantiated at any one instance in time. For example, where the hardware-implemented modules comprise a general- purpose processor configured using software, the general-purpose processor may be configured as respective different hardware-implemented modules at different times. Software may accordingly configure the processor 1202, for example, to constitute a particular hardware-implemented module at one instance of time and to constitute a different hardware-implemented module at a different instance of time.

[0117] Hardware-implemented modules may provide information to, and / or receive information from, other hardware-implemented modules. Accordingly, the described hardware-implemented modules may be regarded as being communicatively coupled. Where multiple of such hardware-implemented modules exist contemporaneously, communications may be achieved through signal transmission (e.g., over appropriate circuits and buses) that connect the hardware- implemented modules. In embodiments in which multiple hardware-implemented modules are configured or instantiated at different times, communications between such hardware-implemented modules may be achieved, for example, through the storage and retrieval of information in memory structures to which the multiple hardware-implemented modules have access. For example, one hardware- implemented module may perform an operation, and may store the output of that operation in a memory device to which it is communicatively coupled. A furtherhardware-implemented module may then, at a later time, access the memory device to retrieve and process the stored output. Hardware-implemented modules may also initiate communications with input or output devices.

[0118] Computing systems or devices referenced herein may include desktop computers, laptops, tablets e-readers, personal digital assistants, smartphones, gaming devices, servers, and the like. The computing devices may access computer-readable media that include computer-readable storage media and data transmission media. In some embodiments, the computer-readable storage media are tangible storage devices that do not include a transitory propagating signal. Examples include memory such as primary memory, cache memory, and secondary memory (e.g., DVD) and other storage devices. The computer-readable storage media may have instructions recorded on them or may be encoded with computer-executable instructions or logic that implements aspects of the functionality described herein. The data transmission media may be used for transmitting data via transitory, propagating signals or carrier waves (e.g., electromagnetism) via a wired or wireless connection.

[0119] The described methods, processes, operations, and associated actions may also be performed in various orders in addition to the order described in this application, in parallel, and / or simultaneously. The described systems are exemplary in nature and may include additional elements and / or omit elements. Furthermore, references to or “one example” of the present disclosure are not intended to be interpreted as excluding the existence of additional embodiments that also incorporate the recited features. It will be understood that when a certain part or process “includes” a certain component or operation, that part or process does not exclude another component or operation. While illustrative examples of the name screening techniques using phonetic embeddings have been described herein including systems, devices, and the like, it is to be understood that various other adaptations and modifications may be made within the spirit and the scope of the examples herein. Additionally, it is appreciated that while specific graphics are shown and described, such graphics are illustrative and exemplary and are not intended to limit the scope of this disclosure.

[0120] The foregoing description has been directed to specific examples. It will be apparent, however, that other variations and modifications may be made to the described examples, with the attainment of some or all of theiradvantages. For instance, it is expressly contemplated that the components and / or elements described herein can be implemented as software being stored on a tangible (non-transitory) computer-readable medium, devices, and memories (e.g., disks / CDs / RAM / EEPROM / etc.) having program instructions executing on a computer, hardware, firmware, or a combination thereof. Further, methods describing the various functions and techniques described herein can be implemented using computer-executable instructions that are stored or otherwise available from computer readable media. Such instructions can comprise, for example, instructions and data which cause or otherwise configure a general- purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions. Portions of computer resources used can be accessible over a network. The computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, firmware, or source code. Examples of computer-readable media that may be used to store instructions, information used, and / or information created during methods according to described examples include magnetic or optical disks, flash memory, USB devices provided with non-volatile memory, networked storage devices, and so on. In addition, devices implementing methods according to these disclosures can comprise hardware, firmware and / or software, and can take any of a variety of form factors. Typical examples of such form factors include laptops, smart phones, small form factor personal computers, personal digital assistants, and so on. Functionality described herein also can be embodied in peripherals or add-in cards. Such functionality can also be implemented on a circuit board among different chips or different processes executing in a single device, by way of further example. Instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are means for providing the functions described in these disclosures. Accordingly, this description is to be taken only by way of example and not to otherwise limit the scope of the examples herein. Therefore, it is the object of the appended claims to cover all such variations and modifications as come within the true spirit and scope of the examples herein.

[0121] In addition, the description of the disclosure is provided to enable a person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be readily apparent to those skilled in the art, andthe generic principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Throughout this disclosure the term “example” or “exemplary” indicates an example or instance and does not imply or require any preference for the noted example. Thus, the disclosure is not to be limited to the examples and designs described herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0122] Illustrative aspects of this disclosure include:

[0123] Statement 1. A method that includes operations for language agnostic-name screening. The method includes operations for: preprocessing a name of interest from a query to derive a standardized representation in a predetermined format (e.g., reorder and optionally remove / omit various elements of the name of interest); transliterating the name of interest in the standardized representation to a predetermined language; parsing the name of interest as transliterated to generate one or mor name components, the name components representing portions of the name in the predetermined language; computing multiple measures of similarity to compare each name component with components associated with prescreened entities of a database, the measures of similarity including phonetic comparison between components of the name of interest and components of prescreened entities of the database; and outputting a response to the query including an identification, decision, or recommendation for the name of interest based on the results of the multiple measures of similarity.

[0124] Statement 2. The method of statement 1 , further including generating at least one phonetic phrase based on respective groupings of characters in the direction of pronunciation for the name.

[0125] Statement 3. The method of any one of statements 1-2, further including generating an embedded phrase to represent the at least one phonetic phrase in a vector space.

[0126] Statement 4. The method any one of statements 1-3, further including accessing a plurality of embedded candidate phrases in the vector space based on a relative proximity to the at least one phonetic phrase, the plurality of embedded candidate phrases corresponding to prescreened entities.

[0127] Statement 5. The method any one of statements 1-4, further including computing a degree of similarity between the embedded phrase and an embedded candidate phrase associated with at least one prescreened name.

[0128] Statement 6. The method of any one of statements 1-5, wherein to transliterate the name for screening, the processor tokenizes the name for screening into a one or more N-grams in the first language, inputs the one or more N-grams in the first language into a deep learning model for transliteration of the one or more N-grams from the first language into one or more N-grams from the language, implements reverse tokenization of the one or more N-grams in the language into one or more words in the language

[0129] Statement 7. The method of any one of statements 1-6, further comprising: accessing an index comprising a plurality of embedded phrases in a vector space and an embedded phrase associated with the name for searching the index and selecting one or more entities referenced by the index based on the embedded phrase for searching the index.

[0130] Statement 8. The method of any one of statements 1-7, further comprising generating a plurality of variants of the name in the language; and supplementing the plurality of name components with the plurality of variants.

[0131] Statement 9. An apparatus comprising a processor configured to execute one or more processes, and memory configured to store a process executable by the processor. The process, when executed, is operable to perform operations according to any of statements 1-8.

[0132] Statement 10. A non-transitory, computer-readable medium storing instructions encoded thereon. The instructions, when executed by one or more processors, cause the one or more processors to perform operations according to any of statements 1-8.

[0133] Additional aspects of this disclosure are set out in the independent claims and preferred features are set out in the dependent claims. Features of one aspect may be applied to each aspect alone or in combination with other aspects. In addition, while certain operations in the claims are provided in a particular order, it is appreciated that such order is not required unless the context otherwise indicates.

[0134] It should be understood from the foregoing that, while particular embodiments have been illustrated and described, various modifications can be made thereto without departing from the spirit and scope of the invention as will be apparent to those skilled in the art. Such changes and modifications are within the scope and teachings of this invention as defined in the claims appended hereto.

Claims

CLAIMSWhat is claimed is:

1. A system for phonetic comparison between entities or names, comprising: a processor in communication with a memory, the memory including instructions, which, when executed, cause the processor to: access a name for screening; create a standardized name in a language based on the name, the language having a first alphabet of characters and a direction of pronunciation; generate at least one phonetic phrase based on respective groupings of characters in the direction of pronunciation for the standardized name; generate an embedded phrase to represent the at least one phonetic phrase in a vector space; access a plurality of embedded candidate phrases in the vector space based on a relative proximity to the at least one phonetic phrase in the vector space, the plurality of embedded candidate phrases corresponding to prescreened entities; and compute a degree of similarity between the embedded phrase associated with the name and at least one embedded candidate phrase of the plurality of embedded candidate phrases.

2. The system of claim 1, the memory further including instructions, which, when executed, cause the processor to: transliterate the name for screening from a first language to the language, wherein the processor: tokenizes the name for screening into a one or more N-grams in the first language, inputs the one or more N-grams in the first language into a deep learning model for transliteration of the one or more N-grams from the first language into one or more N-grams from the language,implements reverse tokenization of the one or more N-grams in the language into one or more words in the language.

3. The system of claim 1, the memory further including instructions, which, when executed, cause the processor to: generate the embedded phrase to represent the at least one phonetic phrase in the vector space by extraction of phonetic features based on pronunciation or phonetic properties of the at least one phonetic phrase.

4. The system of claim 2, the memory further including instructions, which, when executed, cause the processor to: create the standardized name in the first language by processing the name for screening, the processing of the name for screening including removal of titles and stop words and parsing of the name into individual components.

5. The system of claim 1, wherein the name for screening is associated with an entity defining an individual, an organization, or any combination thereof.

6. The system of claim 1, the memory further including instructions, which, when executed, cause the processor to: access an index comprising a plurality of embedded phrases in a vector space and an embedded phrase for searching the index; and select a one or more entities referenced by the index based on the embedded phrase for searching the index.

7. The system of claim 1 , wherein the relative proximity of the plurality of embedded candidate phrases in the vector space to the at least one phonetic phrase is based on a K-Nearest Neighbors algorithm.

8. A method of language-agnostic phonetic name comparison, comprising:standardizing a name into a script defining a standardized representation of the name associated with a language; parsing the script of the name into a plurality of name components in the language; accessing a plurality of prescreened names associated with entities including one or more prescreened components defining prescreened phonetic embeddings; and computing, for each component of the plurality of name components, a phonetic similarity with the one or more prescreened components using one or more measures of similarity, including: generating a phonetic embedding from phonetic features extracted from each component, the phonetic embedding defining a vector, and measuring a similarity in a vector space between the phonetic embedding of the component and the prescreened phonetic embeddings.

9. The method of claim 8, further comprising: generating a plurality of variants of the name in the language; and supplementing the plurality of name components with the plurality of variants.

10. The method of claim 8, further comprising: weighting one or more components of the plurality of name components to modify the phonetic similarity.

Citation Information

Patent Citations

  • System, method and computer program product for matching textual strings using language-biased normalisation, phonetic representation and correlation functions

    US20040024760A1

  • Method and system for associating data records in multiple languages

    US20090089332A1

  • Name indexing for name matching systems

    US20100153396A1

  • Parsing culturally diverse names

    US20120016660A1