Method, device and system for text string embedding and authentication
By generating an embedding lookup table and using graph model to train the embedding model, the problem of inaccurate text string authentication in eKYC is solved, and more accurate text string embedding and authentication is achieved.
Patent Information
- Application Number
- CN202110193206.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-12
- Filing Date
- 2021-02-20
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-02-20
AI Technical Summary
The prior art cannot accurately identify and verify the same or similar expressions in the text string in the electronic understanding customer (eKYC) process, resulting in inaccurate authentication.
By generating an embedding lookup table, using the graph model to split the text string into words and associate it with the device identifier, the embedding model is trained to construct text string embeddings containing semantic meaning and context information.
Improve the accuracy of text string authentication, can identify and verify text strings of different expression forms, and enhance the authentication capabilities in eKYC processing.
Smart Images

Figure CN112800412B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates broadly but not exclusively to methods, apparatuses, and systems for text string embedding and authentication. Background Art
[0002] Electronic Know Your Customer (eKYC) is a digital due diligence process carried out by enterprises to verify the identity of their customers and to assess the potential risk of illegal intent in a business relationship. In the eKYC process, an enterprise must perform authentication to verify the personal information of a user. The personal information of the user contains data in the form of text strings, such as user address, user name, etc.
[0003] The prior art uses word comparison techniques based on text distance metrics, such as edit distance, for text string authentication. However, word comparison techniques cannot capture the semantic meaning of words and / or phrases, and thus cannot identify the similarity between various expressions of the same or similar text strings. According to the word comparison techniques in the prior art, for example, the text string "Bishan, District 129, 04-01, Singapore 570129" is considered different from another text string "Bishan, Blk 129, 04-01, SG", although they refer to the same address. Therefore, subsequent authentication is inaccurate.
[0004] Word embedding techniques (such as word2vec) are capable of extracting semantic meaning from words or phrases. However, in the context of eKYC, text strings (such as addresses, names, etc.) are captured from a photo of an identity (ID) card by OCR technology or entered by a user, and do not contain the context information of the words in the text string (such as address words in an address, words in a name). Therefore, current word embedding techniques cannot produce satisfactory authentication results.
[0005] Therefore, there is a need to provide methods, apparatuses, and systems that can generate accurate embeddings for text strings to improve authentication accuracy in the eKYC process. Summary of the Invention
[0006] According to a first embodiment of the present disclosure, a method for text string embedding is provided. The method includes: accessing historical data stored including text strings and device identifiers, wherein each of the device identifiers is associated with one or more of the text strings; and generating an embedding lookup table including a plurality of keywords and a plurality of values, wherein generating the embedding lookup table includes: splitting the text strings in the historical data into a plurality of words; generating a graph by using the plurality of words and the device identifiers by: representing each of the plurality of words as a first type of node; representing each of the device identifiers as a second type of node; and for each of the second type of nodes, constructing edges by linking the second type of node representing one of the device identifiers with nodes in the first type of nodes, wherein the words represented by the linked nodes in the first type of nodes are included in one or more text strings associated with one of the device identifiers; using the generated graph to train an embedding model; and constructing the embedding lookup table based on the trained embedding model, wherein each of the plurality of keywords in the embedding lookup table includes one of the first type of nodes and the second type of nodes, and each of the plurality of values in the embedding lookup table includes a vector corresponding to one of the first type of nodes and the second type of nodes.
[0007] According to a second embodiment of the present disclosure, a method for text string authentication is provided. The method includes: receiving a first text string from a user; splitting the first text string into a plurality of first words; and authenticating the first text string by using one or more values in an embedding lookup table generated according to any one of the foregoing embodiments, the one or more values being associated with each of the plurality of first words in the embedding lookup table corresponding to the first text string.
[0008] According to a third embodiment of the present disclosure, a text string embedding device is provided. The text string embedding device includes: a first storage device configured to store historical data including text strings and device identifiers, wherein each of the device identifiers is associated with one or more of the text strings; a training device coupled to the first storage device and configured to generate an embedding lookup table including a plurality of keywords and a plurality of values, the training device being configured to: access the stored historical data; generate the embedding lookup table, wherein generating the embedding lookup table includes: splitting the text strings in the stored historical data into a plurality of words; generating a graph using the plurality of words and the device identifiers by: representing each of the plurality of words as a first type of node; representing each of the device identifiers as a second type of node; and for each of the second type of nodes, constructing an edge by linking the second type of node representing one of the device identifiers to a node in the first type of nodes, the words represented by the linked nodes in the first type of nodes being included in one or more text strings associated with one of the device identifiers; training an embedding model using the generated graph; and constructing the embedding lookup table based on the trained embedding model, wherein each of the plurality of keywords in the embedding lookup table includes one of the first type of nodes and the second type of nodes, and each of the plurality of values in the embedding lookup table includes a vector corresponding to one of the first type of nodes and the second type of nodes; and a second storage device configured to store the embedding query table.
[0009] According to a fourth embodiment of the present specification, a text string authentication device is provided, which is coupled to a text string embedding device according to any one of the foregoing embodiments, the text string embedding device including an embedding lookup table generated according to any one of the foregoing embodiments, wherein the text string authentication device is configured to: receive a first text string; split the first text string into a plurality of first words; and authenticate the first text string using one or more values in the embedding lookup table, the one or more values being associated with keywords corresponding to each of the plurality of first words in the embedding lookup table corresponding to the first text string. Description of the Drawings
[0010] The embodiments and implementation manners are provided by way of example only. For those of ordinary skill in the art, the embodiments and implementation manners will be better understood and more readily apparent from the following written description and in conjunction with the accompanying drawings, wherein:
[0011] Figure 1 is a schematic diagram of a system for text string authentication according to an embodiment. As depicted in this embodiment, the system includes a text string embedding device and a text string authentication device coupled to the text string embedding device.
[0012] Figure 2Schematic diagram of a computing device according to an embodiment. As described herein, the computing device can be implemented as a text string embedding device, a text string authentication device, or other components in a system for text string authentication.
[0013] Figure 3 Flowchart showing steps in a text string embedding method according to an embodiment.
[0014] Figure 4 Flowchart showing steps in a text string authentication method according to an embodiment.
[0015] Figure 5 Shows historical data including a text string and a device identifier according to an embodiment. In this embodiment, the text string is an address. As shown in this embodiment, each address is associated with one of the device identifiers. It is also shown in this embodiment that each address is split into multiple address words, and each of the multiple address words is associated with one of the device identifiers.
[0016] Figure 6 Shows the use of Figure 5 Example of a graph generated using multiple address words of an address and device identifiers as shown.
[0017] Figure 7 Shows an example of an embedding lookup table. In this embodiment, the embedding lookup table is constructed based on an embedding model trained using a graph as shown in Figure 6 Shown.
[0018] Figure 8 Shows a block diagram of a computer system suitable for performing at least some steps of the text string embedding method or the text string authentication method described herein. According to the Figures 1 to 7 Embodiment shown, the computer system is also suitable for use as a text string embedding device, a text string authentication device, or a system for text string authentication.
[0019] Those skilled in the art will understand that the elements in the figures are shown for simplicity and clarity and are not necessarily drawn to scale. For example, the dimensions of some elements in the illustration, block diagram, or flowchart may be exaggerated relative to other elements to help improve the understanding of this embodiment. Detailed Description
[0020] The embodiments will be described by way of example only, with reference to the accompanying drawings. Like reference numerals and symbols in the drawings represent the same elements or equivalents.
[0021] Some of the portions described below are presented explicitly or implicitly in terms of algorithms and functions or symbolic representations of operations on data within a computer memory. These algorithmic descriptions and function or symbolic representations are the means by which those skilled in the data processing arts most effectively convey the substance of their work to others skilled in the art. Here, an algorithm is generally considered to be a self-consistent sequence of steps leading to a desired result. These steps are those requiring physical manipulation of physical quantities (such as electrical, magnetic, or optical signals that can be stored, transferred, combined, compared, and otherwise manipulated).
[0022] Unless otherwise explicitly stated and as will be apparent from the following, it will be understood that throughout the specification, discussions using terms such as "access", "generate", "split", "represent", "construct", "link", "train", "configure", "calculate", "apply", "predict", "generate", "execute", "receive", "split", "extract", "authenticate", etc., refer to the actions and processes of a computer system or similar electronic device that manipulate and transform data represented as physical quantities within the computer system into other data similarly represented as physical quantities within the computer system or other information storage, transmission, or display device.
[0023] Also disclosed herein are apparatuses for performing the operations of the described methods. Such apparatuses may be specially constructed for the required purposes or may comprise a computer or other devices selectively activated or reconfigured by a computer program stored in the computer. The algorithms and displays presented herein have no inherent connection to any particular computer or other apparatus. According to the teachings herein, various machines may be used with the program. Alternatively, the construction of more specialized apparatuses for performing the required method steps may be appropriate. The structure of a computer suitable for performing the various methods / processes described herein will be presented by the following description. In this document, the terms "apparatus" and "device" may be used interchangeably.
[0024] Furthermore, a computer program is implicitly disclosed herein, since it will be apparent to those skilled in the art that the various steps of the methods described herein can be implemented by computer code. The computer program is not intended to be limited to any particular programming language and its implementation. It will be understood that various programming languages and their codes can be used to implement the teachings of the present disclosure contained herein. Moreover, the computer program is not intended to be limited to any particular control flow. There are many other variations of the computer program that can use different control flows without departing from the scope hereof.
[0025] In addition, one or more steps of a computer program can be executed in parallel rather than sequentially. Such a computer program can be stored on any computer-readable medium. The computer-readable medium can include storage devices such as magnetic or optical disks, memory chips, or other storage devices suitable for interfacing with a computer. The computer-readable medium can also include hardwired media as exemplified in Internet systems, or wireless media as exemplified in GSM mobile phone systems. When the computer program is loaded and executed on such a computer, it effectively produces an apparatus for implementing the steps of any method described herein.
[0026] In this document, it has been observed that when two or more addresses are provided to the service provider's eKYC system using the same device, the two or more addresses are likely to be the same, similar, or related addresses because device sharing typically occurs among user groups sharing the same or related spaces, such as family members living in the same unit, neighbors living in the same apartment building or nearby apartment buildings, classmates sharing a classroom or campus building, etc. In this way, if address A and address A' are provided by the same device, the address terms in address A can be used to provide semantic meaning and / or context information for the address terms in address A', and vice versa.
[0027] In view of the above, embodiments of this document provide methods, devices, and systems for associating the device identifier of a user device with address terms, so that semantic meaning and context information can be used for address terms to facilitate the generation of accurate embeddings of addresses and improve the accuracy of authentication in eKYC processing.
[0028] In addition, the present methods, devices, and systems generate a graph with the device identifier and associated address terms as nodes and the links between the nodes as edges. In this way, the semantic meaning and context information of the address terms contained in the nodes and edges are learned through an embedding model to form an address embedding. The address embedding can be used for authentication as described herein. Those skilled in the art will understand that the address embedding can also be used for similar address searches, address comparisons, etc.
[0029] In addition to addresses, the above observations in this document also apply to other text strings, such as user names, etc.
[0030] As a supplement or alternative to the device identifier of the user device, other types of user personal information, such as the user's phone number, the user's account at the service provider, etc., can also be used to associate with the words in the address terms or other text strings to provide context information for the words in the address terms or other text strings.
[0031] In view of the above, the present method, device, and system can also generate a graph with device identifiers, the user's telephone number and / or the user's account, and associated words in the address word or other text strings as nodes, and the links between the nodes as edges. In this way, the embedding model is used to learn the semantic meanings and context information of the words (e.g., address words or words in other text strings) contained in the text strings in the nodes and edges to form an embedding of the text strings. The embedding of the text strings can be used for authentication of text strings (e.g., addresses, names, etc.) as described herein or for quantification of the similarity of text strings.
[0032] Figure 1 FIG. 4 shows a schematic diagram of a system 100 for text string authentication according to an embodiment herein. The system 100 may be part of an eKYC system ( Figure 1 not shown) of a service provider. Those skilled in the art can also understand that the system 100 can be implemented as an eKYC system without substantial modification.
[0033] As described in this embodiment, the system 100 includes a text string embedding device 102 and a text string authentication device 112 coupled to the text string embedding device 102.
[0034] As Figure 1 shown, the embedding device 102 includes a training device 106 coupled to a first storage device 104. The first storage device 104 stores historical data including text strings such as user addresses, user names, etc. The historical data also includes data associated with the text strings, such as device identifiers, the user's telephone number, and the user's account at the service provider. In some embodiments, the device identifier may be generated by the system 100 when receiving an address from the user device previously. In this way, each address in the historical data is associated with the device identifier corresponding to the device that provided the address. The device identifier may have a format such as "Device 1", "Device 2", etc. Those skilled in the art will understand that the device identifier may have other formats.
[0035] Based on the serial number of the user device, different devices can be identified and different device identifiers can be assigned to them. Similarly, a device that provides multiple addresses at different time points can be identified as the same device and the same device identifier can be assigned to it. In this way, one or more addresses provided by the same device can be associated with the same device identifier in the historical data.
[0036] As an alternative to the system 100, the device identifier may be generated by other components in the eKYC system ( Figure 1 not shown) and sent to the first storage device 104 for storage. For simplicity, the generation of the device identifier is not further discussed in detail herein.
[0037] In some embodiments, when a user provides an address from a user device to a service provider, the user's name is also provided along with the address. The name is linked to the address in the historical data and can also receive the text string embeddings described herein.
[0038] In some embodiments, as a supplement or alternative to the device identifier corresponding to the user device, the user's phone number and / or the user's account at the service provider can also be associated with the address and / or the name to provide context information for the words in these text strings.
[0039] The text string embedding device 102 is configured to perform a text string embedding method. Figure 3 An embodiment 300 of the text string embedding method is shown. As Figure 3 shown, the text string embedding device 102 is configured to at least perform the following steps in the text string embedding method 300. Each step of the method can be performed sequentially, or in parallel or in any order when applicable.
[0040] Method 300 includes: step 302, accessing historical data stored including text strings and device identifiers, wherein each device identifier is associated with one or more text strings. In some embodiments, the text strings include addresses. In these embodiments, each device identifier is associated with one or more addresses.
[0041] Method 300 further includes: step 304, generating an embedding lookup table including a plurality of keywords and a plurality of values. The step 304 of generating the embedding lookup table includes: step 304a, splitting the text strings in the historical data into a plurality of words; step 304b, generating a graph using the plurality of words and the device identifiers; step 304c, using the generated graph to train an embedding model; step 304d, constructing the embedding lookup table based on the trained embedding model, wherein each keyword in the plurality of keywords of the embedding lookup table includes one of a first type of node and a second type of node, and each value in the plurality of values of the embedding lookup table includes a vector corresponding to the one of the first type of node and the second type of node.
[0042] The step 304b of generating a graph using the plurality of words and the device identifiers includes: representing each word in the plurality of words as a first type of node; representing each device identifier as a second type of node; for each second type of node, constructing an edge by linking the second type of node representing one of the device identifiers to a node in the first type of nodes, and the word represented by the linked node in the first type of nodes is included in one or more text strings associated with the one of the device identifiers.
[0043] In some embodiments, the text string includes an address, and the plurality of words includes a plurality of address words. In these embodiments, in step 304a, the text string embedding device 102 is configured to split the address in the historical data into a plurality of address words. In step 304b, the text string embedding device 102 is configured to generate a graph by using the plurality of address words and the device identifier as follows: represent each address word in the plurality of address words as a first type of node; represent each device identifier as a second type of node; for each second type of node, construct an edge by linking the second type of node representing one of the device identifiers to a node in the first type of nodes, and the address word represented by the linked node in the first type of nodes is included in one or more addresses associated with one of the device identifiers.
[0044] In some embodiments, the text string includes a name, and the plurality of words includes a plurality of name words. In these embodiments, in step 304a, the text string embedding device 102 is configured to split the name in the historical data into a plurality of name words. In step 304b, the text string embedding device 102 is configured to represent each name word in the plurality of name words as a first type of node; represent each device identifier as a second type of node; for each second type of node, construct an edge by linking the second type of node representing one of the device identifiers to a node in the first type of nodes, and the name word represented by the linked node in the first type of nodes is included in one or more names associated with one of the device identifiers.
[0045] In some embodiments, the text string includes an address and a name, and the plurality of words includes a plurality of address words and name words. In these embodiments, in step 304a, the text string embedding device 102 is configured to split the address and the name in the historical data into a plurality of address words and name words. In step 304b, the text string embedding device 102 is configured to generate a graph by using the plurality of address words, name words, and device identifier as follows: represent each address word or name word in the plurality of address words and name words as a first type of node; represent each device identifier as a second type of node; for each second type of node, construct an edge by linking the second type of node representing one of the device identifiers to a node in the first type of nodes, and the address word represented by the linked node in the first type of nodes is included in one or more addresses associated with one of the device identifiers, or the name word represented by the linked node in the first type of nodes is included in one or more names associated with one of the device identifiers.
[0046] In some embodiments, as a supplement to or an alternative for the device identifier, the user's phone number and / or the user account at the service provider can also be associated with the address and / or name in the text string to provide context information for the words in the text string. In these embodiments, at step 304b, the text string embedding device 102 is configured to generate a graph by using multiple address words, name words, and the device identifier as follows: representing each address word or name word among the multiple address words and name words as a first type of node; representing each device identifier, the user's phone number, and / or the user account as a second type of node; for each second type of node, constructing an edge by linking the second type of node representing one of the device identifier, the user's phone number, and / or the user account to a node in the first type of nodes, where the address word represented by the linked node in the first type of nodes is included in one or more addresses associated with one of the device identifier, the user's phone number, and / or the user account, or the name word represented is included in one or more names associated with one of the device identifier, the user's phone number, and / or the user account.
[0047] Refer to the following Figures 5 to 7 to describe steps 302 and 304. Figures 5 to 7 An embodiment is described where the text string only includes an address. That is, an embodiment where the text string is an address. In some other embodiments, the text string can also include a name, etc.
[0048] In some embodiments of step 302, the training device 106 in the text string embedding device 102 is configured to access historical data stored in the first storage device 104 of the text string embedding device.
[0049] The historical data includes the text string and the device identifier as described above. In Figure 5 an embodiment of the historical data 500 stored in the first storage device 104 is described. In this embodiment, the text string only includes an address. In this embodiment, an exemplary part of the historical data 500 includes: address #1 "Bishan District 129 Singapore 570129" and its associated device identifier "device 1"; address #2 "Bishan District 131 06 - 02 Singapore 570131" and its associated device identifier "device 2"; address #3 "Bishan Blk 129 SG" and its associated device identifier "device 1".
[0050] In as Figure 5In the historical data shown, each device identifier is associated with one or more addresses. For example, in embodiment 500 of the historical data, the device identifier "Device 2" is associated with address #2, while the device identifier "Device 1" is associated with address #1 and address #3. It can be seen that address #1 and address #3 are associated with the shared device represented by the device identifier "Device 1". Such information, that is, the information of multiple addresses associated with the shared device, is advantageously used herein to construct context information and provide semantic meaning for the address terms in the addresses. For example, since address #1 and address #3 are associated with the same device, the addresses "Bishan District 129 Singapore 570129" and "Bishan Blk129SG" are very likely to be the same or related because they can be provided by family members living in the same unit, neighbors living in the same apartment building or nearby apartment buildings, classmates sharing the same classroom or campus building, etc. In this way, the semantic meaning and context information of one address are presented by another address associated with the same device identifier.
[0051] In step 304, the training device 106 in the text string embedding device 102 is configured to generate a text string embedding lookup table for text string embedding based on the addresses and device identifiers in the historical data. The embedding lookup table includes a plurality of keywords and a plurality of values. Step 304 of generating the embedding lookup table includes steps 304a, 304b, 304c, and 304d described above and below.
[0052] In step 304a, the training device 106 in the text string embedding device 102 is configured to split the addresses in the historical data 500 into a plurality of address terms.
[0053] Refer to Figure 5In an embodiment of the historical data 500, the training device 106 splits the addresses #1, #2, and #3 in the historical data 500 into multiple address words. The address words include: words, abbreviations of words, numbers, etc. For example, the address "Bishan District 129 Singapore 570129" is split into address words: "Bishan", "District", "129", "Singapore", "570129"; the address "Bishan District 131 06-02 Singapore 570131" is split into address words: "Bishan", "District", "131", "06-02", "Singapore", "570131"; the address "Bishan Blk 129 SG" is split into address words: "Bishan", "Blk", "129", "SG". In this way, multiple address words "Bishan", "District", "129", "Singapore", "570129", "131", "06-02", "Blk", "129", "SG" are separated from the addresses. In some embodiments, for simplicity, duplicate entries of the same address word can be deleted from the multiple address words. It can be seen that the address words "Singapore", "Bishan", and "District" appear in multiple addresses. That is, the address word "Singapore" is included in two addresses (i.e., address #1 and address #2); the address word "Bishan" is included in all three addresses (i.e., address #1, address #2, and address #3); the address word "District" is included in two addresses (i.e., address #1 and address #2).
[0054] Since the address is associated with the corresponding device identifier in the historical data 500, the address words separated from the address must also be associated with the corresponding device identifier. For example, each address word "Bishan", "District", "129", "Singapore", and "570129" separated from the address "Bishan District 129 Singapore 570129" is associated with the device identifier "Device 1". Similarly, each address word "Bishan", "District", "131", "06-02", "Singapore", and "570131" separated from the address "Bishan District 131 06-02 Singapore 570131" is associated with the device identifier "Device 2", and each address word "Bishan", "Blk", "129", and "SG" separated from the address "Bishan Blk 129 SG" is associated with the device identifier "Device 1".
[0055] For address words (such as "Singapore", "Bishan", and "District") included in multiple addresses (such as address #1 and address #2) of an address, since the multiple addresses are associated with the corresponding device identifiers (for example, "Address 1" is associated with "Device 1" and "Address 2" is associated with "Device 2"), these address words must also be respectively associated with the corresponding multiple device identifiers. For example, each address word "Singapore", "Bishan", and "District" is associated with two device identifiers "Device 1" and "Device 2".
[0056] As described above, the address terms "Bishan", "District", "129", "Singapore", and "570129" separated from address #1, and the address terms "Bishan", "Blk", "129", and "SG" separated from address 3 associated with the same device identifier "Device 1" are considered to have the same or related semantic and geographical meanings. If such a relationship is learned, the semantic meaning and context information of the address terms can be advantageously provided. This article saves such a relationship together with the addresses and device identifiers in the graph generated in step 304b to achieve an overall and accurate text string embedding.
[0057] In step 304b, the training device 106 in the text string embedding device 102 is configured to generate a graph using multiple address terms of the address (such as those split in step 304a) and the device identifier associated with the multiple address terms. Since the address terms and the device identifier have different attributes, they are captured as different types of nodes in the graph.
[0058] An embodiment of such a graph 600 is shown in Figure 6 . Referring to Figure 6 , the steps of generating the graph 600 in step 304b include the following steps:
[0059] First, the training device 106 in the text string embedding device 102 represents each address term in the multiple address terms of each address as a first type of node. As Figure 6 shown, each address term "Bishan", "District", "129", "Singapore", "570129" in the address "Bishan District 129 Singapore 570129" is represented as a first type of node. Similarly, each address term "Bishan", "District", "131", "06-02", "Singapore", "570131" in the address "Bishan District 131 06-02 Singapore 570131" is represented as a first type of node; each address term "Bishan", "Blk", "129", "SG" in the address "Bishan Blk 129 SG" is also represented as a first type of node.
[0060] As described above, in some embodiments, an address term may appear multiple times in multiple addresses. Such an address term is captured only once in the graph 600 and is represented as a single first type of node. For example, the address term "Singapore" included in the two addresses "Bishan District 129 Singapore 570129" and "Bishan District 131 06-02 Singapore 570131" is represented as a single first type of node in the graph 600. As Figure 6As shown, Figure 600 includes 10 nodes of the first type, namely, node 602 representing "Singapore", node 604 representing "SG", node 606 representing "Bishan", node 608 representing "Blk", node 610 representing "Area", node 612 representing "129", node 614 representing "570129", node 616 representing "06-02", node 618 representing "131", and node 620 representing "570131".
[0061] Secondly, the training device 106 in the text string embedding device 102 represents each device identifier as a node of the second type. As Figure 6 shown, Figure 600 includes 2 nodes of the second type, namely, node 622 representing the device identifier "Device 1" and node 624 representing the device identifier "Device 2".
[0062] As described above, each of the address words "Singapore", "Bishan", and "Area" included in the multiple addresses is associated with a corresponding device identifier. In this regard, among the above 10 nodes of the first type, node 602 representing "Singapore" is associated with two nodes of the second type, namely, node 622 representing the device identifier "Device 1" and node 624 representing the device identifier "Device 2". Similarly, node 606 representing "Bishan" and node 610 representing "Area" are also both associated with node 622 representing the device identifier "Device 1" and node 624 representing the device identifier "Device 2".
[0063] Thirdly, for each node of the second type, the training device 106 in the text string embedding device 102 links the node of the second type representing one of the device identifiers with a node in the nodes of the first type, and the address word represented by the linked node in the nodes of the first type is included in one or more addresses associated with one of the device identifiers. As Figure 6As shown, for the second - type node 622 representing the device identifier "Device 1", the training device 106 constructs edges 626, 628, 630, 632, 634, 636, 638 by linking the second - type node 622 with the first - type nodes 602, 604, 606, 608, 610, 612, 614. The address words "Singapore", "SG", "Bishan", "Blk", "Area", "129", "570129" represented by the first - type nodes 602, 604, 606, 608, 610, 612, 614 are included in one or more addresses (i.e., Address #1 and Address #2) associated with the device identifier "Device 1" represented by the node 622. Similarly, for the second - type node 624 representing the device identifier "Device 2", the training device 106 constructs edges 640, 642, 646, 648, 650, 652 by linking the second - type node 624 with the first - type nodes 602, 606, 610, 616, 618, 620. The address words "Singapore", "Bishan", "Area", "06 - 02", "131", "570131" represented by the first - type nodes 602, 606, 610, 616, 618, 620 are included in one or more addresses (i.e., Address #2) associated with the device identifier "Device 2" represented by the node 624.
[0064] The graph 600 generated in the above - mentioned manner includes the address words represented by the first - type nodes, the device identifiers represented by the second - type nodes, the relationships between the device identifiers and the address words represented by the edges, and the relationships between the address words of one address and the address words of another address linked by the shared device identifiers. As described above, the relationships between the device identifiers and the address words represented by the edges and the relationships between the address words of one address and the address words of another address linked by the shared device identifiers are conducive to constructing semantic meanings and context information for the address words in this article.
[0065] In step 304c, the training device 106 in the text string embedding device 102 uses the generated graph 600 to train the embedding model.
[0066] In some embodiments, step 304c of using the generated graph 600 to train the embedding model includes the following sub - steps.
[0067] First, in sub-step a, for each node 602, 604, 606, 608, 610, 612, 614, 616, 618, 620, 622, 624 in the generated graph 600, the training device 106 can calculate a vector based on information of the neighbors of the node, where the neighbors of the node include one or more nodes linked to the node within a predetermined number of edges. The predetermined number of edges can be determined based on the accuracy requirement for text string authentication or other downstream applications that utilize text string embeddings. For example, the predetermined number of edges can be 5.
[0068] In some embodiments, the device identifier can include geographical information of the corresponding device as an attribute. The geographical information can include GPS information of the corresponding device. Such geographical information can be captured by the system 100 or other components in the eKYC system. In these embodiments, the information of the neighbors of the node includes geographical information corresponding to one or more device identifiers "Device 1", "Device 2" represented by one or more second-type nodes 622, 624 among the neighbors of the node.
[0069] In some embodiments, when the training device 106 calculates the vector of a first-type node representing an address word with an occurrence frequency lower than a threshold, the training device 106 can calculate the vector of the node as the vector of an unknown word. For example, the threshold can be 5% or a predetermined value based on actual needs. The calculation and training of the vector of an unknown word are easily understandable to those skilled in the art.
[0070] Subsequently, in sub-step b, the training device 106 can apply random walk on all nodes 602, 604, 606, 608, 610, 612, 614, 616, 618, 620, 622, 624 in the generated graph 600 to generate a node sequence.
[0071] Thereafter, in sub-step c, for each node within a node sequence in the generated node sequence, the training device 106 can predict the neighbors of the nodes in the node sequence to form a predicted node distribution.
[0072] Subsequently, in sub-step d, the training device 106 can train the embedding model based on the predicted node distribution and the true node distribution. The true node distribution can be obtained from the actual node sequence.
[0073] The above sub-steps a, c, and d of step 304c can be iterated for all nodes 602, 604, 606, 608, 610, 612, 614, 616, 618, 620, 622, 624 in the generated graph 600 to train the embedding model of all address words in the generated graph 600.
[0074] In an embodiment of the iterative process, the training device 106 updates the vector of each node with information of the neighbors of the node. Subsequently, the training device 106 predicts the neighbors of the nodes in the node sequence. The neighbors of the nodes in the node sequence are different from the neighbors of the nodes in FIG. 600. Thereafter, the training device 106 trains the embedding model based on the predicted node distribution and the true node distribution to update the vector of each node.
[0075] In some embodiments, the graph 600 generated in step 304c can be trained by a heterogeneous neural network because the address words and device identifiers in the graph 600 are regarded as two types of nodes. In this regard, the embedding model can be a General Attribute Multivariate Heterogeneous Network Embedding (GATNE) model, a Hierarchical Attention Network (HAN) model, or a Heterogeneous Graph Neural Network (HetGNN) model. Those skilled in the art can understand that other graph neural networks can also be used to train the embedding model herein.
[0076] In step 304d, the training device 106 in the text string embedding device 102 constructs an embedding lookup table based on the trained embedding model. Each keyword in the multiple keywords of the embedding lookup table includes one node from the first type of nodes and the second type of nodes. Each value in the multiple values of the embedding lookup table includes a vector corresponding to one node from the first type of nodes and the second type of nodes. As described above, the vector corresponding to one node from the first type of nodes and the second type of nodes is trained in step 304c.
[0077] An embodiment 700 of the embedding lookup table is shown in Figure 7 Embedding lookup table 700 includes multiple keywords and multiple values. As Figure 7 shown, each keyword in the multiple keywords of the embedding lookup table 700 includes one node from the first type of nodes representing the address words "Singapore", "SG", "Bishan", "Blk", "Area", "129", "131", "570129", "570131", "06-02" and the second type of nodes representing the device identifiers "Device 1", "Device 2".
[0078] As Figure 7 shown, each value in the multiple values of the embedding lookup table 700 includes a vector corresponding to one node from the first type of nodes and the second type of nodes. In some embodiments, as described above, the vector is trained in step 304c.
[0079] The embedded lookup table 700 may also include keywords with address words having a low occurrence frequency. The low occurrence frequency may be an occurrence frequency lower than a threshold. Based on actual needs, the threshold may be 5% or a predetermined value. Keywords with address words having a low occurrence frequency may be regarded as keywords for unknown words. Keywords for unknown words have corresponding vectors, the values of which are, for example, "vector_UNK". In step 304c, the vector "vector_UNK" is trained together with other vectors. Based on actual requirements in different scenarios, the vector "vector_UNK" may be equal to 0 or various other values.
[0080] Compared with the prior art, the embedded lookup table generated in the above embodiments of the method provides a more accurate text string embedding.
[0081] As Figure 1 shown, the embedded lookup table generated by the training device 106 may be sent to the second storage device 108 for storage. The second storage device 108 is coupled to the training device 106 and the text string authentication device 112, and is used to authenticate new text strings (such as new addresses, new names, etc.) in downstream applications.
[0082] In Figure 1 the shown embodiment, the first storage device 104 and the second storage device 108 are included in the text string embedding device 102. Those skilled in the art can understand that in some other embodiments, the first storage device 104 and the second storage device 108 may be external hardware components coupled to the text string embedding device 102. In addition, in some alternative embodiments, the first storage device 104 and the second storage device 108 may be implemented by a single storage device, or be implemented as hardware components included in the text string embedding device 102, or be implemented as external hardware components communicable with the text string embedding device 102.
[0083] As Figure 1 shown, the system 100 may also include an input device 110 for receiving new text strings. The input device 110 is coupled to the first storage device 104 for storing new text strings. The input device 110 is also coupled to the text string authentication device 112. The text string authentication device 112 is coupled to the second storage device 108, and the second storage device 108 stores the embedded lookup table.
[0084] The text string authentication device 112 is configured to execute a text string authentication method. Figure 4 An embodiment 400 of the text string authentication method is shown. As Figure 4 shown, the text string authentication device 112 is configured to at least execute the following steps in the text string authentication method 400. Each step of the method may be executed sequentially; or in parallel or in any order when applicable.
[0085] Method 400 includes: step 402 of receiving a first text string from a user.
[0086] Method 400 further includes: step 404 of splitting the first text string into a plurality of first words.
[0087] Method 400 further includes: step 406 of authenticating the first text string using one or more values in a generated embedding lookup table as described herein, the one or more values in the embedding lookup table being associated with keywords corresponding to each word of the plurality of first words of the first text string in the embedding lookup table.
[0088] In some embodiments, in step 402, the text string authentication device 112 is configured to receive a first text string 116 sent from an input device 110 coupled to the text string authentication device 112. In an embodiment, the first text string 116 is a first address 116. The first address 116 may be provided by the user.
[0089] In some embodiments, in step 404, the text string authentication device 112 is configured to split the first address 116 into a plurality of first address words.
[0090] In some embodiments, in step 406, the text string authentication device 112 is configured to authenticate the first address 116 using one or more values in an embedding lookup table 700, the one or more values in the embedding lookup table being associated with keywords corresponding to each address word of the plurality of first address words of the first address 116 in the embedding lookup table 700.
[0091] In some embodiments, step 406 of authenticating the first address 116 further includes the following sub-steps. Each sub-step of step 406 may be executed sequentially; or in parallel or in any order when applicable.
[0092] Step 406 includes: sub-step 406a of extracting one or more values in the embedding lookup table, the one or more values being associated with keywords corresponding to each address word of the plurality of first address words of the first address 116, wherein any unknown word among the plurality of first address words of the first address 116 is mapped to a keyword of the unknown word, and the keyword of the unknown word has an unknown word vector as its value.
[0093] Step 406 further includes: sub-step 406b of generating a vector of the first address 116 by performing a summation calculation or an average calculation on the values of the extracted plurality of first address words.
[0094] In some embodiments, step 406 of authenticating the first address 116 further includes the following sub-steps. Each sub-step of step 406 may be executed sequentially; or in parallel or in any order when applicable.
[0095] Step 406 further includes: sub-step 406c, receiving a second text string 118 of the user from an official source.
[0096] Step 406 further includes: sub-step 406d, splitting the second text string 118 into a plurality of second words.
[0097] Step 406 further includes: sub-step 406e, extracting one or more values from an embedding lookup table, where the one or more values are associated with keywords corresponding to each word in the plurality of second words of the second text string 118 in the embedding lookup table, and any unknown word in the plurality of second words of the second text string 118 is mapped to a keyword of the unknown word, and the keyword of the unknown word has an unknown word vector as its value.
[0098] Step 406 further includes: sub-step 406f, generating a vector of the second text string 118 by performing a summation calculation or an average calculation on the values of the extracted plurality of second words.
[0099] In some embodiments, the second text string 118 is a second address 118. The plurality of second words are a plurality of second address words.
[0100] In some embodiments, step 406 of authenticating the first address 116 further includes: sub-step 406g, authenticating the first address 116 based on the similarity between the vector of the first address 116 and the vector of the second address 118. In some embodiments, in step 406g, the similarity between the vector of the first address 116 and the vector of the second address 118 can be calculated based on cosine similarity. In some examples, if the similarity is greater than a threshold, the first address 116 can be authenticated as a real address. According to actual requirements, the threshold can be 0.85 or other predetermined values.
[0101] In some embodiments, the official source mentioned in sub-step 406c includes a government database or a third-party database authorized by the government, which provides official personal information of the user, including identity document (ID) number, name, address, etc. The second address received from the official source can be obtained based on the user ID number provided to the system 100 in the eKYC process.
[0102] The above paragraphs describe embodiments of steps 402, 404, and 406 in the case where the text string is an address. Based on these embodiments, it can be understood by those skilled in the art that steps 402, 404, and 406 are similar in the case where the text string is a name.
[0103] Through the text string embedding implemented herein, more accurate results can be achieved for text string authentication between a text string provided by a user (e.g., an address) and an official text string of the user from an official source (e.g., an address).
[0104] In addition, through the text string embedding implemented in this article, the above text string authentication method can be modified to: a method for identifying potential fraudsters by calculating the similarity between the address vector of a known fraudster and another address vector of a suspected fraudster based on the text string embedding lookup table constructed in this article. Those skilled in the art can understand that the text string embedding can also be used in downstream applications, such as similar address searches, address comparisons, etc.
[0105] As Figure 1 shown, the system 100 may further include an output device 114 coupled to the text string authentication device 112. The result of the text string authentication can be sent from the text string authentication device 112 to the output device 114 for display. Alternatively or additionally, the result of the text string authentication can be sent from the text string authentication device 112 to a transmitter for sending to the user's device.
[0106] Figure 2 A schematic diagram of the device 200 is shown. As described in this article, the device 200 can be implemented as the text string embedding device 102, the text string authentication device 112, or other components for text string authentication in the system 100.
[0107] The device 200 includes at least a processor 202 and a memory 204. The processor 202 and the memory 204 are interconnected. The memory 204 includes computer program code ( Figure 2 not shown in the figure). The memory 204 and the computer program code are configured to, together with the processor 202, cause the device 200 to perform the steps for text string embedding or text string authentication as described in this article.
[0108] Figure 8 A block diagram of a computer system 800 suitable for use as the text string embedding device 102, the text string authentication device 112, or the system 100 for text string authentication is shown. The following description of the computer system / computing device 800 is provided only by way of example and is not intended to be limiting.
[0109] As Figure 8 shown, the exemplary computing device 800 includes a processor 804 for executing software routines. Although a single processor is shown for clarity, the computing device 800 may also include a multiprocessor system. The processor 804 is connected to a communication facility 806 for communicating with other components of the computing device 800. The communication facility 806 may include, for example, a communication bus, a cross-bar, or a network.
[0110] The computing device 800 also includes a main memory 808 such as random access memory (RAM) and an auxiliary memory 810. The auxiliary memory 810 may include, for example, a hard disk drive 812 and / or a removable storage drive 814, and the removable storage drive 814 may include a magnetic tape drive, an optical disk drive, etc. The removable storage drive 814 reads from and / or writes to a removable storage unit 818 in a well-known manner. The removable storage unit 818 may include magnetic tapes, optical disks, etc. that are read from and written to by the removable storage drive 814. As will be understood by those skilled in the relevant art, the removable storage unit 818 includes a computer-readable storage medium in which computer-executable program code instructions and / or data are stored.
[0111] In an alternative embodiment, the auxiliary memory 810 may additionally or alternatively include other similar devices for allowing a computer program or other instructions to be loaded into the computing device 800. Such devices may include, for example, a removable storage unit 822 and an interface 820. Examples of the removable storage unit 822 and the interface 820 include a removable storage chip (e.g., EPROM or PROM) and an associated slot, and other removable storage units 822 and interfaces 820 that allow software and data to be transferred from the removable storage unit 822 to the computer system 800.
[0112] The computing device 800 also includes at least one communication interface 824. The communication interface 824 allows software and data to be transferred between the computing device 800 and external devices via a communication path 826. In various embodiments, the communication interface 824 allows data to be transferred between the computing device 800 and a data communication network such as a public data or private data communication network. The communication interface 824 can be used to exchange data between different computing devices 800 that form part of an interconnected computer network. Examples of the communication interface 824 may include a modem, a network interface (such as an Ethernet card), a communication port, an antenna with associated circuitry, etc. The communication interface 824 can be wired or can be wireless. The software and data transferred through the communication interface 824 are in the form of signals, which can be electrical signals, electromagnetic signals, optical signals, or other signals that can be received by the communication interface 824. These signals are provided to the communication interface via the communication path 826.
[0113] Optionally, the computing device 800 further includes a display interface 802 and an audio interface 832. The display interface 802 performs operations for providing images to an associated display 830, and the audio interface 832 performs operations for playing audio content via an associated speaker 834.
[0114] As used herein, the term "computer program product" can refer in part to a removable storage unit 818, a removable storage unit 822, a hard disk installed in the hard disk drive 812, or a carrier wave that transports software via a communication path 826 (wireless link or cable) to the communication interface 824. A computer-readable storage medium refers to any non-transitory tangible storage medium that provides recorded instructions and / or data to a computing device 800 for execution and / or processing. Examples of such storage media include floppy disks, magnetic tapes, CD-ROMs, DVDs, Blu-ray TM optical discs, hard disk drives, ROMs or integrated circuits, USB memories, magneto-optical discs, or computer-readable cards such as PCMCIA cards, whether these devices are internal or external to the computing device 800. Examples of transient or non-tangible computer-readable transmission media that can also participate in providing software, applications, instructions, and / or data to the computing device 800 include radio or infrared transmission channels, as well as a network connected to another computer or network device, and the Internet or an intranet including email transmissions and information recorded on a website, etc.
[0115] The computer program (also referred to as computer program code) is stored in the main memory 808 and / or the auxiliary memory 810. The computer program can also be received via the communication interface 824. Such a computer program, when executed, enables the computing device 800 to perform one or more features of the embodiments discussed herein. In various embodiments, the computer program, when executed, enables the processor 804 to perform the features of the above-described embodiments. Thus, such a computer program represents the controller of the computer system 800.
[0116] The software can be stored in a computer program product and loaded into the computing device 800 using a removable storage drive 814, a hard disk drive 812, or an interface 820. Alternatively, the computer program product can be downloaded to the computer system 800 via the communication path 826. The software, when executed by the processor 804, enables the computing device 800 to perform the functions of the embodiments described herein.
[0117] It should be understood that Figure 8 the embodiments of
[0118] The techniques described herein produce one or more technical effects. In particular, the present disclosure advantageously utilizes information of multiple addresses associated with a shared device to construct context information and provide semantic meanings for address terms in the addresses. Using this information, multiple address terms associated with the same device identifier and additional multiple address terms are considered to have the same or related semantic and geographical meanings. In addition to addresses, the above observations in this article also apply to other text strings, such as user names, etc.
[0119] As a supplement or alternative to the device identifier of the user device, other types of user personal information (e.g., the user's phone number, the user's account at the service provider, etc.) can also be used to be associated with the terms in the address terms or other text strings to provide context information for the terms in the address terms or other text strings.
[0120] The present article preserves the above relationships in a graph to train an embedding model, thereby achieving an overall and accurate text string embedding. Then, this text string embedding is used for the text string authentication described herein. Those skilled in the art will understand that the text string embedding can also be used in downstream applications, such as similar address searches, address comparisons, etc. Compared with traditional text comparisons or word embeddings based on text strings and text string authentication, the current devices, methods, and systems improve the accuracy of text string embedding, and thus improve the accuracy of text string authentication in eKYC processing.
[0121] Those skilled in the art will understand that various changes and / or modifications can be made to the present disclosure shown in the specific embodiments without departing from the scope of the present disclosure as broadly described herein. Therefore, the present embodiment should be considered illustrative in all respects and not restrictive.
Claims
1. A method for text string embedding, comprising: Accessing historical data stored including text strings and device identifiers, wherein each of the device identifiers is associated with one or more of the text strings; and Generating an embedding lookup table including a plurality of keywords and a plurality of values, wherein generating the embedding lookup table includes: Splitting the text strings in the historical data into a plurality of words; Generating a graph by using the plurality of words and the device identifiers as follows: Representing each of the plurality of words as a first type of node; Representing each of the device identifiers as a second type of node; and For each of the second type of nodes, constructing edges by linking the second type of node representing one of the device identifiers to nodes in the first type of nodes, wherein the words represented by the linked nodes in the first type of nodes are included in one or more text strings associated with one of the device identifiers; Using the generated graph to train an embedding model; and Constructing the embedding lookup table based on the trained embedding model, wherein each keyword in the plurality of keywords of the embedding lookup table includes a node in the first type of nodes and the second type of nodes, and each value in the plurality of values of the embedding lookup table includes a vector corresponding to the one node in the first type of nodes and the second type of nodes; Wherein using the generated graph to train the embedding model includes: For each node in the generated graph, calculating the vector based on information of the neighbors of the node, wherein the neighbors of the node include one or more nodes linked to the node within a predetermined number of edges; Applying random walk on all nodes in the generated graph to generate a node sequence; For each node in a node sequence within the generated node sequence, predicting the neighbors of the node in the node sequence to form a predicted node distribution; and Training the embedding model based on the predicted node distribution and the true node distribution.
2. The method according to claim 1, wherein, The text string includes an address, and the plurality of words include a plurality of address words.
3. The method according to claim 2, wherein When the neighbors of the node include one or more of the second type of nodes, the information of the neighbors of the node includes geographical information corresponding to one or more of the device identifiers represented by the one or more second type of nodes.
4. The method according to claim 1, wherein Calculating the vector for each node in the generated graph further includes: If the node is a first type of node representing a word with an occurrence frequency lower than a threshold, calculating the vector of the node as an unknown word vector.
5. The method according to any one of claims 1 to 4, wherein The embedding model includes a Generalized Attribute Multivariate Heterogeneous Network Embedding (GATNE) model, a Hierarchical Attention Network (HAN) model, or a Heterogeneous Graph Neural Network (HetGNN) model.
6. A method for text string authentication, comprising: Receiving a first text string from a user; Splitting the first text string into a plurality of first words; And Authenticating the first text string by using one or more values in an embedding lookup table generated according to any one of claims 1 to 5, the one or more values being associated with each of the plurality of first words in the embedding lookup table corresponding to the first text string.
7. The method according to claim 6, wherein, Authenticating the first text string includes: Extract one or more values from the embedded lookup table, where the one or more values in the embedded lookup table are associated with keywords corresponding to each word in the multiple first words of the first text string in the embedded lookup table, and any unknown word in the multiple first words of the first text string is mapped to a keyword of an unknown word with an unknown word vector as the value; and Generate a vector of the first text string by performing a summation calculation or an average calculation on the values of the extracted multiple first words.
8. The method according to claim 7, wherein Authenticating the first text string further includes: Receiving a second text string of the user from an official source; Splitting the second text string into multiple second words; Extract one or more values from the embedded lookup table, where the one or more values in the embedded lookup table are associated with keywords corresponding to each word in the multiple second words of the second text string in the embedded lookup table, and any unknown word in the multiple second words of the second text string is mapped to a keyword of an unknown word with an unknown word vector as the value; and Generate a vector of the second text string by performing a summation calculation or an average calculation on the values of the extracted multiple second words.
9. The method according to claim 8, further comprising: Authenticating the first text string based on the similarity between the vector of the first text string and the vector of the second text string.
10. A text string embedding device, comprising: A first storage device for storing historical data including text strings and device identifiers, where each of the device identifiers is associated with one or more of the text strings; A training device coupled to the first storage device for generating an embedded lookup table including multiple keywords and multiple values, the training device being configured to: Access the stored historical data; Generate the embedded lookup table, where generating the embedded lookup table includes: Splitting the text strings in the stored historical data into multiple words; Generating a graph by using the multiple words and the device identifiers through: Representing each word in the multiple words as a first type of node; Representing each of the device identifiers as a second type of node; and For each of the second type of nodes, constructing an edge by linking the second type of node representing one of the device identifiers to a node in the first type of nodes, where the words represented by the linked nodes in the first type of nodes are included in one or more text strings associated with one of the device identifiers; Using the generated graph to train an embedding model; and Constructing the embedded lookup table based on the trained embedding model, where each keyword in the multiple keywords of the embedded lookup table includes a node in the first type of nodes and the second type of nodes, and each value in the multiple values of the embedded lookup table includes a vector corresponding to the one node in the first type of nodes and the second type of nodes; and A second storage device for storing the embedded lookup table; Where, when training the embedding model using the generated graph, the training device is configured to: For each node in the generated graph, calculate the vector based on the information of the node's neighbors, where the neighbors of the node include one or more nodes linked to the node within a predetermined number of edges; Apply random walks on all nodes in the generated graph to generate a node sequence; For each node in a node sequence within the generated node sequence, predict the neighbors of the node in the node sequence to form a predicted node distribution; and Train the embedding model based on the predicted node distribution and the true node distribution.
11. The device according to claim 10, wherein, The text string includes an address, and the plurality of words include a plurality of address words.
12. The apparatus according to claim 11, wherein, When the neighbors of the node include one or more of the second type of nodes, the information of the neighbors of the node includes geographical information corresponding to one or more of the device identifiers represented by the one or more second type of nodes.
13. The device according to claim 10, wherein, When calculating the vector of each node in the generated graph, the training device is further configured to: If the node is a first type of node representing a word with an occurrence frequency lower than a threshold, calculate the vector of the node as an unknown word vector.
14. The apparatus according to any one of claims 10 to 13, wherein, The embedding model includes a Generalized Attribute Multivariate Heterogeneous Network Embedding (GATNE) model, a Hierarchical Attention Network (HAN) model, or a Heterogeneous Graph Neural Network (HetGNN) model.
15. A text string authentication device is coupled to the text string embedding device according to any one of claims 10 to 14, and the text string embedding device includes an embedding lookup table generated according to any one of claims 1 to 5, wherein, The text string authentication device is configured to: Receive a first text string; Split the first text string into a plurality of first words; And Authenticate the first text string using one or more values in the embedding lookup table, where the one or more values in the embedding lookup table are associated with keywords corresponding to each of the plurality of first words in the first text string in the embedding lookup table.
16. The device according to claim 15, wherein, The text string authentication device is configured to: Extract one or more values in the embedding lookup table, where the one or more values in the embedding lookup table are associated with keywords corresponding to each of the plurality of first words in the first text string in the embedding lookup table, and any unknown word in the plurality of first words of the first text string is mapped to a keyword of an unknown word with an unknown word vector as the value; and Generate a vector of the first text string by performing a summation calculation or an average calculation on the extracted values of the plurality of first words.
17. The apparatus according to claim 16, wherein, The text string authentication device is further configured to: Receive a second text string of the user from an official source; Split the second text string into a plurality of second words; Extract one or more values in the embedding lookup table, where the one or more values in the embedding lookup table are associated with keywords corresponding to each of the plurality of second words in the second text string in the embedding lookup table, and any unknown word in the plurality of second words of the second text string is mapped to a keyword of an unknown word with an unknown word vector as the value; And Generate a vector of the second text string by performing a summation calculation or an average calculation on the extracted values of the plurality of second words.
18. The device according to claim 17, wherein, The text string authentication device is further configured to: Authenticate the first text string based on the similarity between the vector of the first text string and the vector of the second text string.
Citation Information
Patent Citations
Natural language processing-based multi-language analysis method and device
CN108197109A
Text recommendation method and system based on heterogeneous topic model and word embedding model
CN110851714A