System for classifying words in batch of words

By training and adjusting the neural network, vector pairs of word-based dictionary data sets are realized, and the transmitted words are anonymous and classification without storing the original words, solving the problem of safe monitoring of personal information categories and improving the accuracy and efficiency of information classification.

CN120373302APending Publication Date: 2025-07-25COLLIBRA NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510383728.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-04-10
Filing Date
2020-10-16
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art is difficult to securely anonymize and classify transmitted words in the context of protecting personal information, especially when the original words have not been saved, and it is difficult to effectively monitor information categories.

Method used

Using neural network training and retraining, the distance between vector pairs of word-based dictionary data sets is approximate to the editing distance between corresponding words. Convolutional neural networks and twin convolutional neural networks are used to create vectors for words, and the network is adjusted through similarity metrics and cosine similarity to realize the anonymization and classification of words.

Benefits of technology

It realizes the safe anonymization and classification of transmitted words without storing original words, and improves the accuracy and efficiency of monitoring transmission information categories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373302A_ABST
    Figure CN120373302A_ABST
Patent Text Reader

Abstract

Systems, methods, and media are disclosed. The method includes obtaining a first and a second plurality of words and a first plurality of classifications, a word of the first plurality of words being classified into a classification of the first plurality of classifications; obtaining a first and second plurality of vectors representing the first and second plurality of words; obtaining, for each vector of the second plurality of vectors, a subset of matching vectors of the first plurality of vectors, and one or more scores associated with the subset of matching vectors; classifying the second plurality of words by obtaining a subset of matching classifications associated with the subset of matching vectors and one or more scores associated with the subset of matching vectors; performing, for each classification in the subset of matching classifications: averaging the subset of scores to obtain an averaged score; and reporting at least a portion of the subset of the matching classification and the average score to indicate to the user if the sensitive data is included in the second plurality of words without disclosing the second plurality of words.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of a patent application for invention with the application date of October 16, 2020, application number 202080102064.6 (International Application Number PCT / IB2020 / 059737), and invention title "System for Classifying Words in a Batch of Words".

[0002] Cross - reference to related applications

[0003] This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 008,552, filed on April 10, 2020, the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0004] This patent application relates to text similarity calculation, and more particularly, to word edit distance embedding. Background Art

[0005] Enterprises that process personal information (whether it is employee information, medical data, or financial information) need to protect this information and restrict its use. In addition, any data collected should be anonymized and stored only when necessary. Governments have enacted regulations, such as the European Union's General Data Protection Regulation (GDPR), to help protect personal information. Brief Description of the Drawings

[0006] The systems and methods described herein can be better understood by reference to the following detailed description in conjunction with the accompanying drawings, in which like reference numerals represent the same or functionally similar elements:

[0007] Figure 1 is a flowchart of an operating method of a processor - based word edit distance embedding system according to some embodiments of the present technology;

[0008] Figure 2 is a flowchart showing a method for training a neural network for classifying words according to some embodiments of the present technology;

[0009] Figure 3 shows a two - dimensional matrix for encoding words according to some embodiments of the present technology;

[0010] Figure 4 is a schematic diagram of a training architecture with a siamese convolutional neural network according to some embodiments of the present technology;

[0011] Figure 5 is a flowchart showing a method for setting up a word vector index for classifying words according to some embodiments of the present technology;

[0012] Figure 6 is a diagram showing dictionary word embeddings and associated categories and thresholds;

[0013] Figure 7 is a flowchart showing a method for classifying words according to some embodiments of the present technology;

[0014] Figure 8 is a block diagram showing an overview of a device on which some embodiments can operate;

[0015] Figure 9 is a block diagram showing an overview of an environment in which some embodiments can operate; and

[0016] Figure 10 is a block diagram showing components that can be used in a system employing the disclosed technology in some embodiments.

[0017] The headings provided herein are for convenience only and do not necessarily affect the scope of the embodiments. Additionally, the drawings are not necessarily to scale. For example, the dimensions of some elements in the drawings may be enlarged or reduced to help improve understanding of the embodiments. Further, while the disclosed technology is amenable to various modifications and alternative forms, specific embodiments have been shown by way of example in the drawings and are described in detail below. However, it is not intended to unduly limit the described embodiments. Rather, the embodiments are intended to cover all modifications, combinations, equivalents, and alternatives falling within the scope of the present disclosure. Detailed Description

[0018] Examples of the above systems and methods will now be described in further detail. The following description provides specific details in order to provide a thorough understanding and implementation of these example descriptions. However, those skilled in the relevant art will understand that the processes and techniques discussed herein can be practiced without many of these details. Similarly, those skilled in the relevant art will also understand that the present technology may include many other features not described in detail herein. Additionally, some well-known structures or functions may not be shown or described in detail below to avoid unnecessarily obscuring the relevant description.

[0019] The terms used below should be interpreted in the broadest reasonable manner, even if it is used in conjunction with a detailed description of some specific examples of the embodiments. In fact, some terms may even be emphasized below; however, any term intended to be interpreted in any restricted manner will be clearly and specifically defined in this section.

[0020] Disclosed herein are methods and systems for classifying unknown words using novel neural network and similarity search architectures. In the context of protecting personal information, it is necessary to monitor whether personal information is being transmitted within and / or outside a network. The disclosed techniques enable such transmissions to be monitored by securely anonymizing the words in transit and comparing them to words in a variety of category dictionaries (e.g., countries, names, ethnicities, languages, religions, genders, marital status, etc.) to determine the category of information being transmitted. Advantageously, the disclosed techniques securely anonymize and classify words without ever storing the original words.

[0021] In some embodiments, a neural network is trained and retrained such that the distance between vector pairs based on a dictionary dataset of words approximates an edit-distance-based metric between the corresponding words. Training includes using the trained neural network to create vectors for all dictionary words and storing the vectors in a particular structure. A matching algorithm then creates vectors for all words in a batch, including finding the closest neighboring matches for each word and applying a score based on the distance to the closest dictionary match to determine the classification of each word.

[0022] Figure 1 is a high-level flowchart of an operational method 100 of a processor-based word edit distance embedding system for classifying multiple batches of data (e.g., words) according to some embodiments of the present technology. The method may include training a neural network (NN) with multiple pairs of words at step 102 such that the distance between corresponding vectors approximates an edit-distance-based metric between those words. At step 104, selected vectors for dictionary words are created using the trained NN. The resulting vectors may be stored at step 106. Next, vectors are created for all words in a batch to be classified at step 108. The closest match for each word in the batch is determined at step 110. In some embodiments, at step 112, a score may be applied to each word in the batch based on the distance to the associated closest match.

[0023] Figure 2 illustrates a method 200 for training a neural network for classifying words according to some embodiments of the present technology. In some embodiments, the neural network may be a convolutional neural network (CNN). For example, training the CNN may include encoding each of a plurality of training words in matrix form ( Figure 3 ). The training words may include a dictionary of words belonging to a particular category (e.g., countries, names, ethnicities, languages, religions, genders, marital status, etc.). The training dictionary may include a base dictionary, a complete English dictionary, and a non-word dictionary whose contained data is not similar to the type of data in any similar dictionary.

[0024] At step 204, use a CNN and its siamese to create pairs of training vectors for multiple pairs of encoded training words ( Figure 4 ). Pairs can be randomly selected from the training dictionary for training against large word edit distances, and pairs can be created by adding noise to words for training against small word edit distances.

[0025] At step 206, compute a similarity metric (SM) for each pair of the multiple pairs of training words of the multiple training words. The SM can be computed based on the edit distance (ED) (e.g., Levenshtein ED) as follows:

[0026] SM(word1,word2) = 1 - ED(word1,word2) / max(length(word1),length(word2))

[0027] In some embodiments, the SM can be based on the encoded form of the words.

[0028] At step 208, compare the similarity metric and the cosine similarity for each pair of training words and adjust the CNN (e.g., adjust the weights) based on the comparison to drive the cosine similarity between the pairs to match the similarity metric. Using cosine similarity instead of Euclidean distance between vectors can improve computation time and accuracy. Once the CNN is trained, store it at step 210. In some embodiments, the cosine similarity can be computed by dividing the dot product of the vector pair by the product of their Euclidean norms.

[0029] Using the similarity metric mentioned above instead of the traditional Levenshtein ED improves the search. This can be illustrated with two examples: two 2-letter codes - PL and PS, and two words - Poland and Roland. The ED for both pairs = 1. But the value 1 has a much greater impact on the first pair because it changes half of the word. Thus, when the error in the edit distance is approximately "1", the accuracy using ED for short words is much lower compared to long words. In contrast, the disclosed similarity metric loses precision on longer pairs (especially those with large edit distances), which is acceptable in the disclosed system because it is searching for words within some range, and if it is too large, the exact distance value is not that important. In some embodiments, the minimum similarity threshold for accepting a match is 0.7.

[0030] Figure 3 A two-dimensional matrix 220 for encoding a word (e.g., "Poland") according to some embodiments of the present technology is shown. The disclosed two-dimensional matrix is similar to the matrix described below: Lluís Gómez, and Dimosthenis Karatzas, "LSDE: Levenshtein Space Deep Embedding for Query-by-string Word Spotting", Proceedings of the 2017 International Conference on Document Analysis and Recognition (ICDAR), which is hereby incorporated by reference in its entirety. Matrix 220 includes a first dimension 222 and a second dimension 224. The first dimension 222 includes 26 alphabetic characters and four special characters, and the second dimension 224 includes the character positions in the string. The number of character positions can be limited to, for example, 50 positions, thus generating a 30×50 matrix. It should be noted that matrix 220 can encode alphabetic and non-alphabetic characters, as well as the "empty" character described below.

[0031] Each character in the dataset can be grouped as follows:

[0032] 1. Alphabetic characters - Alphabetic characters are normalized (e.g., é = e), resulting in 26 basic alphabetic characters.

[0033] 2. Digits - Create a "bag" for all digit characters, represented by "#".

[0034] 3. Separating characters - Create a "bag" for all separating characters, represented by "-" (e.g., '', '-', '_', '.', ',', ':', ' / '). However, if any of these characters appears in the dataset more than a threshold minimum number of times, it is treated as an independent character.

[0035] 4. "Empty" positions, represented by "*".

[0036] 5. Characters that appear in the dataset less than the threshold minimum number of times are excluded and represented by "?".

[0037] Figure 4 is a schematic diagram of a training architecture 300 with a Siamese convolutional neural network according to some embodiments of the present technology. The Siamese CNNs 302(1) and 302(2) have the same structure and the same weights. The input to the Siamese CNNs is the same 30×50 matrix 304, as described above with respect to Figure 3Those described. In some embodiments, the CNN application has a convolutional layer 308 with 64 kernels 306 of size 30×3 and a fully connected layer 310 with 128 output neurons. The cosine similarity 312 of each pair of vectors corresponding to word1 and word2 is calculated and compared with the similarity metric between the encoded word1 and word2, resulting in a "loss" 316. The CNN is driven to learn a better transformation of words (e.g., strings) into vector form such that the cosine similarity distance between the vectors equals the similarity metric of their original words. In other words, the loss 316 is approximated to zero.

[0038] Figure 5 is a flowchart showing a method 400 for setting up a word vector index for classifying words according to some embodiments of the present technology. Once the CNN is trained as described above with reference Figures 2 - 4 In step 402, a plurality of dictionary words are encoded into matrix form, such as the matrix described above with reference to Figure 3 In some embodiments, the dictionary words used to train the CNN may include dictionary words for classification. In step 404, a dictionary vector for each of the plurality of encoded dictionary words is created using the trained CNN. In step 406, the resulting dictionary vectors and the categories corresponding to the associated dictionary words are stored. In step 408, the dictionary word vectors can be indexed for efficient search. In some embodiments, the vectors are indexed according to the Facebook AI Similarity Search (FAISS) technique. Figure 6 shows an example data structure 420 of dictionary word vector embeddings 426 and associated categories 428 and thresholds 430. For example, a country dictionary 422 includes the names of a plurality of countries, and a name dictionary 424 includes a plurality of given names. Each individual country and given name corresponds to a vector embedding 426. The category code 428 for countries is "0", and the category code 428 for names is "1". The data structure 420 includes a mapping 430 of category codes and category thresholds to each vector embedding. In some embodiments, the thresholds for each category can be determined experimentally to minimize false positives.

[0039] Figure 7 is a flowchart showing a method 500 for classifying words according to some embodiments of the present technology. Once the CNN is trained and the dictionary word index is ready, words in a batch of words of interest can be classified to determine whether the words in the batch fall into any sensitive data categories, such as personal information. In step 502, a batch of words for classification is received but not stored. In step 504, for example, using the above reference Figure 3The described technique encodes the received words into a matrix form. At step 506, a word vector for each of the multiple encoded batches of words is created using a trained CNN. At step 508, the most matching item and score for each resulting word vector in the batch are found in the index of the dictionary word vectors by searching for the closest neighbors using, for example, the FAISS library. FAISS is a library for efficient similarity search and dense vector clustering. It contains algorithms for searching a set of vectors of any size. At step 510, the average results for each category corresponding to the batch of words are reported.

[0040] Step 510 can be illustrated with an example batch of two words: Holand and Cuba. For each word, the four most matching items and the corresponding scores are obtained as follows:

[0041] Holand - Holand, Poland, Roland, Holland

[0042] Cuba - Cuba, Coby, Curt, Aruba

[0043] Each matching item is associated with a category and the score obtained, as follows:

[0044] Holand - Holand (surname - 1.0), Poland (country - 0.8), Roland (first name - 0.8), Holland (country - 0.8)

[0045] Cuba - Cuba (country - 1.0), Coby (first name - 0.5), Curt (first name - 0.5), Aruba (country - 0.6)

[0046] For each word in the batch, the highest score for each category is selected, as follows:

[0047] Holand - surname: 1.0, country: 0.8, first name: 0.8

[0048] Cuba - country: 1.0, first name: 0.5

[0049] Given the following category thresholds:

[0050] country: 0.8

[0051] surname: 0.8

[0052] first name: 0.7

[0053] The above thresholds are applied to the highest score for each category respectively. If the score is greater than or equal to the corresponding threshold, the score is retained, and if it is less than the threshold, the score is considered 0, as follows:

[0054] Surname: 1.0, Country: 0.8, Given name: 0.8

[0055] Cuba - Country: 1.0, Given name: 0

[0056] Next, the results for each category are averaged as follows:

[0057] Surname: (1 + 0) / 2 = 0.5

[0058] Country: (0.8 + 1) / 2 = 0.9

[0059] Given name: (0.8 + 0) / 2 = 0.4

[0060] And, the results are reported as follows:

[0061] Country: 0.9

[0062] Surname: 0.5

[0063] Given name: 0.4

[0064] Without storing or disclosing the underlying data, the reported results can indicate to the user the likelihood that certain classified data (e.g., country, surname, and given name) is being transmitted on the system.

[0065] Suitable system

[0066] The techniques disclosed herein can be embodied as dedicated hardware (e.g., circuitry), programmable circuitry suitably programmed with software and / or firmware, or a combination of dedicated and programmable circuitry. Accordingly, embodiments can include a machine - readable medium having instructions stored thereon that can be used to cause a computer, microprocessor, processor, and / or microcontroller (or other electronic device) to perform a process. The machine - readable medium can include, but is not limited to, optical disks, compact disk read - only memory (CD - ROM), magneto - optical disks, ROM, random access memory (RAM), erasable programmable read - only memory (EPROM), electrically erasable programmable read - only memory (EEPROM), magnetic or optical cards, flash memory, or other types of media / machine - readable media suitable for storing electronic instructions.

[0067] Several embodiments are discussed in more detail below with reference to the accompanying drawings. Figure 8It is a block diagram showing an overview of a device that can operate some embodiments of the disclosed technology. The device may include hardware components of device 700 that determines a risk score and associated pricing and / or risk ratio. Device 700 may include one or more input devices 720 that provide input to the CPU (processor) 710 and notify it of an action. The action is typically intervened by a hardware controller that interprets the signals received from the input device and transmits the information to the CPU 710 using a communication protocol. Input devices 720 include, for example, a mouse, a keyboard, a touch screen, an infrared sensor, a touchpad, a wearable input device, a camera-based or image-based input device, a microphone, or other user input devices.

[0068] The CPU 710 may be a single processing unit or multiple processing units in the device or distributed across multiple devices. For example, the CPU 710 may be coupled to other hardware devices using a bus, such as a PCI bus or a SCSI bus. The CPU 710 may communicate with a hardware controller of a device (e.g., the display 730). The display 730 can be used to display text and graphics. In some examples, the display 730 provides visual feedback of graphics and text to the user. In some embodiments, the display 730 includes an input device as part of the display, such as when the input device is a touch screen or is equipped with an eye direction monitoring system. In some embodiments, the display and the input device are separate. Examples of display devices include: LCD display screens; LED display screens; projection, holographic, or augmented reality displays (e.g., head-up display devices or head-mounted devices); and so on. Other I / O devices 740 may also be coupled to the processor, such as a network card, a video card, a sound card, USB, FireWire, or other external devices, a camera, a printer, a speaker, a CD-ROM drive, a DVD drive, a disk drive, or a Blu-ray device.

[0069] In some implementations, device 700 further includes a communication device capable of wireless or wired communication with network nodes. The communication device may communicate with another device or server over a network using, for example, the TCP / IP protocol. Device 700 may use the communication device to distribute operations across multiple network devices.

[0070] The CPU 710 may have access to the memory 750. The memory includes one or more of a variety of hardware devices for volatile and non-volatile storage and may include read-only and writable memories. For example, the memory may include random access memory (RAM), CPU registers, read-only memory (ROM), and writable non-volatile memories such as flash memory, hard disk drives, floppy disks, CDs, DVDs, magnetic storage devices, tape drives, device buffers, and so on. The memory is not a propagated signal detached from the underlying hardware; thus, the memory is non-transitory. The memory 750 may include a program memory 760 that stores programs and software such as an operating system 762, a word edit distance embedding platform 764, and other application programs 766. The memory 750 may also include a data memory 770 that may include database information and the like that can be provided to any element of the program memory 760 or the device 700.

[0071] Some embodiments may operate in conjunction with many other general purpose or special purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations suitable for use with the present technology include, but are not limited to, personal computers, server computers, hand-held or laptop devices, cellular telephones, mobile phones, wearable electronics, gaming consoles, tablet devices, multiprocessor systems, microprocessor-based systems, set top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like.

[0072] Figure 9 is a block diagram showing an overview of an environment 800 in which some embodiments of the disclosed technology may operate. The environment 800 may include one or more client computing devices 805A-D, examples of which may include the device 700. The client computing devices 805 may operate in a networked environment connected to one or more remote computers (such as server computing devices 810) using logical connections via a network 830.

[0073] In some embodiments, the server computing device 810 may be an edge server that receives client requests and coordinates the implementation of these requests through other servers (such as servers 820A-C). The server computing devices 810 and 820 may include computing systems such as the device 700. Although each server computing device 810 and 820 is logically shown as one server, the server computing devices may each be a distributed computing environment that includes multiple computing devices located at the same or geographically different physical locations. In some embodiments, each server computing device 820 corresponds to a group of servers.

[0074] The client computing device 805 and the server computing devices 810 and 820 can each act as a server or a client to other server / client devices. The server 810 can be connected to the database 815. The servers 820A-C can each be connected to a corresponding database 825A-C. As described above, each server 820 can correspond to a group of servers, and each of these servers can share a database or can have its own database. The databases 815 and 825 can store information repositories (e.g., store). Although the databases 815 and 825 are logically shown as a single unit, the databases 815 and 825 can each be a distributed computing environment that includes multiple computing devices, can be located within their respective servers, or can be located in the same or geographically different physical locations.

[0075] The network 830 can be a local area network (LAN) or a wide area network (WAN), but can also be other wired or wireless networks. The network 830 can be the Internet or some other public or private network. The client computing device 805 can be connected to the network 830 through a network interface (e.g., through wired or wireless communication). Although the connection between the server 810 and the server 820 is shown as a separate connection, these connections can be any type of local, wide area, wired, or wireless network, including the network 830 or a separate public or private network.

[0076] Figure 10 is a block diagram showing components 900, which in some embodiments can be used in a system employing the disclosed technology. The components 900 include hardware 902, general software 920, and dedicated components 940. As discussed above, a system implementing the disclosed technology can use a variety of hardware, including a processing unit 904 (e.g., CPU, GPU, APU, etc.), working memory 906, storage memory 908, and input and output devices 910. The components 900 can be implemented in a client computing device such as the client computing device 805 or on a server computing device such as the server computing devices 810 or 820.

[0077] The general software 920 can include a variety of applications, including an operating system 922, native programs 924, and a basic input / output system (BIOS) 926. The dedicated components 940 can be sub-components of the general software application 920, such as the native programs 924. The dedicated components 940 can include a preparation module 944, a training module 946, a matching module 948, and components that can be used to transfer data and control the dedicated components, such as an interface 942. In some embodiments, the components 900 can be located in a computing system distributed across multiple computing devices, or can be an interface to a server-based application that executes one or more dedicated components 940.

[0078] Those skilled in the art will understand that the components shown in the above Figures 8 - 10 and in each of the flowcharts discussed above can be changed in various ways. For example, the order of the logic can be rearranged, sub-steps can be executed in parallel, the shown logic can be omitted, other logic can be included, etc. In some embodiments, one or more of the components described above can perform one or more of the processes described herein.

[0079] Although specific embodiments have been shown by way of example in the drawings and described in detail above, other embodiments are possible. For example, in some embodiments, a system for classifying words in a batch of words can include at least one storage device storing instructions for causing at least one processor to create a dictionary vector for each of a plurality of dictionary words using a neural network (NN), store each dictionary vector and a classification identifier corresponding to the associated dictionary word, and create a word vector for each word in the batch of words for classification using the NN. Find the most matching dictionary vector for each word vector and report the classification identifier of the most matching dictionary vector for each word vector in the batch.

[0080] In some embodiments, the dictionary words and the words in the batch are encoded as a 30-character by 50-length matrix. In some embodiments, the NN is a convolutional neural network. In some embodiments, each word vector is created without storing the corresponding word in the batch. In some embodiments, each dictionary vector and the corresponding classification identifier are indexed to facilitate similarity search. In some embodiments, the system can also include training the NN, including: calculating a similarity metric for a plurality of pairs of training words of a plurality of training words; creating a pair of training vectors for each pair of training words using the NN and a twin of the NN; calculating the cosine similarity between each pair of training vectors; and comparing the similarity metric and the cosine similarity of each pair of training words and adjusting the NN based on the comparison. In some embodiments, the training words include dictionary words. In some embodiments, the similarity metric is equal to 1 minus the Levenshtein edit distance divided by the length of the longest word in the pair. In some embodiments, the cosine similarity is calculated by dividing the dot product of the pair of training vectors by the product of their Euclidean norms. In some embodiments, each of one or more of the most matching dictionary vectors has an associated score, and reporting the classification identifier of one or more of the most matching dictionary vectors for each word vector in the batch includes: for each word in the batch, selecting the highest score for each classification identifier; and averaging the selected highest scores for each classification identifier.

[0081] In another representative embodiment, a system for classifying words in a batch of words may include at least one storage device storing instructions for causing at least one processor to train a convolutional neural network (CNN); encode each of a plurality of dictionary words into a matrix form; use the trained CNN to create a dictionary vector for each of the plurality of encoded dictionary words; store each dictionary vector and a classification identifier corresponding to the associated dictionary word; encode each word in the batch of words to be classified into a matrix form; use the trained CNN to create a word vector for each encoded word in the batch; find the most matching dictionary vector for each word vector; and report the classification identifier of the most matching dictionary vector for each word vector in the batch. Training the CNN may include: encoding each of a plurality of training words into a matrix form; calculating a similarity metric for multiple pairs of training words of the plurality of training words; using the CNN and a twin of the CNN to create a pair of training vectors for each pair of the encoded training words corresponding to the multiple pairs of training words; calculating the cosine similarity between each pair of training vectors; comparing the similarity metric of each pair of training words and the cosine similarity, and adjusting the CNN based on the comparison; and storing the trained CNN;

[0082] In a further representative embodiment, a system for classifying words in a batch of words may include at least one storage device storing instructions for causing at least one processor to: train a neural network (NN); use the trained NN to create a dictionary vector for each of a plurality of dictionary words; store each dictionary vector and a classification identifier corresponding to the associated dictionary word; use the trained NN to create a word vector for each word in the batch of words to be classified; find the most matching dictionary vector for each word vector; and report the classification identifier of the most matching dictionary vector for each word vector in the batch.

[0083] Training the NN may include calculating a similarity metric for multiple pairs of training words of the plurality of training words, where the similarity metric is equal to 1 minus the Levenshtein edit distance divided by the length of the longest word in the pair; using the NN and a twin of the NN to create a pair of training vectors for each pair of the multiple pairs of training words; calculating the cosine similarity between each pair of training vectors; and comparing the similarity metric of each pair of training words and the cosine similarity, and adjusting the NN based on the comparison.

[0084] The following examples provide additional embodiments of the present technology.

[0085] Example

[0086] 1. A system for classifying words in a batch of words, comprising:

[0087] At least one storage device storing instructions for causing at least one processor to:

[0088] Create dictionary vectors for each of multiple dictionary words using a neural network (NN);

[0089] Store each dictionary vector and a classification identifier corresponding to the associated dictionary word;

[0090] Create word vectors for each word in a batch of words for classification using the NN;

[0091] Find one or more most matching dictionary vectors for each word vector; and

[0092] Report the classification identifiers of the one or more most matching dictionary vectors for each word vector in the batch.

[0093] 2. The system as described in Example 1, wherein the dictionary words and the words in the batch are encoded as a 30-character by 50-length matrix.

[0094] 3. The system as described in Example 1 or 2, wherein the NN is a convolutional neural network.

[0095] 4. The system as described in any of Examples 1 to 3, wherein each word vector is created without storing the corresponding word in the batch.

[0096] 5. The system as described in any of Examples 1 to 4, wherein each dictionary vector and the corresponding classification identifier are indexed for facilitating similarity search.

[0097] 6. The system as described in any of Examples 1 to 5, wherein each of the one or more most matching dictionary vectors has an associated score, and wherein reporting the classification identifiers of the one or more most matching dictionary vectors for each word vector in the batch includes:

[0098] For each word in the batch, select the highest score for each classification identifier; and

[0099] Average the selected highest scores for each classification identifier.

[0100] 7. The system as described in any of Examples 1 to 6, further comprising training the NN, including:

[0101] Calculate a similarity metric for multiple pairs of training words;

[0102] Use the NN and a twin of the NN to create a pair of training vectors for each pair of the multiple pairs of training words;

[0103] Calculate the cosine similarity between each pair of training vectors; and

[0104] Compare the similarity measure and the cosine similarity of each pair of training words, and adjust the NN based on the comparison.

[0105] 8. The system of claim 7, wherein the training words include the dictionary words.

[0106] 9. The system of claim 7 or 8, wherein the similarity measure is equal to 1 minus the Levenshtein edit distance divided by the length of the longest word in the pair.

[0107] 10. The system of any of claims 7 to 9, wherein the cosine similarity is calculated by dividing the dot product of the training vectors by the product of their Euclidean norms.

[0108] 11. A system for classifying words in a batch of words, comprising:

[0109] At least one storage device storing instructions for causing at least one processor to:

[0110] Train a convolutional neural network (CNN), comprising:

[0111] Encode each of a plurality of training words into matrix form;

[0112] Calculate a similarity measure for multiple pairs of the plurality of training words;

[0113] Create a pair of training vectors for each pair of the encoded training words corresponding to the multiple pairs of training words using the CNN and a twin of the CNN;

[0114] Calculate the cosine similarity between each pair of training vectors;

[0115] Compare the similarity measure and the cosine similarity of each pair of training words, and adjust the CNN based on the comparison; and

[0116] Store the trained CNN;

[0117] Encode each of a plurality of dictionary words into matrix form;

[0118] Create a dictionary vector for each of the plurality of encoded dictionary words using the trained CNN;

[0119] Store each dictionary vector and a classification identifier corresponding to the associated dictionary word;

[0120] Encode each word in a batch of words for classification into matrix form;

[0121] Create a word vector for each encoded word in the batch using the trained CNN;

[0122] Find one or more most matching dictionary vectors for each word vector; and

[0123] Report the classification identifier of the one or more most matching dictionary vectors for each word vector in the batch.

[0124] 12. The system as described in Example 11, wherein the training words include the dictionary words.

[0125] 13. The system as described in Example 11 or 12, wherein the similarity metric is equal to 1 minus the Levenshtein edit distance divided by the length of the longest word in the pair.

[0126] 14. The system as described in any of Examples 11 to 13, wherein the cosine similarity is calculated by dividing the dot product of the pair of training vectors by the product of their Euclidean norms.

[0127] 15. The system as described in any of Examples 11 to 14, wherein the training words, the dictionary words, and the words in the batch are encoded into a 30 - character × 50 - length matrix.

[0128] 16. The system as described in any of Examples 11 to 15, wherein each word vector is created without storing the corresponding word in the batch.

[0129] 17. The system as described in any of Examples 11 to 16, wherein the dictionary vectors and the corresponding classification identifiers are indexed for easy similarity search.

[0130] 18. The system as described in any of Examples 11 to 17, wherein each of the one or more most matching dictionary vectors has an associated score, and wherein reporting the classification identifier of the one or more most matching dictionary vectors for each word vector in the batch includes:

[0131] For each word in the batch, selecting the highest score for each classification identifier; and

[0132] Averaging the selected highest scores for each classification identifier.

[0133] 19. A system for classifying words in a batch of words, comprising:

[0134] At least one storage device storing instructions for causing at least one processor to:

[0135] Train a neural network (NN), including:

[0136] Calculate a similarity metric for multiple pairs of training words, where the similarity metric is equal to 1 minus the Levenshtein edit distance divided by the length of the longest word in the pair; use the NN and the twin of the NN to create a pair of training vectors for each pair of the multiple pairs of training words;

[0137] Calculate the cosine similarity between each pair of training vectors; and compare the similarity metric and the cosine similarity of each pair of training words and adjust the NN based on the comparison;

[0138] Use the trained NN to create a dictionary vector for each word in multiple dictionary words;

[0139] Store each dictionary vector and the classification identifier corresponding to the associated dictionary word;

[0140] Use the trained NN to create a word vector for each word in a batch of words for classification;

[0141] Find one or more most matching dictionary vectors for each word vector; and

[0142] Report the classification identifier of the one or more most matching dictionary vectors of each word vector in the batch.

[0143] 20. The system according to example 19, wherein the cosine similarity is calculated by dividing the dot product of the pair of training vectors by the product of their Euclidean norms.

[0144] 21. The system according to example 19 or 20, wherein each of the one or more most matching dictionary vectors has an associated score, and wherein reporting the classification identifier of the one or more most matching dictionary vectors of each word vector in the batch includes:

[0145] For each word in the batch, select the highest score for each classification identifier; and

[0146] Score the selected highest score for each classification identifier.

[0147] 22. The system according to any of examples 19 to 21, wherein the training words, the dictionary words, and the words in the batch are encoded as a 30-character × 50-length matrix before creating the corresponding vectors using the NN.

[0148] 23. The system according to any of examples 19 to 22, wherein the NN is a convolutional neural network.

[0149] 24. The system according to any of examples 19 to 23, wherein each word vector is created without storing the corresponding word in the batch.

[0150] 25. A system as described in any of Examples 19 to 24, wherein each dictionary vector and corresponding classification identifier are indexed to facilitate similarity search.

[0151] 26. A system as described in any of Examples 19 to 25, wherein the training words include the dictionary words.

[0152] Statement

[0153] The above description and drawings are illustrative and should not be construed as limiting. Many specific details are described to provide a thorough understanding of the technology. However, in some instances, well-known details are not described to avoid obscuring the description. Additionally, various modifications may be made without departing from the scope of the embodiments.

[0154] References to "one embodiment" or "an embodiment" in this specification mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the technology. The appearances of the phrase "in one embodiment" in various places in the specification are not necessarily all referring to the same embodiment, nor are they separate or alternative embodiments mutually exclusive of other embodiments. Additionally, various features are described that may be shown by some embodiments and not by others. Similarly, various requirements are described that may be for some embodiments and not for others.

[0155] The terms used in this specification generally have their ordinary meaning in the art, within the context of the present disclosure, and in the particular context in which each term is used. It is understood that the same thing can be described in multiple ways. Thus, alternative languages and synonyms may be used for any one or more of the terms discussed herein, and whether a term is elaborated or discussed herein does not have any special significance. Synonyms for some terms are provided. The listing of one or more synonyms does not exclude the use of other synonyms. Examples are used anywhere in this specification, including examples of any of the terms discussed herein, for illustrative purposes only and are not intended to further limit the scope and meaning of the technology or any example term. Similarly, the technology is not limited to the various embodiments given in this specification. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this technology pertains. In case of conflict, the present document (including definitions) shall prevail.

Claims

1. A system, comprising: At least one hardware processor; and at least one non - transitory memory storing instructions which, when executed by the at least one hardware processor, cause the system to: obtain a first plurality of words and a first plurality of classifications, wherein the words in the first plurality of words are classified into the classifications in the first plurality of classifications; obtain a second plurality of words, wherein a specific classification associated with the second plurality of words is unknown; obtain a first plurality of vectors representing the first plurality of words and a second plurality of vectors representing the second plurality of words; for each vector in the second plurality of vectors, obtain a subset of matching vectors in the first plurality of vectors and one or more scores associated with the subset of matching vectors; classify the second plurality of words by: obtaining a subset of matching classifications associated with the subset of matching vectors and the one or more scores associated with the subset of matching vectors; selecting the highest score associated with each classification in the subset of matching classifications; obtaining one or more thresholds associated with the subset of matching classifications; comparing the one or more thresholds and the highest score; based on the comparison, setting the scores associated with one or more classifications that do not meet the one or more thresholds to a predetermined value; performing the following steps for each classification in the subset of matching classifications: averaging a subset of scores corresponding to each classification to obtain an average score corresponding to each classification; and reporting at least a portion of the one or more classifications and the average scores corresponding to each classification associated with the second plurality of words, thereby indicating to a user whether sensitive data is included in the second plurality of words without disclosing the second plurality of words.

2. The system of claim 1, comprising instructions to: encode the first plurality of words into a first matrix, Among them, the first matrix including a first dimension and a second dimension, the first dimension indicating a plurality of alphanumeric characters, and the second dimension indicating the position of the characters in the words in the first plurality of words; encode the second plurality of words into a second matrix, wherein the second matrix includes the first dimension and the second dimension; and provide the first matrix and the second matrix to a neural network to obtain the first plurality of vectors and the second plurality of vectors respectively.

3. The system of claim 1, comprising instructions to train a neural network to generate the first plurality of vectors and the second plurality of vectors by the following steps: obtain a pair of words; provide the pair of words to the neural network; obtain from the neural network a first vector representing the first word in the pair of words and a second vector representing the second word in the pair of words; and determine a similarity between the first vector and the second vector, Among them, the similarity increasing as the length associated with the pair of words increases.

4. The system of claim 1, comprising instructions to train a neural network to generate the first plurality of vectors and the second plurality of vectors by the following steps: obtain a pair of words; provide the pair of words to the neural network; Obtain a first vector representing the first word in the pair of words and a second vector representing the second word in the pair of words from the neural network; Determine a first similarity measure between the first vector and the second vector, wherein the first similarity measure increases as the length associated with the pair of words increases; Determine a second similarity measure between the first vector and the second vector, wherein the second similarity measure includes cosine similarity; Determine the difference between the first similarity measure and the second similarity measure; and Adjust the weights associated with the neural network to increase the match between the first similarity measure and the second similarity measure.

5. The system according to claim 4, wherein, The first similarity measure is equal to 1 minus the edit distance divided by the length of the longest word in the pair of words.

6. The system according to claim 4, wherein, The second similarity measure is calculated by dividing the dot product of the training vector pair by the product of their Euclidean norms.

7. A method, comprising: Obtain a first plurality of words and a first plurality of classifications, wherein the words in the first plurality of words are classified into the classifications in the first plurality of classifications; Obtain a second plurality of words, wherein the specific classification associated with the second plurality of words is unknown; Obtain a first plurality of vectors representing the first plurality of words and a second plurality of vectors representing the second plurality of words; For each vector in the second plurality of vectors, obtain a subset of matching vectors in the first plurality of vectors and one or more scores associated with the subset of matching vectors; Classify the second plurality of words by: Obtain a subset of matching classifications associated with the subset of matching vectors and the one or more scores associated with the subset of matching vectors; For each classification in the subset of matching classifications, perform the following steps: Average the subset of scores corresponding to each classification to obtain the average score corresponding to each classification; and Report at least a portion of the subset of matching classifications and the average scores corresponding to each classification indicator associated with the second plurality of words, thereby indicating to the user whether sensitive data is included in the second plurality of words without disclosing the second plurality of words.

8. The method of claim 7, comprising: Encode the first plurality of words into a first matrix, wherein the first matrix includes a first dimension and a second dimension, the first dimension indicating a plurality of alphanumeric characters, and the second dimension indicating the position of the character in the words of the first plurality of words; Encode the second plurality of words into a second matrix, wherein the second matrix includes the first dimension and the second dimension; and Provide the first matrix and the second matrix to a neural network to obtain the first plurality of vectors and the second plurality of vectors respectively.

9. The method of claim 7, comprising training a neural network to produce the first plurality of vectors and the second plurality of vectors by the following steps: Obtain a pair of words; Provide the pair of words to the neural network; Obtain a first vector representing the first word in the pair of words and a second vector representing the second word in the pair of words from the neural network; and Determine the similarity between the first vector and the second vector, Among them, wherein the similarity increases as the length associated with the pair of words increases.

10. The method according to claim 7, comprising training a neural network to generate the first plurality of vectors and the second plurality of vectors by the following steps: Obtain a pair of words; Provide the pair of words to the neural network; Obtain a first vector representing the first word in the pair of words and a second vector representing the second word in the pair of words from the neural network; Determine a first similarity measure between the first vector and the second vector, Among them, wherein the first similarity measure increases as the length associated with the pair of words increases; Determine a second similarity measure between the first vector and the second vector, wherein the second similarity measure includes cosine similarity; Determine the difference between the first similarity measure and the second similarity measure; and Adjust the weights associated with the neural network to increase the match between the first similarity measure and the second similarity measure.

11. The method according to claim 10, wherein, The first similarity measure is equal to 1 minus the edit distance divided by the length of the longest word in the pair of words.

12. The method according to claim 10, wherein, The second similarity measure is calculated by dividing the dot product of the first vector and the second vector by the product of their Euclidean norms.

13. The method according to claim 7, comprising: Classify the second plurality of words by: Obtain a subset of the matching classifications associated with the subset of the matching vectors and one or more scores associated with the subset of the matching vectors; Select the highest score associated with each classification in the subset of the matching classifications; Obtain one or more thresholds associated with the subset of the matching classifications; Compare the one or more thresholds and the highest score; Based on the comparison, set the scores associated with one or more classification indicators that do not meet the one or more thresholds to a predetermined value; and For each classification in the subset of the matching classifications, perform the following steps: Average the subset of scores corresponding to each classification to obtain the average score corresponding to each classification.

14. A non-transitory computer-readable storage medium, including instructions recorded thereon, wherein, When the instructions are executed by at least one data processor of the system, cause the system to: Obtain a first plurality of words and a first plurality of classifications, wherein the words in the first plurality of words are classified into the classifications in the first plurality of classifications; Obtain a second plurality of words, wherein the specific classification associated with the second plurality of words is unknown; Obtain a first plurality of vectors representing the first plurality of words and a second plurality of vectors representing the second plurality of words; For each vector in the second plurality of vectors, obtain a subset of the matching vectors in the first plurality of vectors and one or more scores associated with the subset of the matching vectors; Classify the second set of words by: Obtain a subset of the matching classifications associated with the subset of the matching vectors and the one or more scores associated with the subset of the matching vectors; Perform the following steps for each classification in the subset of the matching classifications: Average the subset of scores corresponding to each classification to obtain an average score corresponding to each classification; and Report at least a portion of the subset of the matching classifications and the average scores corresponding to each classification indicator associated with the second plurality of words, thereby indicating to the user whether sensitive data is included in the second plurality of words without disclosing the second plurality of words.

15. The non-transitory computer-readable storage medium according to claim 14, comprising instructions to: Encode the first plurality of words into a first matrix, Among them, The first matrix includes a first dimension and a second dimension, the first dimension indicating a plurality of alphanumeric characters, and the second dimension indicating the position of the character in the words of the first plurality of words; Encode the second plurality of words into a second matrix, wherein, the second matrix includes the first dimension and the second dimension; and Provide the first matrix and the second matrix to a neural network to obtain the first plurality of vectors and the second plurality of vectors respectively.

16. The non-transitory computer-readable storage medium according to claim 14, comprising instructions to train a neural network to generate the first plurality of vectors and the second plurality of vectors by the following steps: Obtain a pair of words; Provide the pair of words to the neural network; Obtain from the neural network a first vector representing the first word in the pair of words and a second vector representing the second word in the pair of words; and Determine the similarity between the first vector and the second vector, Among them, The similarity increases as the length associated with the pair of words increases.

17. The non-transitory computer-readable storage medium according to claim 14, comprising instructions to train a neural network to generate the first plurality of vectors and the second plurality of vectors by the following steps: Obtain a pair of words; Provide the pair of words to the neural network; Obtain from the neural network a first vector representing the first word in the pair of words and a second vector representing the second word in the pair of words; Determine a first similarity measure between the first vector and the second vector, wherein, the first similarity measure increases as the length associated with the pair of words increases; Determine a second similarity measure between the first vector and the second vector, wherein, the second similarity measure includes cosine similarity; Determine the difference between the first similarity measure and the second similarity measure; and Adjust the weights associated with the neural network to increase the match between the first similarity measure and the second similarity measure.

18. The non-transitory computer-readable storage medium according to claim 17, wherein, The first similarity measure is equal to 1 minus the edit distance divided by the length of the longest word in the pair of words.

19. The non-transitory computer-readable storage medium according to claim 17, wherein, The second similarity measure is calculated by dividing the dot product of the first vector and the second vector by the product of their Euclidean norms.

20. The non-transitory computer-readable storage medium according to claim 14, comprising instructions to: classify the second plurality of words by: obtaining a subset of the matching classifications associated with a subset of the matching vectors and the one or more scores associated with the subset of the matching vectors; selecting the highest score associated with each classification in the subset of the matching classifications; obtaining one or more thresholds associated with the subset of the matching classifications; comparing the one or more thresholds and the highest score; based on the comparison, setting the scores associated with one or more classification indicators that do not meet the one or more thresholds to a predetermined value; and performing the following steps for each classification in the subset of the matching classifications: averaging a subset of scores corresponding to each classification to obtain the average score corresponding to each classification.