Processing of classification field values in machine learning applications

Generate high-dimensional numerical representations of classification values ​​through embedding technology and process these representations in auxiliary neural networks, which solves the problem that categorical variable representations in machine learning models are difficult to maintain relevant information, and realizes efficient use of computing resources.

CN119990369APending Publication Date: 2025-05-13EXPEDIA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510082875.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-03-13
Filing Date
2020-03-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

It is difficult to maintain relevant information in machine learning models, while avoiding the use of excessive computing resources.

Method used

High-dimensional numerical representations of classification values ​​are generated by using embedding techniques and processed in auxiliary neural networks to generate low-dimensional features passed to the main neural network.

Benefits of technology

Effectively maintaining relevant information of categorical variables reduces the excessive combination growth of the main neural network and improves the efficiency of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990369A_ABST
    Figure CN119990369A_ABST
Patent Text Reader

Abstract

Systems and methods for processing classification field values in machine learning applications, particularly neural networks, are disclosed. Classification field values are typically converted to vectors prior to being passed to the neural network. However, low-dimensional vectors limit the ability of the network to understand correlations between contextually, semantically, or featured similar values. On the contrary, the high-dimensional vector may suppress the neural network, resulting in the network finding a correlation with respect to individual dimension values, which may be false. The disclosure relates to a hierarchical neural network comprising a primary network and one or more secondary networks. Classification field values are processed in the secondary network to reduce the dimensions of the values prior to being processed by the primary network. This enables contextual, semantic and feature dependencies to be identified without overloading the entire network.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of a patent application with application number 2020800206309, application date March 10, 2020, and invention name “Processing of classification field values ​​in machine learning applications”. Background Art

[0002] In general, machine learning is a data analysis application that seeks to automate the construction of analytical models. Machine learning has been applied to various fields in an effort to understand data correlations that may be difficult or impossible to detect using well-defined models. For example, machine learning has been applied to machine learning systems 118s to model how various data fields known at the time of a transaction (e.g., cost, account identifier, transaction location, purchased items) are related to the percentage probability of transaction fraud. The values ​​of these fields related to historical data and subsequent fraud rates are passed through a machine learning algorithm to generate a statistical model. When a new transaction is attempted, the value of the field can be passed through the model to generate a numerical value indicating the percentage probability of a new transaction fraud. Many machine learning models are known in the art, such as neural networks, decision trees, regression algorithms, and Bayesian algorithms.

[0003] One problem that arises in machine learning is the representation of categorical variables. Categorical variables are those variables that usually take one of a finite set of possible values, where each value represents a specific individual or group. For example, a categorical variable may include a color (e.g., "green", "blue", etc.) or a location (e.g., "Seattle", "New York", etc.). Typically, categorical variables do not imply sorting. On the contrary, ordinal values ​​are used to represent sorting. For example, scores (e.g., "1", "2", "3", etc.) can be ordinal values. Machine learning algorithms are typically developed to introduce numerical representations of data. However, in many cases, machine learning algorithms are constructed to assume that the numerical representations of data are ordinal. This leads to wrong conclusions. For example, if the colors "green", "blue", and "red" are represented as values ​​1, 2, and 3 in a machine learning algorithm, the algorithm may assume that the average value of "green" and "red" (represented as half of the sum of 1 and 3) is equal to 2, or "blue". This wrong conclusion leads to errors in the model output.

[0004] The difficulty in representing categorical variables often stems from the dimensionality of the variable. As nominal terms, two categorical values ​​can represent correlations in a wide variety of abstract dimensions that are easy for humans to recognize but difficult for machines to represent. For example, "ship" and "ship" are easily seen as strongly correlated by humans, but this correlation is difficult to represent to machines. Various attempts have been made to reduce the abstract dimensions of categorical variables to concrete numerical forms. For example, it is common practice to simplify each categorical value into a number that represents the correlation with the final correlation value. For example, in the context of fraud detection, any name that has been associated with fraud can be assigned a high value, while names that are not associated with fraud can be assigned a low value. This approach is disadvantageous because slight changes in names can escape detection, and because users with common names may be inaccurately accused of fraud. In contrast, in the case of converting each categorical value into a multidimensional value (trying to specifically represent the abstract dimension of the variable), the complexity of the machine learning model increases rapidly. For example, a machine learning algorithm can often treat each dimension of a value as a different "feature" - a value that will be compared with other different values ​​to indicate the correlation of a given output. As the number of model features increases, the complexity of the model also increases. However, in many cases, the individual values ​​of a multidimensional categorical variable cannot be compared individually. For example, if the name "John Doe" is converted into a vector of n values, the correlation between the first of these n values ​​and the network address that initiated the transaction may not have predictive value. Therefore, comparing each of the n values ​​to the network address may result in excessive and inefficient use of computing resources. (In contrast, comparing the set of n values ​​representing the name "John Doe" as a whole to a range of network addresses may have predictive value—if this name is associated with fraud and, for example, originates from an address in a country where fraud is prevalent). Therefore, representing categorical variables as low-dimensional values ​​(e.g., a single value) is computationally efficient, but causes the model to ignore interactions between similar categorical variables. In contrast, representing categorical variables as high-dimensional values ​​is computationally inefficient. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] Embodiments of various inventive features will now be described with reference to the following drawings. Reference numerals are reused throughout the drawings to indicate the corresponding relationship between the referenced elements. The drawings are provided to illustrate the example embodiments described herein and are not intended to limit the scope of the present disclosure.

[0006] Figure 1 A block diagram of a machine learning system 118 is shown that applies a neural network machine learning algorithm to categorical variables in historical transaction data to facilitate predicting transaction fraud.

[0007] Figure 2AA block diagram showing the illustrative generation and flow of data for initializing a fraud detection machine learning model within a network environment, according to some embodiments.

[0008] Figure 2B A block diagram is shown of an illustrative generation and flow of data for utilizing a machine learning system 118 within a network environment, in accordance with some embodiments.

[0009] FIG. 3A to FIG. 3B A visual representation of an exemplary neural network architecture used by the machine learning system 118 is shown in accordance with some embodiments.

[0010] Figure 4 A general architecture of a computing device configured to perform a fraud detection method according to some embodiments is shown.

[0011] Figure 5 A flow chart illustrating an exemplary fraud detection method according to some embodiments. DETAILED DESCRIPTION

[0012] In general, aspects of the present disclosure relate to efficient processing of categorical variables in machine learning models to maintain relevant information about the categorical variables while limiting or eliminating the excessive computing resources required to analyze the relevant information within the machine learning model. Embodiments of the present disclosure can be illustratively used to detect when multiple similar categorical variable values ​​indicate fraud, thereby allowing detection of fraud attempts for other similar categorical variable values. For example, embodiments of the present disclosure can detect a strong correlation between fraud and the use of the names "John Doe" and "John Dohe", and therefore predict that the use of the name "Jon Doe" may also be fraudulent. In order to efficiently process categorical variables, embodiments of the present disclosure utilize "embedding" to generate high-dimensional numerical representations of categorical values. Embedding is a known technique in machine learning that attempts to reduce the dimensionality of a value (e.g., a categorical value) while maintaining important relevant information about the value. These high-dimensional numerical representations are then processed as features of auxiliary neural networks (e.g., input to auxiliary neural networks). The output of each auxiliary neural network is used as a feature of the main neural network, together with other features (e.g., non-categorical variables) to produce an output result, such as a model that provides a percentage probability of transaction fraud. By processing high-dimensional numerical representations in separate auxiliary networks, the interaction of the various dimensions of such representations with other features (e.g., non-categorical variables) is limited, thereby reducing or eliminating excessive combinatorial growth of the entire network. The output of each auxiliary network is limited to represent categorical features in appropriate dimensions based on other data to be analyzed with it. For example, two variables that are usually not semantically or contextually related (e.g., the name and time of a transaction) can be processed as low-dimensional values ​​in the main network (e.g., a single value, each representing a feature of the main network). Variables that are highly semantically or contextually related (e.g., two values ​​of a name variable) can be processed in high dimensions. Variables that are somewhat semantically or contextually related (e.g., names and email addresses, which may overlap in content but differ in overall form) can be processed in intermediate dimensions, for example, by combining the outputs of two initial auxiliary networks and inputting them into an intermediate auxiliary network, and then feeding the outputs of the intermediate auxiliary network into the main neural network. This combination of networks can produce a hierarchical neural network. By using this "layer" of the network, the level of these interactions can be controlled relative to the expected semantic or contextual relevance of the interactions of the features on the neural network, thereby enabling machine learning based on high-dimensional representations of categorical variables without incurring excessive computational resource usage of existing models.

[0013] As described above, in order to process categorical variables, an initial conversion of the variables to numerical values ​​is usually performed. According to an embodiment of the present disclosure, embedding can be used to generate a high-dimensional representation of the variable. As used herein, dimension generally refers to the number of numerical values ​​used to represent categorical values. For example, representing the color value "blue" as a numerical value "1" can be considered a single-dimensional value. Representing the value "blue" as a vector "[1, 0]" can be considered a two-dimensional value, etc.

[0014] An example of an embedding is a "word-level" embedding (also called a "word-level representation"), which attempts to convert words into multidimensional values, with the distance between values ​​indicating the relevance between the words. For example, the words "ship" and "ship" can be converted into values ​​with low distances in a multidimensional space (because both involve watercraft). Similarly, word-level embeddings can convert "shipping" and "mailing" into values ​​with low distances in a multidimensional space (because both are related to sending packages). However, the same word-level embedding can convert "ship" and "mail" into values ​​with high distances in a multidimensional space. Therefore, word-level embeddings can maintain high-level relevant information about human-readable words while representing words in digital form. Word-level embeddings are well known in the art, so they will not be described in detail. However, in short, word-level embeddings typically rely on previous applications of machine learning to a corpus of words. For example, a machine learning analysis performed on a published text may indicate that "dog" and "cat" frequently appear near the word "pet" in the text, and are therefore related. Therefore, the multidimensional representations of "dog" and "cat" according to the embedding can be close in the multidimensional space. An example of a word level embedding algorithm is the one developed by Google TM A "word2vec" algorithm was developed that takes a word as input and produces a multi-dimensional value ("vector") that attempts to preserve contextual information about the word. Other word-level embedding algorithms are known in the art, any of which may be used in conjunction with the present disclosure. In some embodiments, word-level embeddings may be supplemented with historical transaction data to determine contextual relationships between specific words in the context of potentially fraudulent transactions. For example, a corpus of words and data indicating word correspondences and associated fraud (e.g., from historical records indicating the use of each word in a transaction data field, and whether the transaction was ultimately determined to be fraudulent) may be trained in a neural network. The output of the neural network may be a multi-dimensional representation that represents the contextual relationships of the words in the context of the transaction, rather than in a general corpus. In some cases, training of the network that determines word-level embeddings occurs prior to independently training a fraud detection model as described herein. In other cases, training of the network that determines word-level embeddings occurs simultaneously with training a fraud detection model as described herein. For example, training of a neural network that provides word-level embeddings may be represented as an auxiliary network to a layered neural network.

[0015] Another example of an embedding is a "character-level" embedding (also called a "character-level representation"), which attempts to convert a word into a multidimensional value that represents the individual characters in the word (as opposed to a representation used by the semantics of the word as in a word-level embedding). For example, given the general structure of overlapping characters and words, a character-level embedding can convert the words "hel lo" and "yel low" into values ​​that are close to each other in a multidimensional space. Character-level embeddings can be used to capture small variations in categorical values ​​that are uncommon (or unused) in common speech. For example, the two usernames "Johnpdoe" and "Jonhdoe" may not be represented in a corpus, and therefore, a word-level embedding may not be sufficient to represent the usernames. However, a character-level embedding will likely convert both usernames into similar multidimensional values. Like word-level embeddings, character-level embeddings are well known in the art and will not be described in detail. An example of a word-level embedding algorithm is the "seq2vec" algorithm, which takes a string as input and produces a multidimensional value (a "vector") that attempts to retain contextual information about the objects within the string. Although seq2vec models are often applied similarly to "word2vec" to describe contextual information between words, the model can also be trained to identify individual characters as objects, thereby finding contextual information between characters. In this way, character-level embedding models can be viewed as similar to word-level embedding models, because these models take a corpus of strings (e.g., a general word corpus in a given language, a word corpus used in the context of potentially fraudulent transactions, etc.) as input and output a multidimensional representation that attempts to preserve contextual information between characters (e.g., so that characters that appear close to each other in the corpus are assigned vector values ​​that are close to each other in the multidimensional space). Other word-level embedding algorithms are known in the art, any of which can be used in conjunction with the present disclosure.

[0016] After obtaining high-dimensional representations for each value of a given categorical variable (e.g., the name of a person who has made a transaction), these representations can be passed into an auxiliary neural network to generate outputs (e.g., neurons), which in turn are used as features for a subsequent neural network (e.g., an intermediate network or a primary network). A separate auxiliary network can be established for each categorical variable (e.g., name, email address, location, etc.), and the output of each categorical variable can be constrained relative to the number of inputs, which is generally equal to the number of dimensions in the high-dimensional representation of the variable values. For example, in the case where the name is represented as a 100-dimensional vector, the auxiliary network can take the 100 dimensions of each name as 100 input values ​​and produce 3 to 5 neuron outputs. These outputs effectively represent low-dimensional representations of the categorical variable values, which can be passed into a subsequent neural network. The output of the primary network is established as the desired outcome (e.g., a binary classification of whether the transaction is fraudulent or not). The auxiliary network and the primary network are then trained simultaneously such that the output of the auxiliary network represents a low-dimensional representation specific to the desired output (e.g., a binary classification as fraud or non-fraud or a multi-class classification with fraud / abuse types), rather than a generalized low-dimensional representation achieved through embedding (which relies on an established model rather than being trained simultaneously). Therefore, the low-dimensional representation of the categorical variable produced by the auxiliary neural network is expected to maintain semantic or contextual information related to the desired end result without the need to feed a high-dimensional representation into the primary model (which, as described above, would otherwise incur costs associated with attempting to model one or more high-dimensional representations in a single model). Advantageously, utilizing the low-dimensional output of the auxiliary network with the primary network allows a user to test the interaction and correlation of categorical variables with non-categorical variables using fewer computational resources than existing methods.

[0017] As will be understood by those skilled in the art from this disclosure, the embodiments disclosed herein improve the ability of a computing system to perform machine learning associated with categorical variables in an efficient manner. Specifically, embodiments of the present disclosure improve the efficiency of computing resource usage of such a system by utilizing a combination of a main machine learning model and one or more auxiliary models, the auxiliary model enabling the processing of categorical variables as high-dimensional representations while limiting the interaction of those high-dimensional representations with other features passed to the main model. In addition, the currently disclosed embodiments solve technical problems inherent in computing systems; specifically, the limited computing resources used to perform machine learning, and the inefficiency caused by attempting to perform machine learning on high-dimensional representations of categorical variables within the main model. These technical problems are solved by various technical solutions described herein, including using auxiliary models to process high-dimensional representations of categorical variables and providing the output as features to the main model. Therefore, the present disclosure generally represents an improvement over existing data processing systems and computing systems.

[0018] Although embodiments of the present disclosure are described with reference to specific machine learning models, such as neural networks, other machine learning models may be utilized in accordance with the present disclosure.

[0019] The foregoing aspects of the present disclosure and many of the attendant advantages will become more readily appreciated as they become better understood when reference is made to the following description, taken in conjunction with the accompanying drawings.

[0020] Figure 1 1 is a block diagram illustrating an environment 100 in which a machine learning system 118 applies a neural network machine learning algorithm to categorical variables and non-categorical variables in historical data to facilitate classification of later data. Specifically, the machine learning system 118 processes historical data by generating a neural network model that includes both a primary network and an auxiliary network that processes a high-dimensional representation of the categorical variables before passing the output to the primary network. In an illustrative embodiment, the machine learning system 118 processes historical transaction data to generate a binary classification of new outstanding transactions as fraudulent or non-fraudulent. However, in other embodiments, other types of data may be processed to generate other classifications, including binary or non-binary classifications. For example, multiple output nodes of the primary network may be configured such that the network outputs values ​​for use in a multivariate classification system. Figure 1 Environment 100 is depicted as including a user device 102 , a transaction system 106 , and a machine learning system 118 , which can all communicate with each other via a network 114 .

[0021] The transaction system 106 illustratively represents a network-based transaction facilitator that operates to service requests from clients (via user devices 102) to initiate transactions. Transactions can be illustratively the purchase or acquisition of physical objects, non-physical objects, services, etc. Many different types of network-based transaction facilitators are known in the art. Therefore, the operational details of the transaction system 106 can vary between embodiments and are not discussed here. However, for the purpose of discussion, it is assumed that the transaction system 106 maintains historical data that associates various fields related to transactions with the final results of the transactions (e.g., fraud or non-fraud). The fields of each transaction can vary, and can include, for example, fields for transaction time and transaction amount, fields that identify one or more parties to the transaction (e.g., name, birthday, account identifier or user name, email address, mailing address, Internet Protocol (IP) address, etc.), items involved in the transaction (e.g., characteristics of the items, such as the departure and arrival airports of the purchased flight, the brand of the purchased items, etc.), payment information for the transaction (e.g., the type of payment instrument used or credit card number), or other constraints on the transaction (e.g., whether the transaction is refundable). The outcome of each transaction can be determined by monitoring those transactions after they are completed, such as by monitoring "chargebacks" for transactions that were later reported as fraudulent by the impersonating individual. The historical transaction data is illustratively stored in a data store 110, which can be a hard disk drive (HDD), a solid state drive (SSD), a network attached storage (NAS), or any other persistent or substantially persistent data storage device.

[0022] The user device 102 generally represents a device that interacts with the trading system in order to request a transaction. For example, the trading system 106 may provide a user interface, such as a graphical user interface (GUI), through which a client using the user device 102 may submit a transaction request and data fields associated with the request. In some cases, the data fields associated with the request may be independently determined by the trading system 106 (e.g., by independently determining the time of day, by referencing profile information to retrieve data about the customer associated with the request, etc.). The user device 102 may include any number of different computing devices. For example, each user device 102 may correspond to a laptop or tablet computer, a personal computer, a wearable computer, a personal digital assistant (PDA), a hybrid PDA / mobile phone, or a mobile phone.

[0023] User device 102 and transaction system 106 may interact via network 114. Network 114 may be any wired network, wireless network, or combination thereof. In addition, network 114 may be a personal area network, a local area network, a wide area network, a global area network (e.g., the Internet), a cable network, a satellite network, a cellular telephone network, or a combination thereof. Although shown as a single network 114, in some embodiments, Figure 1The components can communicate over a number of potentially different networks.

[0024] As mentioned above, transaction system 106 generally wishes to detect fraudulent transactions before completing the transaction. Figure 1 , the transaction system 106 is shown in communication with a machine learning system 118, which operates to assist in fraud detection by generating a fraud detection model. Specifically, the machine learning system 118 is configured to process high-dimensional representations of categorical variables using auxiliary neural networks, whose outputs are used as features of the main neural network, and the outputs of the main neural network in turn represent the classification of transactions as fraud or non-fraud (the classification can be modeled as, for example, a percentage probability of fraud occurring). To facilitate the generation of the model, the machine learning system includes a vector conversion unit 126, a modeling unit 130, and a risk detection unit 134. The vector conversion unit 126 may include computer code for converting categorical field values ​​(e.g., name, email address, etc.) into high-dimensional numerical representations of those field values. Each high-dimensional numerical representation may take the form of a set of numerical values, generally referred to herein as a vector. In one embodiment, as described above, the categorical field values ​​are converted into numerical representations by using an embedding technique such as word-level or character-level embedding. The modeling unit 130 may represent code for generating and training a machine learning model (e.g., a hierarchical neural network), wherein the high-dimensional numerical representation is first passed through one or more auxiliary neural networks before being passed to the main network. The trained model may then be used by the risk detection unit 134, which may include computer code for passing new field values ​​of an attempted transaction into the trained model to classify the likelihood that the transaction is fraudulent.

[0025] refer to FIG. 2A to FIG. 2B , shows an illustrative interaction for the operation of the machine learning system 118 to generate, train, and utilize a hierarchical neural network that includes one or more auxiliary networks whose outputs serve as features of the main neural network. Specifically, Figure 2A An illustrative interaction for generating and training such a layered neural network is shown, and Figure 2B An illustrative interaction for using a trained network to predict the fraud likelihood of an attempted transaction is shown.

[0026] The interaction begins at (1), where the transaction system 106 sends historical transaction data to the machine learning system 118. In some embodiments, the historical transaction data may include raw data of past transactions that have been processed or submitted to the transaction system 106. For example, the historical data may be a list of all transactions conducted on the transaction system 106 over the course of a three-month period, as well as fields associated with the transactions, such as the transaction time and transaction amount, fields identifying one or more parties to the transaction (e.g., name, birthday, account identifier or username, email address, mailing address, Internet Protocol (IP) address, etc.), items involved in the transaction (e.g., characteristics of the items, such as the departure and arrival airports of the purchased flight, the brand of the purchased item, etc.), payment information for the transaction (e.g., the type of payment instrument used or credit card number), or other constraints on the transaction (e.g., whether the transaction is refundable). The historical data is illustratively "tagged" or annotated with the results of the transaction relative to the desired classification. For example, each transaction may be labeled as "fraudulent" or "non-fraudulent". In some embodiments, the historical data may be stored and transmitted in the form of a text file, table, or other data storage format.

[0027] At (2), the machine learning system 118 obtains neural network hyperparameters for the desired neural network. The hyperparameters may be specified, for example, by an operator of the trading system 106 or the machine learning system 118. In general, the hyperparameters may include those fields within the historical data that should be considered categorical, and the embeddings applied to the field values. The hyperparameters may also include the overall desired structure of the neural network, in terms of auxiliary networks, main networks, and intermediate networks (if any). For example, the hyperparameters may specify, for each categorical field, the number of hidden layers of the auxiliary network associated with the categorical field and the number of units in these layers, as well as the number of output neurons of the auxiliary network. The hyperparameters may similarly specify the number of hidden layers of the main network, the number of units in each such layer, and other non-categorical features to be provided to the main network. If an intermediate network is used between the output of the auxiliary network and the input (“feature”) of the main network, the hyperparameters may specify the structure of such an intermediate network. Various additional hyperparameters known in the art regarding neural networks may also be specified.

[0028] At (3), the machine learning system 118 (e.g., vector conversion unit 126) converts the categorical field values ​​from the historical data into corresponding high-dimensional numerical representations (vectors) specified by the hyperparameters. Illustratively, each categorical field value may be processed according to at least one of a word-level embedding or a character-level embedding as described above to convert the string representation of the field value into a vector. Although a single embedding for a given categorical field is illustratively described, in some cases the same field is represented by different embeddings, each of which is passed to a different auxiliary neural network. For example, a name field may be represented by both word-level and character-level embeddings in order to assess semantic / contextual information (e.g., repeated use of words to mean similar things) and character relationship information (e.g., slight variations in the characters used for names).

[0029] Thereafter, at (4), the machine learning system 118 (e.g., via the modeling unit 130) generates and trains a neural network according to the hyperparameters. Illustratively, for each categorical field specified in the hyperparameters, the modeling unit 130 may generate an auxiliary network that takes as input the value of the vector representation of the field value and provides as output a node set as input to a subsequent network. The number of nodes output by each auxiliary network may be specified in the hyperparameters and may typically be less than the dimension of the vector representation employed by the auxiliary network. Thus, the output of the node set itself may be viewed as a low-dimensional representation of the categorical field value. The modeling unit 130 may combine the output of each auxiliary network in a manner specified in the hyperparameters. For example, the output of each auxiliary network may be used directly as an input to the primary network, or may be used as the output of one or more intermediate networks, the output of which is in turn an input to the primary network. The modeling unit 130 may also provide one or more non-categorical fields as input to the primary network.

[0030] After generating the network structure, the modeling unit 130 can train the network using at least a portion of the historical transaction data. General training of the defined neural network structure is known in the art and will not be described in detail herein. However, in brief, the modeling unit 130 can, for example, divide the historical data into multiple data sets (e.g., training sets, validation sets, and test sets) and process the data sets using a hierarchical neural network (the entire network, including the auxiliary network, the main network, and any intermediate networks) to determine the weights applied to the input data at each node. As a final result, a final model can be generated that takes the fields from the proposed transaction as input and outputs the probability of placing these fields in a given category (e.g., fraud or non-fraud).

[0031] Figure 2BA block diagram is shown of an illustrative generation and flow of data for utilizing a machine learning system 118 within a networked environment in accordance with some embodiments. The data flow may begin when (5) a user requests, via a user device 102, to initiate a transaction on a transaction system 106. For example, the user may attempt to purchase an item from an online website of a commercial retailer. To assist in determining whether to allow the transaction, at (6), the transaction system 106 submits the transaction information (e.g., including the fields discussed above) to the machine learning system 118. The machine learning system 118 (e.g., via the risk detection unit 134) may then apply a previously learned model to the transaction information to obtain a likelihood that the transaction is fraudulent. At (8), the machine learning system 118 sends a final risk score to the transaction system 106 so that the transaction system 106 may determine whether to allow the transaction. Illustratively, the transaction system may establish a threshold likelihood such that any attempted transaction above the threshold is rejected or retained for further processing (e.g., manual or automatic verification).

[0032] FIG. 3A to FIG. 3B is a visual representation of an exemplary hierarchical neural network that may be generated and trained by the machine learning system 118 based at least in part on examining historical data over a period of time, according to some embodiments. Figure 3A A layered neural network with a single auxiliary network connected to the main network is shown. Figure 3B A hierarchical neural network with multiple auxiliary networks, intermediate networks, and a main network is shown.

[0033] Specifically, in Figure 3A , an exemplary hierarchical neural network 300 is shown that includes a single categorical field (e.g., a "name" field) processed by an auxiliary network (shown as a shaded node), the output of which is passed as an input (or feature) into the main network. The auxiliary network includes an input node 302 corresponding to the categorical field value (e.g., "John Doe" for a transaction entry). The auxiliary network also includes a vector layer 304 that represents the value of the categorical field converted to a multi-dimensional vector by embedding. Each node in the vector layer 304 illustratively represents a single numerical value in a vector created by applying the embedding to the categorical field value. Thus, in Figure 3A In the example, embedding the categorical field values ​​may produce a 5-dimensional vector, each value of which is passed to each node in the vector layer 304. In practice, the categorical field values ​​may be converted into a very high-dimensional vector (e.g., 100 dimensions or more), and thus the vector layer 304 may have a larger dimension than the vector layer 304. Figure 3A Although input node 302 is shown for completeness, in some cases, the auxiliary network may exclude the input node because the categorical field values ​​may have been previously converted into vectors. Therefore, vector layer 304 may serve as the input layer of the auxiliary network.

[0034] In addition, the hierarchical network 300 includes a main network (shown as unshaded nodes). The output of the auxiliary network represents the input or features 307 to the main network. In addition, the main network obtains a set of additional features from the non-categorical field 306 (e.g., which can be formed by operator-defined transformations of the non-categorical field values). The main network features 307 are passed through the hidden layer 308 to reach the output node 310. In some embodiments, the output 310 is a final score indicating the fraud likelihood of a given categorical field value 302 and other non-categorical field values ​​306 (e.g., transaction price, transaction time, or other numerical data).

[0035] like Figure 3A As shown, the number of outputs of the auxiliary neural network can be selected to be low relative to the size of the vector layer 304. In one embodiment, the output of the auxiliary network is set to between three and five neurons. Relative to other techniques for incorporating categorical fields into the network 300, the overall complexity of the network 300 can be reduced by using an auxiliary network with a low-dimensional output. For example, in a conventional neural network architecture that relies on simple embedding and cascading, the categorical value can be converted into a 50-dimensional vector by embedding, and the vector can be cascaded with other features of the network, resulting in 50 features being added to the network. As the number of features grows, the complexity of the network and the time required to generate and train the network also increase. Therefore, cascading may be impractical and inefficient, especially when considering multiple categorical values. This inefficiency is exacerbated by the configuration of the neural network to consider features independently rather than as a group. Therefore, the addition of a vector of 50 features will unnecessarily cause the network to find the correlation between each of these 50 features and other non-categorical features-the correlation may be spurious.

[0036] Compared to traditional neural network techniques that rely on simple embedding and concatenation of categorical features with other non-categorical features, network 300 does not concatenate the vector representation of the categorical field with other non-categorical features, but processes the categorical field via an auxiliary network. By avoiding traditional concatenation, network 300 can keep the entire vector as a semantic unit, and by processing each value in the vector separately without losing the semantic relationship. Advantageously, network 300 can avoid learning unnecessary and meaningless interactions between each value, and does not inadvertently impose unnecessary complexity and invalid relationship and interaction mapping.

[0037] Figure 3B An exemplary hierarchical neural network 311 is shown having multiple auxiliary networks 312, intermediate networks 314, and a main network 316. Many elements of network 311 are similar to Figure 3A The network 300 will therefore not be described again. However, in contrast to the network 300, Figure 3BThe network 311 includes three auxiliary networks, network 312A to network 312C. Each network illustratively corresponds to a categorical field, which is converted to a high-dimensional vector by embedding before dimensionality reduction by the corresponding auxiliary network 312. The output of the auxiliary network 312 is taken as the input of the intermediate network 314, which again reduces the dimension of the output. The use of the intermediate network 314 is beneficial, for example, it can enable the detection of correlations between multiple categorical field values ​​without attempting to detect correlations with non-categorical field values. For example, the intermediate network 314 can be used to detect higher-level correlations between a user's name, an email address, and a mailing address (for example, so that when these three fields are related in some way, the possibility of fraud is greater or less). The output of the intermediate network 314 generally loses information related to the input to the network 314, and therefore the main network does not need to attempt to detect higher-level correlations between the user's name and other non-categorical fields (e.g., transaction amount). Therefore, the layered network 311 enables the interaction of different fields to be controlled, thereby limiting the network to only check those expected correlations rather than false correlations.

[0038] Figure 4 A general architecture of a computing device configured to perform a fraud detection method according to some embodiments is shown. Figure 4 The general architecture of the machine learning system 118 shown in FIG. 1 includes an arrangement of computer hardware and software that can be used to implement various aspects of the present disclosure. The hardware can be implemented on a physical electronic device, as discussed in more detail below. The machine learning system 118 can include more than Figure 4 Those shown may be more (or less) than those shown. However, it is not necessary to show all of these generally conventional elements in order to provide an enabling disclosure. In addition, Figure 4 The general architecture shown in can be used to implement Figure 1 One or more other components shown in .

[0039] As shown, the machine learning system 118 includes a processing unit 490, a network interface 492, a computer-readable medium drive 494, and an input / output device interface 496, all of which can communicate with each other via a communication bus. The network interface 492 can provide a connection to one or more networks or computing systems. The processing unit 490 can therefore receive information and instructions from other computing systems or services via the network 114. The processing unit 490 can also communicate with the memory 480 and also provide output information for an optional display (not shown) via the input / output device interface 496. The input / output device interface 496 can also accept input from an optional input device (not shown).

[0040] Memory 480 may contain computer program instructions (grouped into units in some embodiments) that processing unit 490 executes to implement one or more aspects of the present disclosure. Memory 480 corresponds to one or more layers of memory devices, including (but not limited to) RAM, 3D XPOINT memory, flash memory, magnetic storage, etc.

[0041] The memory 480 may store an operating system 484 that provides computer program instructions for use by the processing unit 490 in the general management and operation of the machine learning system 118. The memory 480 may also include computer program instructions and other information for implementing aspects of the present disclosure. For example, in one embodiment, the memory 480 includes a user interface unit 482 that generates a user interface (and / or instructions thereof) for display on a computing device, such as via a navigation and / or browsing interface (e.g., a browser or application installed on the computing device).

[0042] In addition to and / or in combination with the user interface unit 482, the memory 480 may include a vector conversion unit 126 configured to convert the categorical fields into vector representations. The vector conversion unit 126 may include lookup tables, mappings, etc. to facilitate these conversions. For example, in the case where the vector conversion unit 126 implements the word2vec algorithm, the unit 126 may include a lookup table that enables conversion of each word in the dictionary into a corresponding vector, which may be generated by training the word2vec algorithm separately for a corpus of words. The unit 126 may include similar lookup tables or mappings to facilitate character-level embeddings, such as tables or mappings generated by implementing the seq2vec algorithm.

[0043] The memory 480 may also include a modeling unit 130 configured to generate and train a hierarchical neural network. The memory 480 may also include a risk detection unit 134 to pass transaction data through a trained machine learning model to detect fraud.

[0044] Figure 5 is a flow chart illustrating an example routine 500 for processing categorical field values ​​in a machine learning application by using an auxiliary network. The routine 500 may be, for example, Figure 1 More specifically, the routine 500 illustrates the interactions for generating and training a hierarchical neural network to classify events or items. Figure 5 In the context of , the routine 500 will be described with reference to classifying transactions as fraudulent or non-fraudulent based on historical transaction data. However, other types of data may also be processed via the routine 500.

[0045] The routine 500 begins at block 510, where the machine learning system 118 receives labeled data. The labeled data may include, for example, a list of past transactions from the trading system 106, labeled according to whether the transaction is fraudulent. In some embodiments, the historical data may include a past record of all transactions that occurred through the trading system 106 within a period of time (e.g., within the past 12 months).

[0046] The routine 500 then continues to box 515, where the system 118 obtains hyperparameters for the hierarchical neural network to be trained based on the labeled data. The hyperparameters may include, for example, indications of which fields of the labeled data are categorical, and appropriate embeddings to be applied to the categorical field values ​​to produce high-dimensional vectors. The hyperparameters may also include the desired structure of the auxiliary networks created for each categorical value, such as the number of hidden layers or output nodes to be included in each auxiliary network. In addition, the hyperparameters may specify the desired hierarchical structure of the hierarchical neural network, such as whether one or more auxiliary networks should be merged via an intermediate network before being passed to the main network, and the size and structure of the intermediate network. The hyperparameters may also include parameters for the main network, such as the number of hidden layers and the number of nodes in each layer.

[0047] At block 520, the machine learning system 118 converts the categorical field values ​​(as represented in the labeled data) into vectors, as indicated within the hyperparameters. Implementation of block 520 may include embedding the field values ​​according to a predetermined transformation. In some examples, these transformations may occur during training of the hierarchical network, so it may not be necessary to implement block 520 as a different block.

[0048] At block 525, the machine learning system 118 generates and trains a hierarchical neural network, including an auxiliary network, a main network, and an intermediate network (if specified in the hyperparameters) for each categorical field value identified within the hyperparameters. Examples of models that can be generated are described in Figure 3A and Figure 3B , as described above. In one embodiment, based on the hyperparameters, the network is programmatically generated by initially generating an auxiliary network for each categorical value, merging the outputs of those auxiliary networks via intermediate networks (if specified within the hyperparameters), and combining the outputs of the auxiliary networks (or alternatively one or more intermediate networks) with non-categorical feature values ​​as input to the main network. Thus, while the hyperparameters may specify the overall structural considerations of the hierarchical network, in some instances the network itself need not be explicitly modeled by a human operator. After the network is generated, the machine learning system 118 trains the network via labeled data in accordance with traditional neural network training. As a result, a model is generated that produces a categorical value as an output (e.g., the risk that a transaction is fraudulent) for a given record of an input field.

[0049] Once the machine learning model is generated and trained in block 525, the machine learning system 118 receives new transaction data in block 530. In some embodiments, the new transaction data may correspond to a new transaction initiated by a user on the transaction system 106, which the transaction system 106 sends to the machine learning system 118 for review. At block 535, the system 118 processes the received data via the generated and trained hierarchical model to generate a classification value (e.g., the risk that the transaction is fraudulent). At block 545, the system 118 then outputs the classification value (e.g., to the transaction system 106). Thus, the transaction system 106 can utilize the classification value to determine, for example, whether to allow or deny the transaction. The routine 500 then ends.

[0050] Embodiments of the present disclosure may be described in light of the following terms:

[0051] Clause 1. A system for processing categorical field values ​​in a machine learning application, comprising:

[0052] a data store including marked transaction records, each record corresponding to a transaction and including a value for a respective field within a set of fields associated with the transaction and marked to indicate whether the transaction is determined to be fraudulent;

[0053] One or more processors configured with computer-executable instructions to at least:

[0054] obtaining hyperparameters for a hierarchical neural network, the hyperparameters identifying at least a categorical field within the set of fields and an embedding process to be used to convert a value of the categorical field into a multidimensional vector;

[0055] Generate a multidimensional vector of the categorical field by converting the field value of the categorical field within the record according to the embedding process;

[0056] generating an auxiliary neural network that takes the multi-dimensional vectors as input and outputs a low-dimensional representation of the vector for each vector;

[0057] generating a hierarchical neural network including at least the auxiliary neural network and a main neural network, wherein the main neural network takes as input a combination of the low dimensional representation output by the auxiliary neural network and one or more values ​​of non-categorical fields within the set of fields, and wherein the main neural network outputs a binary classification indicating a likelihood that a single transaction corresponding to the input record is fraudulent;

[0058] Training the hierarchical neural network based on the labeled transaction data to generate a trained model;

[0059] processing new transaction records according to the trained model to determine the likelihood that the new transaction is fraudulent; and

[0060] The likelihood that the new transaction is fraudulent is output.

[0061] Clause 2. A system according to clause 1, wherein the classification field represents at least one of a name, a user name, an email address, or a mailing address of each party to each transaction.

[0062] Clause 3. A system according to clause 1, wherein the non-categorical field represents an ordinal value or a numerical value for each transaction.

[0063] Clause 4. A system according to clause 3, wherein the ordinal value includes at least one of a transaction volume or a transaction time of the transaction.

[0064] Clause 5. A system according to clause 1, wherein the embedding process represents at least one of word-level embedding or character-level embedding.

[0065] Clause 6. A computer-implemented method comprising:

[0066] obtaining marked transaction records, each record corresponding to a transaction and including a value for a respective field within a set of fields associated with the transaction and marked to indicate whether the transaction is determined to be fraudulent;

[0067] obtaining hyperparameters for a hierarchical neural network, the hyperparameters identifying at least a categorical field within the set of fields and an embedding process to be used to convert a value of the categorical field into a multidimensional vector;

[0068] generating the multidimensional vector;

[0069] Generate a hierarchical neural network including at least an auxiliary neural network and a main neural network, wherein:

[0070] The auxiliary neural network takes the multidimensional vector as input and for each

[0071] vector outputs a low-dimensional representation of the vector; and

[0072] the primary neural network taking as input a combination of the low dimensional representation output by the auxiliary neural network and one or more values ​​of non-categorical fields within the set of fields, and wherein the primary neural network outputs a binary classification indicating a likelihood that a single transaction corresponding to the input record is fraudulent;

[0073] Training the hierarchical neural network based on the labeled transaction records to generate a trained model;

[0074] processing new transaction records according to the trained model to determine the likelihood that the new transaction is fraudulent; and

[0075] The likelihood that the new transaction is fraudulent is output.

[0076] Clause 7. A computer-implemented method according to clause 6, wherein the hyperparameter identifies one or more additional categorical fields within the set of fields, and wherein the hierarchical neural network includes an additional auxiliary neural network for each of the one or more additional categorical fields, and the output of each additional auxiliary neural network represents an additional input to the main neural network.

[0077] Clause 8. A computer-implemented method according to Clause 7, wherein the low-dimensional representation is represented by a set of output neurons of the auxiliary neural network.

[0078] Clause 9. The computer-implemented method of clause 7, wherein generating the multidimensional vector comprises, for each value of the categorical field, referencing a lookup table that identifies the corresponding multidimensional vector.

[0079] Clause 10. A computer-implemented method according to clause 7, wherein the lookup table is generated by pre-applying a machine learning algorithm to a corpus of values ​​of the categorical field.

[0080] Clause 11. A computer-implemented method according to clause 7, wherein the hierarchical neural network further includes an intermediate neural network, which provides the low-dimensional representation output by the auxiliary neural network to the main neural network.

[0081] Clause 12. A computer-implemented method according to clause 11, wherein the intermediate neural network further reduces the dimensionality of the low dimensional representation output by the auxiliary neural network before providing the low dimensional representation to the main neural network.

[0082] Clause 13. A computer-implemented method according to clause 7, wherein the embedding process represents at least one of word-level embedding or character-level embedding.

[0083] Clause 14. A non-transitory computer-readable medium comprising computer-executable instructions that, when executed by a computing system, cause the computing system to:

[0084] Obtaining labeled records, each record including values ​​of each field in a field set and a classification of the labeled record;

[0085] obtaining hyperparameters for a hierarchical neural network, the hyperparameters identifying at least a categorical field within the set of fields and an embedding to be used to convert a value of the categorical field into a multidimensional vector;

[0086] Generate a hierarchical neural network including at least an auxiliary neural network and a main neural network, wherein:

[0087] The auxiliary neural network takes as input the multidimensional vector of the categorical field in the field set, wherein the multidimensional vector is obtained by transforming the value of the categorical field according to the embedding process, and wherein the auxiliary neural network

[0088] Outputting a low-dimensional representation of the multi-dimensional vector; and

[0089] The primary neural network takes as input a combination of the low dimensional representation output by the auxiliary neural network and one or more values ​​of non-categorical fields within the field set, and wherein the primary neural network outputs a binary classification for the input record;

[0090] Training the hierarchical neural network based on the labeled records to generate a trained model;

[0091] processing new records according to the trained model to determine a classification for the new records; and

[0092] Output the classification of the new record.

[0093] Clause 15. The non-transitory computer-readable medium of Clause 14, wherein the categorical fields represent qualitative values ​​and the non-categorical fields represent quantitative values.

[0094] Clause 16. A non-transitory computer-readable medium according to clause 14, wherein the hierarchical neural network is constructed to prevent the identification of correlations between values ​​of the non-categorical field and individual values ​​of the multidimensional vector during training, and to allow the identification of correlations between values ​​of the non-categorical field and individual values ​​of the low-dimensional representation during training.

[0095] Clause 17. A non-transitory computer-readable medium according to clause 14, wherein the hyperparameter identifies one or more additional classification fields within the set of fields, and wherein the hierarchical neural network includes an additional auxiliary neural network for each of the one or more additional classification fields, and the output of each additional auxiliary neural network represents an additional input to the main neural network.

[0096] Clause 18. The non-transitory computer-readable medium of clause 14, wherein the hierarchical neural network further comprises an intermediate neural network that provides the low-dimensional representation output by the auxiliary neural network to the main neural network.

[0097] Clause 19. The non-transitory computer-readable medium of clause 18, wherein the intermediate neural network further reduces the dimensionality of the low dimensional representation output by the auxiliary neural network before providing the low dimensional representation to the main neural network.

[0098] Clause 20. The non-transitory computer-readable medium of Clause 14, wherein the classification is a binary classification.

[0099] Depending on the embodiment, certain actions, events, or functions of any process or algorithm described herein may be performed in a different order, may be added, combined, or omitted entirely (e.g., not all described operations or events are necessary to practice the algorithm). In addition, in some embodiments, operations or events may be performed concurrently rather than sequentially, for example, via multithreading, interrupt handling, or on one or more computer processors or processor cores or other parallel architectures.

[0100] The various illustrative logic blocks, modules, routines, and algorithmic steps described in conjunction with the embodiments disclosed herein may be implemented as electronic hardware, or as a combination of electronic hardware and executable software. To clearly illustrate this interchangeability, various illustrative components, blocks, modules, and steps have been described above generally in terms of their functionality. Whether this functionality is implemented as hardware or as software running on hardware depends on the specific application and the design constraints imposed on the overall system. The described functionality may be implemented in different ways for each specific application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.

[0101] In addition, the various illustrative logic blocks and modules described in conjunction with the embodiments disclosed herein may be implemented or executed by a machine, such as a similarity detection system, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The similarity detection system may be or include a microprocessor, but in an alternative, the similarity detection system may be or include a controller, a microcontroller or a state machine, a combination thereof, etc. configured to estimate and transmit prediction information. The similarity detection system may include a circuit configured to process computer executable instructions. Although primarily described herein with respect to digital technology, the similarity detection system may also primarily include analog components. For example, some or all of the prediction algorithms described herein may be implemented in analog circuits or mixed analog and digital circuits. The computing environment may include any type of computer system, including but not limited to a microprocessor-based computer system, a mainframe computer, a digital signal processor, a portable computing device, a computing engine within an appliance, to name a few.

[0102] The elements of the methods, processes, routines or algorithms described in conjunction with the embodiments disclosed herein may be directly embodied in hardware, in a software module executed by a similarity detection system, or in a combination of the two. The software module may reside in a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, a hard disk, a removable disk, a CD-ROM, or any other form of non-temporary computer-readable storage medium. An exemplary storage medium may be coupled to the similarity detection system so that the similarity detection system can read information from the storage medium and can write information to the storage medium. In an alternative, the storage medium may be integrated into the similarity detection system. The similarity detection system and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In an alternative, the similarity detection system and the storage medium may reside in a user terminal as discrete components.

[0103] Conditional language used herein, such as, among others, "can", "may", "could", "for example", etc., unless otherwise specifically stated, or otherwise understood in the context of use, is generally intended to indicate that certain embodiments include certain features, elements, and / or steps, while other embodiments do not include these features, elements, and / or steps. Therefore, such conditional language is generally not intended to imply that features, elements, and / or steps are necessary for one or more embodiments in any way, or that one or more embodiments must include logic for determining whether these features, elements, and / or steps are included in any particular embodiment or will be performed in any particular embodiment with or without other input or prompts. The terms "include", "comprise", "have", etc. are synonyms and are used inclusively in an open-ended manner, and do not exclude additional elements, features, actions, operations, etc. In addition, the term "or" is used in its inclusive sense (rather than its exclusive sense), so that when, for example, used to connect a list of elements, the term "or" represents one, some, or all of the elements in the list.

[0104] Unless specifically stated otherwise, disjunctive language such as the phrase "at least one of X, Y, or Z" is otherwise understood in the context as generally used for presentation, and an item, term, etc. may be X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z). Thus, such disjunctive language is generally not intended to, and should not, imply that certain embodiments require at least one of X, at least one of Y, or at least one of Z to each be present.

[0105] Unless expressly stated otherwise, articles such as "a" or "an" should generally be interpreted as including one or more of the described items. Thus, a phrase such as "a device configured to" is intended to include one or more of the described devices. Such one or more described devices may also be collectively configured to perform the described content. For example, "a processor configured to perform statements A, B, and C" may include a first processor configured to perform statement A in cooperation with a second processor configured to perform statements B and C.

[0106] Although the above detailed description has shown, described and pointed out the novel features applied to various embodiments, it is understood that various omissions, substitutions and changes can be made to the form and details of the described devices or algorithms without departing from the spirit of the present disclosure. As can be appreciated, certain embodiments described herein can be embodied in a form that does not provide all the features and benefits set forth herein, because some features can be used or practiced separately from other features. The scope of certain embodiments disclosed herein is indicated by the appended claims rather than by the preceding description. All changes within the equivalent meaning and range of the claims will be included within their scope.

Claims

1. A computer-implemented method comprising: obtaining marked transaction records, each record corresponding to a transaction and including a value for a respective field within a set of fields associated with the transaction and marked to indicate whether the transaction is determined to be fraudulent; obtaining hyperparameters for a hierarchical neural network, the hyperparameters identifying at least a categorical field within the set of fields and an embedding process to be used to convert a value of the categorical field into a multidimensional vector; generating the multidimensional vector; Generate a hierarchical neural network including at least an auxiliary neural network and a main neural network, wherein: The auxiliary neural network takes the multi-dimensional vectors as input and outputs a low-dimensional representation of the vector for each vector; and the primary neural network taking as input a combination of the low dimensional representation output by the auxiliary neural network and one or more values ​​of non-categorical fields within the set of fields, and wherein the primary neural network outputs a binary classification indicating a likelihood that a single transaction corresponding to an input record is fraudulent; Training the hierarchical neural network based on the labeled transaction records to generate a trained model; processing new transaction records according to the trained model to determine the likelihood that the new transaction is fraudulent; and outputting the likelihood that the new transaction is fraudulent, wherein the hyperparameter identifies one or more additional categorical fields within the set of fields, and wherein the hierarchical neural network includes an additional auxiliary neural network for each of the one or more additional categorical fields, the output of each additional auxiliary neural network representing an additional input to the main neural network.

2. The computer-implemented method of claim 1 , wherein: The low-dimensional representation is represented by a set of output neurons of the auxiliary neural network.

3. The computer-implemented method of claim 1 or 2, wherein: Generating the multidimensional vector includes, for each value of the categorical field, referencing a lookup table that identifies a corresponding multidimensional vector.

4. The computer-implemented method of claim 3, wherein: The lookup table is generated by pre-applying a machine learning algorithm to a corpus of values ​​for the categorical field.

5. The computer-implemented method according to any one of claims 1 to 4, wherein: The hierarchical neural network also includes an intermediate neural network that provides the low-dimensional representation output by the auxiliary neural network to the main neural network.

6. The computer-implemented method of claim 5, wherein: The intermediate neural network further reduces the dimensionality of the low dimensional representation output by the auxiliary neural network before providing the low dimensional representation to the main neural network.

7. The computer-implemented method of claim 6, wherein: The embedding process represents at least one of word-level embedding or character-level embedding.

8. A computing system comprising: processor; as well as a data storage device comprising computer executable instructions which, when executed by the computing system, cause the computing system to: Obtaining labeled records, each record including values ​​of respective fields in a field set and being labeled for a classification of the record; obtaining hyperparameters for a hierarchical neural network, the hyperparameters identifying at least a categorical field within the set of fields and an embedding to be used to convert a value of the categorical field into a multidimensional vector; Generate a hierarchical neural network including at least an auxiliary neural network and a main neural network, wherein: The auxiliary neural network takes as input a multidimensional vector of a categorical field in the field set, the multidimensional vector being obtained by transforming the value of the categorical field according to the embedding process, and wherein the auxiliary neural network outputs a low-dimensional representation of the multidimensional vector for each multidimensional vector; and The primary neural network takes as input a combination of the low dimensional representation output by the auxiliary neural network and one or more values ​​of non-categorical fields within the field set, and wherein the primary neural network outputs a binary classification for the input record; Training the hierarchical neural network based on the labeled records to generate a trained model; processing new records according to the trained model to determine a classification for the new records; and outputting the classification of the new record, wherein the hyperparameter identifies one or more additional categorical fields within the set of fields, and wherein the hierarchical neural network includes an additional auxiliary neural network for each of the one or more additional categorical fields, the output of each additional auxiliary neural network representing an additional input to the main neural network.

9. The system according to claim 8, wherein: The categorical fields represent qualitative values ​​and the non-categorical fields represent quantitative values.

10. The system according to claim 8 or 9, wherein: The hierarchical neural network is constructed to prevent the identification of correlations between the values ​​of the non-categorical field and the respective values ​​of the multi-dimensional vector during training, and to allow the identification of correlations between the values ​​of the non-categorical field and the respective values ​​of the low-dimensional representation during training.

11. A system according to any one of claims 8 to 10, wherein: The hierarchical neural network also includes an intermediate neural network that provides the low-dimensional representation output by the auxiliary neural network to the main neural network.

12. The system according to claim 11, wherein: The intermediate neural network further reduces the dimensionality of the low dimensional representation output by the auxiliary neural network before providing the low dimensional representation to the main neural network.

13. A system according to any one of claims 8 to 12, wherein: The classification is a binary classification.