Machine learning model training using feature augmentation
By leveraging historical data from similar users to generate augmented features, the machine learning model improves fraud risk prediction accuracy for new users, addressing the challenge of limited data availability and adapting to changing behaviors.
Patent Information
- Application Number
- PCT/CN2024/106672
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-22
- Publication Date
- 2026-01-29
AI Technical Summary
Existing machine learning models struggle to accurately predict risks for new users with limited or no historical data, particularly in sectors like e-commerce and financial services, where rapid and accurate fraud detection is crucial.
Enhance machine learning model training by generating augmented features from historical data of similar existing users based on behavioral traits, using a cosine similarity metric to identify linked users and incorporating metrics like suspend seller count, suspend rate, and model scores to improve prediction accuracy.
The approach enriches the data set for new users, enabling more accurate fraud risk predictions and adaptability to evolving user behaviors, enhancing security and operational integrity in environments with limited direct historical data.
Smart Images

Figure CN2024106672_29012026_PF_FP_ABST
Abstract
Description
MACHINE LEARNING MODEL TRAINING USING FEATURE AUGMENTATIONTECHNICAL FIELD
[0001] The present disclosure generally relates to data processing using machine learning technologies. More particularly, various embodiments described herein provide for systems, methods, techniques, instruction sequences, and devices that facilitate machine learning model training using feature augmentation.BACKGROUND
[0002] Machine learning technologies have become integral to a wide range of applications, from predictive analytics to automated decision-making systems. These technologies rely on models that are trained using vast amounts of data to predict outcomes or make decisions based on new, unseen data. The effectiveness of these models largely depends on the quality and the detail of the data on which they are trained. One of the challenges in machine learning is making predictions on unexpected and / or anomaly events or patterns that may pose a risk or threat to a system. To address this, there is a growing interest in methods that can enhance the existing data features with additional, relevant information for making more accurate predictions and / or decisions.SUMMARY
[0003] A behavioral feature associated with a first user is identified. A machine learning (ML) model is used to generate an embedding vector of the first user based on the behavioral feature. A cosine similarity metric is used to identify a second user based on the embedding vector of the first user. One or more augmented features representing the first user are generated based on the one or more features of the second user. A second ML model is trained based on the one or more augmented features.BRIEF DESCRIPTION OF THE DRAWINGS
[0004] In the drawings, which are not necessarily drawn to scale, like numerals may describe similar components in different views. To easily identify the discussion of any particular element or act, the most significant digit or digits in a reference number refer to the figure number in which that element is first introduced. Some embodiments are illustrated by way of examples, and not limitations, in the accompanying figures.
[0005] FIG. 1 is a block diagram showing an example data system that includes a data management system, according to various embodiments of the present disclosure.
[0006] FIG. 2 is a block diagram illustrating an example data management system that facilitates machine learning model training for anomaly prediction using feature augmentation, according to various embodiments of the present disclosure.
[0007] FIG. 3 is a flowchart illustrating an example method for facilitating machine learning model training for anomaly prediction using feature augmentation, according to various embodiments of the present disclosure.
[0008] FIG. 4 is a flowchart illustrating an example method for facilitating machine learning model training for anomaly prediction using feature augmentation, according to various embodiments of the present disclosure.
[0009] FIG. 5 is a flowchart illustrating an example method for facilitating machine learning model training for anomaly prediction using feature augmentation n, according to various embodiments of the present disclosure.
[0010] FIG. 6 is a block diagramillustrating a representative software architecture, which may be used in conjunction with various hardware architectures herein described, according to various embodiments of the present disclosure.
[0011] FIG. 7 is a block diagram illustrating components of a machine able to read instructions from a machine storage medium and perform any one or more of the methodologies discussed herein according to various embodiments of the present disclosure.DETAILED DESCRIPTION
[0012] The description that follows includes systems, methods, techniques, instruction sequences, and computing machine program products that embody illustrative embodiments of the present disclosure. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of embodiments. It will be evident, however, to one skilled in the art that the present inventive subject matter may be practiced without these specific details.
[0013] Reference in the specification to “one embodiment” or “an embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present subject matter. Thus, the appearances of the phrase “in one embodiment” or “in an embodiment” appearing in various places throughout the specification are not necessarily all referring to the same embodiment.
[0014] For purposes of explanation, specific configurations and details are set forth in order to provide a thorough understanding of the present subject matter. However, it will be apparent to one of ordinary skill in the art that embodiments of the subject matter described may be practiced without the specific details presented herein, or in various combinations, as described herein. Furthermore, well-known features may be omitted or simplified in order not to obscure the described embodiments. Various embodiments may be given throughout this description. These are merely descriptions of specific embodiments. The scope or meaning of the claims is not limited to the embodiments given.
[0015] Various embodiments include systems, methods, and non-transitory computer-readable media that facilitate machine learning model training for anomaly prediction using feature augmentation, according to various embodiments of the present disclosure. In the field of machine learning, the challenge of accurately predicting risks (e.g., fraud risks) for new users is significant, especially when there is limited or no historical data available about these users. This scenario is common in various sectors, including e-commerce and financial services, where the ability to quickly and accurately assess risk can prevent fraud and save resources. Various embodiments address this challenge by adopting an approach that enhances the prediction capabilities of machine learning models through the augmentation of data features. This approach leverages historical data from existing users who have similar behavioral traits to the new users, thereby enriching the data set used for making predictions for new users. This approach can similarly be applied to the prediction of anomaly events and / or patterns associated with ongoing transactions.
[0016] Identification of Similar Existing Users: Existing users who exhibit similar behavioral features can be identified based on behavioral data collected for a new user. Example behavioral data can include, without limitation, page view sequences (from login to checkout) and dwell times on web pages. By identifying existing users with similar behaviors, their data can be utilized as a reference point for the new user.
[0017] Generation of Augmented Features: Augmented features can be generated for the new user. Augmented features are derived from the data of the identified existing users and can include metrics such as suspension count, suspension rate, and scores from existing fraud risk models. The augmentation of these features provides a richer set of data points that the machine learning model can use to assess the new user.
[0018] Training of the Machine Learning Model: Based on the augmented features, a tree-based machine learning model is trained to predict fraud risks.
[0019] Prediction of Fraud Risks: The trained model is then used to predict the fraud risk associated with new users. By utilizing the augmented features, the model can make more informed predictions in the absence of direct historical data about the new users.
[0020] Continuous Improvement and Adaptation: In various embodiments, the trained model is designed to continuously improve and adapt based on new data and outcomes. As more data becomes available and as user behaviors evolve, the model is retrained or adjusted to maintain its accuracy and relevance.
[0021] The approach described in various embodiments is particularly useful in environments where the rapid assessment of new users is critical to maintaining security and operational integrity. For example, in online marketplaces, new sellers can be assessed for fraud risk before they are allowed to list products. Similarly, in financial services, new account holders can be evaluated to prevent fraudulent activities. Various embodiments do not just provide a static solution but allow for dynamic updating and learning, making the system robust against evolving fraud tactics and changes in user behavior patterns.
[0022] In summary, various embodiments provide a sophisticated and practical approach to enhancing the prediction of fraud risks in new users by using machine learning models trained with augmented features. This approach addresses the significant challenge of making accurate risk assessments in situations where direct historical data is lacking. By leveraging data from similar existing users, the system boosts its predictive accuracy, offering a valuable asset in combating fraud. This approach is adaptable and can be implemented in various sectors where quick and accurate risk assessment is crucial.
[0023] In various embodiments, augmented features at least include suspend seller count, suspend rate, and the average, maximum and minimum model scores. Adding the augmented features to the training data significantly improves model performance compared with the current mass registration model.
[0024] Suspend seller count: This metric is calculated based on the number of suspended sellers out of the nearest 100 neighbors (e.g., existing similar users) . For instance, a suspend count of 5 indicates that there are 5 suspended users in the 100 similar existing users.
[0025] Suspend rate: This metric is calculated by dividing the suspend count by the total number of neighbors considered. For example, there are 1 suspended user out of a total of 100 existing users, resulting in a suspension rate of 1.00%.
[0026] Avg / Max / Min model score: These metrics represent the average, maximum, and minimum scores generated by a machine learning model (e.g., a seller registration ML model) . For instance, 100 nearest neighbors (e.g., similar existing users) are identified based on the similarity determination described herein. If the highest model score among the 100 users is 990, then the maximum model score is assigned a value of 990.
[0027] In various embodiments, the data management system identifies behavioral features associated with a user, such as a new user (e.g., the first user) . A new user can refer to an individual who has not previously engaged with the platform and is accessing it for the first time or within a relatively short timeframe. New users can be identified based on unique identifiers such as cookies, IP addresses, etc. Behavioral features can refer to data attributes or characteristics that capture patterns of behavior exhibited by individuals or entities within a specific context. These features are derived from observed behaviors or interactions and are used as input variables in machine learning models to predict outcomes, classify entities, or uncover patterns. Behavioral features, as described herein, can include page view sequence features and dwell time. A page view sequence feature refers to a data attribute that represents the sequence of pages or screens viewed by a user during a specific session or interaction with a digital platform, such as a website or application. Dwell time represents the amount of time users spend engaging (e.g., viewing) with a particular piece of content, such as a webpage, before moving on to another piece of content or another webpage.
[0028] In various embodiments, the data management system uses a machine learning (ML) model (e.g., the first ML model) to generate an embedding vector of the new user based on the behavioral features described herein. The first ML model can be a deep learning ML model, such as a Transformer model. An embedding vector can represent the user in a high-dimensional numerical space, where similar users are located closer together and dissimilar users are farther apart. An embedding vector of a user captures various aspects of the user's behavior, preferences, or characteristics.
[0029] In various embodiments, the data management system uses a cosine similarity metric to identify one or more existing similar users (e.g., the second user) based on the embedding vector of the new user. Such identified existing similar users are also referred to as linked users. A cosine similarity metric measures the cosine of the angle between two vectors in a multi-dimensional space. A higher cosine similarity value indicates greater similarity between the users, while a lower value indicates less similarity.
[0030] In various embodiments, the data management system generates one or more augmented features representing the new user (e.g., the first user) based on one or more features of the existing user (e.g., the second user) . Augmented features incorporate additional data points beyond the original features to enhance the predictive capabilities of machine learning models. Augmented features at least include suspend seller count, suspend rate, and Avg / Max / Min model scores.
[0031] In various embodiments, the data management system trains a second ML based on the augmented features representing the new user and uses the trained second ML model to predict the new user's fraud risk. The second ML model can be a tree-based ML model that corresponds to an application of the eXtreme Gradient Boosting (XGBoost) algorithm.
[0032] In various embodiments, the data management system can use the cosine similarity metric to identify a plurality of users (e.g., existing users who are similar to the new user) based on the embedding vector of the new user. Existing users can correspond to (or be associated with) historical data whereas new users do not have an association with historical data. The data management system generates augmented features that represent the new user based on a plurality of features of the plurality of users. The data management system then trains the tree-based ML model based on the augmented features to perform the fraud risk prediction of the first user.
[0033] In various embodiments, the data management system determines a suspension count based on historical data of the plurality of existing users. The suspension count represents the number of users from the plurality of users with a suspended account. The data management system calculates a suspension rate based on the suspension count and identifies a plurality of model scores based on the historical data of the plurality of users. Each model score represents a determination (e.g., fraud risk determination) of a user from the plurality of existing users. The data management system assigns one or more of the suspension count, the suspension rate, and the plurality of model scores as augmented features to the new user. The data management system trains and uses a tree-based ML model (e.g., the second ML model) to perform the fraud risk prediction of the first user based on one or more of the suspension count, the suspension rate, and the plurality of model scores.
[0034] In various embodiments, the data management system retrieves a plurality of embedding vectors of the plurality of existing users with historical data. The data management system uses the cosine similarity metric to identify the existing user (e.g., the second user) based on the embedding vector of the first user and the plurality of embedding vectors of the plurality of existing users.
[0035] Reference will now be made in detail to embodiments of the present disclosure, examples of which are illustrated in the appended drawings. The present disclosure may, however, be embodied in many different forms and should not be construed as being limited to the embodiments set forth herein.
[0036] FIG. 1 is a block diagram showing an example data system 100 that includes a data management system 122 (also referred to as system 122) , according to various embodiments of the present disclosure. By including the data management system 122, the data system 100 can facilitate machine learning model training for anomaly prediction using feature augmentation. As shown, the data system 100 includes one or more client devices 102, a server system 108, and a network 106 (e.g., Internet, wide-area-network (WAN) , local-area-network (LAN) , wireless network) that communicatively couples them together. Each client device 102 can host a number of applications, including a client software application 104. The client software application 104 can communicate data with the server system 108 via a network 106. Accordingly, the client software application 104 can communicate and exchange data with the server system 108 via network 106.
[0037] The server system 108 provides server-side functionality via the network 106 to the client software application 104. While certain functions of the data system 100 are described herein as being performed by the data management system 122 on the server system 108, it will be appreciated that the location of certain functionality within the server system 108 is a design choice. For example, it may be technically preferable to initially deploy certain technology and functionality within the server system 108, but to later migrate this technology and functionality to the client software application 104.
[0038] The server system 108 supports various services and operations that are provided to the client software application 104 by the data management system 122. Such operations include transmitting data from the data management system 122 to the client software application 104, receiving data from the client software application 104 at the data management system 122, and the data management system 122 processing data generated by the client software application 104. Data exchanges within the data system 100 may be invoked and controlled through operations of software component environments available via one or more endpoints, or functions available via one or more user interfaces of the client software application 104, which may include web-based user interfaces provided by the server system 108 for presentation at the client device 102.
[0039] With respect to the server system108, an Application Programming Interface (API) server 110 and a web server 112 is coupled to an application server 116, which hosts the data management system 122. The application server 116 is communicatively coupled to a database server 118, which facilitates access to a database 120 that stores data associated with the application server 116, including data that may be generated or used by the data management system 122.
[0040] The API server 110 receives and transmits data (e.g., API calls, commands, requests, responses, and authentication data) between the client device 102 and the application server 116. Specifically, the API server 110 provides a set of interfaces (e.g., routines and protocols) that can be called or queried by the client software application 104 in order to invoke the functionality of the application server 116. The API server 110 exposes various functions supported by the application server 116 including, without limitation, user registration; login functionality; data object operations (e.g., generating, storing, retrieving, encrypting, decrypting, transferring, access rights, licensing) ; and / or user communications.
[0041] The data management system 122 operates to facilitate machine learning model training for anomaly prediction using feature augmentation, according to various embodiments of the present disclosure. The server system 108, or the data management system 122 may extract user data from one or more third-party platforms (e.g., third-party social media platforms) . The extracted data may be open-source poster data associated with targeted influencers on the one or more third-party platforms 124 and may include user profile data, activity data, and media posted (either created and / or shared) by the one or more influencers. The media (or media data) include text, image, video, audio, and metadata. Example metadata may include hashtags and labels.
[0042] Through one or more web-based interfaces (e.g., web-based user interfaces) , the web server 112 can support various functionality of the data management system 122 of the application server 116.
[0043] FIG. 2 is a block diagram illustrating an example data management system 200 that facilitates machine learning model training for anomaly prediction using feature augmentation, according to various embodiments of the present disclosure. For some embodiments, the data management system 200 represents an example of the data management system 122 described with respect to FIG. 1. As shown, the data management system 200 comprises a behavioral feature identifying component 210, An embedding vector generating component 220, a similar user identifying component 230, an augmented feature generating component 240, a model training component 250, and a risk prediction component 260. According to various embodiments, one or more of the behavioral feature identifying component 210, the embedding vector generating component 220, the similar user identifying component 230, the augmented feature generating component 240, the model training component 250, and the risk prediction component 260 are implemented by one or more hardware processors 202. Data generated by one or more of the behavioral feature identifying component 210, the embedding vector generating component 220, the similar user identifying component 230, the augmented feature generating component 240, the model training component 250, and the risk prediction component 260 may be stored in a database (or datastore) 270 of the data management system 200.
[0044] The behavioral feature identifying component 210 is configured to identify behavioral features associated with a user, such as a new user (e.g., the first user) . A new user can refer to an individual who has not previously engaged with the platform and is accessing it for the first time or within a relatively short timeframe. New users can be identified based on unique identifiers such as cookies, IP addresses, etc. Behavioral features can refer to data attributes or characteristics that capture patterns of behavior exhibited by individuals or entities within a specific context. Behavioral features can include page view sequence features and dwell time.
[0045] The embedding vector generating component 220 is configured to use a machine learning (ML) model (e.g., the first ML model) to generate an embedding vector of the new user based on the behavioral features described herein. The first ML model can be a deep learning ML model, such as a Transformer model.
[0046] The similar user identifying component 230 is configured to use a cosine similarity metric to identify one or more existing similar users (e.g., the second user) based on the embedding vector of the new user. Such identified existing similar users are also referred to as linked users.
[0047] The augmented feature generating component 240 is configured to generate one or more augmented features representing the new user (e.g., the first user) based on one or more features of the existing user (e.g., the second user) . Augmented features incorporate additional data points beyond the original features to enhance the predictive capabilities of machine learning models. Augmented features at least include suspend seller count, suspend rate, and average, maximum, and minimum model scores. The model training component 250 is configured to train a second ML based on the augmented features representing the new user and uses the trained second ML model to predict the new user's fraud risk. The second ML model can be a tree-based ML model that corresponds to an application of the eXtreme Gradient Boosting (XGBoost) algorithm.
[0048] The risk prediction component 260 is configured to use the trained tree-based ML model to predict fraud risks in new users.
[0049] FIG. 3 is a flowchart illustrating an example method 300 for facilitating machine learning model training for anomaly prediction using feature augmentation, according to various embodiments of the present disclosure. It will be understood that example methods described herein may be performed by a machine in accordance with some embodiments. For example, method 300 can be performed by the data management system 122 described with respect to FIG. 1, the data management system 200 described with respect to FIG. 2, or individual components thereof. An operation of various methods described herein may be performed by one or more hardware processors (e.g., central processing units or graphics processing units) of a computing device (e.g., a desktop, server, laptop, mobile phone, tablet, etc. ) , which may be part of a computing system based on a cloud architecture. Example methods described herein may also be implemented in the form of executable instructions stored on a machine-readable medium or in the form of electronic circuitry. For instance, the operations of method 300 may be represented by executable instructions that, when executed by a processor of a computing device, cause the computing device to perform method 300. Depending on the embodiment, an operation of an example method described herein may be repeated in different ways or involve intervening operations not shown. Though the operations of example methods may be depicted and described in a certain order, the order in which the operations are performed may vary among embodiments, including performing certain operations in parallel.
[0050] At operation 302, a processor identifies behavioral features associated with a user, such as a new user (e.g., the first user) . A new user can refer to an individual who has not previously engaged with the platform and is accessing it for the first time or within a relatively short timeframe. New users can be identified based on unique identifiers such as cookies, IP addresses, etc. Behavioral features can refer to data attributes or characteristics that capture patterns of behavior exhibited by individuals or entities within a specific context. These features are derived from observed behaviors or interactions and are used as input variables in machine learning models to predict outcomes, classify entities, or uncover patterns. Behavioral features, as described herein, can include page view sequence features and dwell time.
[0051] At operation 304, a processor uses a machine learning (ML) model (e.g., the first ML model) to generate an embedding vector of the new user based on the behavioral features described herein. The first ML model can be a deep learning ML model, such as a Transformer model. An embedding vector can represent the user in a high-dimensional numerical space, where similar users are located closer together and dissimilar users are farther apart. An embedding vector of a user captures various aspects of the user's behavior, preferences, or characteristics.
[0052] At operation 306, a processor uses a cosine similarity metric to identify one or more existing similar users (e.g., the second user) based on the embedding vector of the new user. Such identified existing similar users are also referred to as linked users. A cosine similarity metric measures the cosine of the angle between two vectors in a multi-dimensional space. A higher cosine similarity value indicates greater similarity between the users, while a lower value indicates less similarity.
[0053] At operation 308, a processor generates one or more augmented features representing the new user (e.g., the first user) based on one or more features of the existing user (e.g., the second user) . Augmented features incorporate additional data points beyond the original features to enhance the predictive capabilities of machine learning models. Augmented features at least include suspend seller count, suspend rate, and average, maximum, and minimum model scores.
[0054] At operation 310, a processor trains a second ML based on the augmented features representing the new user and uses the trained second ML model to predict the new user's fraud risk. The second ML model can be a tree-based ML model that corresponds to an application of the eXtreme Gradient Boosting (XGBoost) algorithm.
[0055] At operation 312, a processor uses the tree-based ML model to perform risk (e.g., fraud risk) predictions for new users. This approach addresses the significant challenge of making accurate risk assessments in situations where direct historical data is lacking. By leveraging data from similar existing users, the system boosts its predictive accuracy, offering a valuable asset in combating fraud.
[0056] Though not illustrated, method 300 can include an operation where a graphical user interface is displayed (or caused to be displayed) by the hardware processor. For instance, the operation can cause a client device (e.g., the client device 102 communicatively coupled to the data management system 122) to display the graphical user interface. This operation for displaying the graphical user interface can be separate from operations 302 through 312 or, alternatively, form part of one or more of operations 302 through 312.
[0057] FIG. 4 is a flowchart illustrating an example method 400 for facilitating machine learning model training for anomaly prediction using feature augmentation, according to various embodiments of the present disclosure. It will be understood that example methods described herein may be performed by a machine in accordance with some embodiments. For example, method 400 can be performed by the data management system 122 described with respect to FIG. 1, the data management system 200 described with respect to FIG. 2, or individual components thereof. An operation of various methods described herein may be performed by one or more hardware processors (e.g., central processing units or graphics processing units) of a computing device (e.g., a desktop, server, laptop, mobile phone, tablet, etc. ) , which may be part of a computing system based on a cloud architecture. Example methods described herein may also be implemented in the form of executable instructions stored on a machine-readable medium or in the form of electronic circuitry. For instance, the operations of method 400 may be represented by executable instructions that, when executed by a processor of a computing device, cause the computing device to perform method 400. Depending on the embodiment, an operation of an example method described herein may be repeated in different ways or involve intervening operations not shown. Though the operations of example methods may be depicted and described in a certain order, the order in which the operations are performed may vary among embodiments, including performing certain operations in parallel. Operations in method 400 can be performed dependently or independently from operations in method 300.
[0058] At operation 402, a processor uses the cosine similarity metric to identify a plurality of users (e.g., existing users similar to the new user) based on the embedding vector of the new user. Existing users can correspond to (or be associated with) historical data whereas new users do not have an association with direct historical data.
[0059] At operation 404, a processor generates augmented features that represent the new user based on a plurality of features of the plurality of users. Features of an existing user can include, without limitation, account features (e.g., name, email, password) , profile features (e.g., shipping address, payment methods, order history) , product browsing features (e.g., browsing history) , shopping cart features (e.g., shopping cart history) , checkout process features (e.g., selected shipping options, applied discounts or coupons, payment information entries) , wish list features, review and rating features, etc.
[0060] At operation 406, a processor trains a second ML based on the augmented features representing the new user and uses the trained second ML model to predict the new user's fraud risk. The second ML model can be a tree-based ML model that corresponds to an application of the eXtreme Gradient Boosting (XGBoost) algorithm.
[0061] Though not illustrated, method 400 can include an operation where a graphical user interface can be displayed (or caused to be displayed) by the hardware processor. For instance, the operation can cause a client device (e.g., the client device 102 communicatively coupled to the data management system 122) to display the graphical user interface. This operation for displaying the graphical user interface can be separate from operations 402 through 406 or, alternatively, form part of one or more of operations 402 through 406.
[0062] FIG. 5 is a flowchart illustrating an example method 500 for facilitating machine learning model training for anomaly prediction using feature augmentation, according to various embodiments of the present disclosure. It will be understood that example methods described herein may be performed by a machine in accordance with some embodiments. For example, method 500 can be performed by the data management system 122 described with respect to FIG. 1, the data management system 200 described with respect to FIG. 2, or individual components thereof. An operation of various methods described herein may be performed by one or more hardware processors (e.g., central processing units or graphics processing units) of a computing device (e.g., a desktop, server, laptop, mobile phone, tablet, etc. ) , which may be part of a computing system based on a cloud architecture. Example methods described herein may also be implemented in the form of executable instructions stored on a machine-readable medium or in the form of electronic circuitry. For instance, the operations of method 500 may be represented by executable instructions that, when executed by a processor of a computing device, cause the computing device to perform method 500. Depending on the embodiment, an operation of an example method described herein may be repeated in different ways or involve intervening operations not shown. Though the operations of example methods may be depicted and described in a certain order, the order in which the operations are performed may vary among embodiments, including performing certain operations in parallel. Operations in method 500 can be performed dependently or independently from operations in method 300 and method 400.
[0063] At operation 502, a processor determines a suspension count based on historical data of the plurality of existing users. The suspension count represents the number of users from the plurality of users with a suspended account.
[0064] At operation 504, a processor calculates a suspension rate based on the suspension count. A suspension rate is calculated by dividing the suspend count by the total number of existing users considered. For example, there are 10 suspended users out of a total of 1000 existing users, resulting in a suspension rate of 1.00%.
[0065] At operation 506, a processor identifies a plurality of model scores (e.g., average, maximum, and minimum model scores) based on the historical data of the plurality of users. Each model score represents a determination (e.g., fraud risk determination) of a user from the plurality of existing users. For instance, 100 nearest neighbors (e.g., similar existing users) are identified based on the similarity determination described herein. If the highest model score among the 100 users is 990, then the max score is assigned a value of 990.
[0066] At operation 508, a processor assigns one or more of the suspension count, the suspension rate, and the plurality of model scores as augmented features to the new user. When dealing with a new user without historical data, augmenting features can be used to create a provisional user profile, which is used to assess the risks associated with the new user.
[0067] At operation 510, a processor trains a tree-based ML model (e.g., the second ML model) based on one or more of the suspension count, the suspension rate, and the plurality of model scores and performs fraud risk prediction for new users, including the first user.
[0068] Though not illustrated, method 500 can include an operation where a graphical user interface can be displayed (or caused to be displayed) by the hardware processor. For instance, the operation can cause a client device (e.g., the client device 102 communicatively coupled to the data management system 122) to display the graphical user interface. This operation for displaying the graphical user interface can be separate from operations 502 through 510 or, alternatively, form part of one or more of operations 502 through 510.
[0069] Example 1 is a system comprising: one or more hardware processors; and at least one machine-storage medium for storing instructions that, when executedby the one or more hardware processors, cause the one or more hardware processors to perform operations comprising: identifying a behavioral feature associated with a first user; generating, using a first machine learning (ML) model, an embedding vector of the first user based on the behavioral feature; identifying, using a cosine similarity metric, a second user based on the embedding vector of the first user; generating one or more augmented features representing the first user based on one or more features of the second user; and training the second ML model based on the one or more augmented features.
[0070] In Example 2, the subject matter of Example 1 includes, wherein the operations comprise: performing, using the trained second ML model, a fraud risk prediction of the first user.
[0071] In Example 3, the subject matter of Examples 1-2 includes, wherein the operations comprise: identifying, using the cosine similarity metric, a plurality of users based on the embedding vector of the first user, the plurality of users including the second user; generating the one or more augmented features that represent the first user based on a plurality of features of the plurality of users; and training the second ML model based on the one or more augmented features representing the first user.
[0072] In Example 4, the subject matter of Example 3 includes, wherein the plurality of users comprises a plurality of existing users with historical data, and wherein the first user corresponds to a new user without historical data.
[0073] In Example 5, the subject matter of Examples 3-4 includes, wherein the operations comprise: determining a suspension count based on historical data of the plurality of users, the suspension count representing a number of users from the plurality of users with a suspended account; calculating a suspension rate based on the suspension count; and identifying a plurality of model scores based on the historical data of the plurality of users, each model score representing a fraud risk determination of a user from the plurality of users.
[0074] In Example 6, the subject matter of Example 5 includes, wherein the operations comprise: assigning one or more of the suspension count, the suspension rate, and the plurality of model scores as the one or more augmented features to the first user; and training the second ML model based on one or more of the suspension count, the suspension rate, and the plurality of model scores.
[0075] In Example 7, the subject matter of Examples 1-6 includes, wherein the behavioral feature comprises a page view sequence feature, the page view sequence feature corresponding to a plurality of webpages, each webpage being associated with a dwell time; and wherein the dwell time represents an amount of time taken by the first user to view a webpage before moving on to another webpage.
[0076] In Example 8, the subject matter of Examples 1-7 includes, wherein the operations comprise: retrieving a plurality of embedding vectors of a plurality of users, the plurality of users including the second user, each user in the plurality of users being associated with historical data; and identifying, using the cosine similarity metric, the second user based on the embedding vector of the first user and the plurality of embedding vectors of the plurality of users.
[0077] In Example 9, the subject matter of Examples 1-8 includes, wherein the second ML model comprises a tree-based ML model, and wherein the tree-based ML model corresponds to an application of eXtreme Gradient Boosting (XGBoost) algorithm.
[0078] In Example 10, the subject matter of Examples 1-9 includes, wherein the first ML model comprises a deep learning ML model.
[0079] Example 11 is a method comprising: identifying, by at least one hardware processor, a behavioral feature associated with a first user; generating, using a first machine learning (ML) model, an embedding vector of the first user based on the behavioral feature; identifying, using a cosine similarity metric, a second user based on the embedding vector of the first user; generating one or more augmented features representing the first user based on one or more features of the second user; and training the second ML model based on the one or more augmented features.
[0080] In Example 12, the subject matter of Example 11 includes, performing, using the trained second ML model, a fraud risk prediction of the first user.
[0081] In Example 13, the subject matter of Examples 11-12 includes, identifying, using the cosine similarity metric, a plurality of users based on the embedding vector of the first user, the plurality of users including the second user; generating the one or more augmented features that represent the first user based on a plurality of features of the plurality of users; and training the second ML model based on the one or more augmented features representing the first user.
[0082] In Example 14, the subject matter of Example 13 includes, wherein the plurality of users comprises a plurality of existing users with historical data, and wherein the first user corresponds to a new user without historical data.
[0083] In Example 15, the subject matter of Examples 13-14 includes, determining a suspension count based on historical data of the plurality of users, the suspension count representing a number of users from the plurality of users with a suspended account; calculating a suspension rate based on the suspension count; and identifying a plurality of model scores based on the historical data of the plurality of users, each model score representing a fraud risk determination of a user from the plurality of users.
[0084] In Example 16, the subject matter of Example 15 includes, assigning one or more of the suspension count, the suspension rate, and the plurality of model scores as the one or more augmented features to the first user; and training the second ML model based on one or more of the suspension count, the suspension rate, and the plurality of model scores.
[0085] In Example 17, the subject matter of Examples 11-16 includes, wherein the behavioral feature comprises a page view sequence feature, the page view sequence feature corresponding to a plurality of webpages, each webpage being associated with a dwell time; and wherein the dwell time represents an amount of time taken by the first user to view a webpage before moving on to another webpage.
[0086] In Example 18, the subject matter of Examples 11-17 includes, retrieving a plurality of embedding vectors of a plurality of users, the plurality of users including the second user, each user in the plurality of users being associated with historical data; and identifying, using the cosine similarity metric, the second user based on the embedding vector of the first user and the plurality of embedding vectors of the plurality of users.
[0087] In Example 19, the subject matter of Examples 11-18 includes, wherein the second ML model comprises a tree-based ML model, wherein the tree-based ML model corresponds to an application of eXtreme Gradient Boosting (XGBoost) algorithmand wherein the first ML model comprises a deep learning ML model.
[0088] Example 20 is a machine-storage medium for storing instructions that, when executed by one or more hardware processors, cause the one or more hardware processors to perform operations comprising: identifying a behavioral feature associated with a first user; generating, using a first machine learning (ML) model, an embedding vector of the first user based on the behavioral feature; identifying, using a cosine similarity metric, a second user based on the embedding vector of the first user; generating one or more augmented features representing the first user based on one or more features of the second user; and training the second ML model based on the one or more augmented features.
[0089] FIG. 6 is a block diagram illustrating an example of a software architecture 602 that may be installed on a machine, according to some example embodiments. FIG. 6 is merely a non-limiting example of a software architecture, and it will be appreciated that many other architectures may be implemented to facilitate the functionality described herein. The software architecture 602 may be executing on hardware such as a machine 700 of FIG. 7 that includes, among other things, processors 710, memory 730, and input / output (I / O) components 750. A representative hardware layer 604 is illustrated and can represent, for example, the machine 700 of FIG. 7. The representative hardware layer 604 comprises one or more processing units 606 having associated executable instructions 608. The executable instructions 608 represent the executable instructions of the software architecture 602. The hardware layer 604 also includes memory or storage modules (storage components) 610, which also have the executable instructions 608. The hardware layer 604 may also comprise other hardware 612, which represents any other hardware of the hardware layer 604, such as the other hardware illustrated as part of the machine 700.
[0090] In the example architecture of FIG. 6, the software architecture 602 may be conceptualized as a stack of layers, where each layer provides particular functionality. For example, the software architecture 602 may include layers such as an operating system 614, libraries 616, frameworks / middleware 618, applications 620, and a presentation layer 644. Operationally, the applications 620 or other components within the layers may invoke API calls 624 through the software stack and receive a response, returned values, and so forth (illustrated as messages 626) in response to the API calls 624. The layers illustrated are representative in nature, and not all software architectures have all layers. For example, some mobile or special-purpose operating systems may not provide a frameworks / middleware 618 layer, while others may provide such a layer. Other software architectures may include additional or different layers.
[0091] The operating system 614 may manage hardware resources and provide common services. The operating system614 may include, for example, a kernel 628, services 630, and drivers 632. The kernel 628 may act as an abstraction layer between the hardware and the other software layers. For example, the kernel 628 may be responsible for memory management, processor management (e.g., scheduling) , component management, networking, security settings, and so on. The services 630 may provide other common services for the other software layers. The drivers 632 may be responsible for controlling or interfacing with the underlying hardware. For instance, the drivers 632 may include display drivers, camera drivers, drivers, flash memory drivers, serial communication drivers (e.g., Universal Serial Bus (USB) drivers) , drivers, audio drivers, power management drivers, and so forth depending on the hardware configuration.
[0092] The libraries 616 may provide a common infrastructure that may be utilized by the applications 620 and / or other components and / or layers. The libraries 616 typically provide functionality that allows other software components / modules to perform tasks in an easier fashion than by interfacing directly with the underlying operating system 614 functionality (e.g., kernel 628, services 630, or drivers 632) . The libraries 616 may include system libraries 634 (e.g., C standard library) that may provide functions such as memory allocation functions, string manipulation functions, mathematic functions, and the like. In addition, the libraries 616 may include API libraries 636 such as media libraries (e.g., libraries to support presentation and manipulation of various media formats such as MPEG4, H. 264, MP3, AAC, AMR, JPG, and PNG) , graphics libraries (e.g., an OpenGL framework that may be used to render 2D and 3D graphic content on a display) , database libraries (e.g., SQLite that may provide various relational database functions) , web libraries (e.g., WebKit that may provide web browsing functionality) , and the like. The libraries 616 may also include a wide variety of other libraries 638 to provide many other APIs to the applications 620 and other software components / modules.
[0093] The frameworks 618 (also sometimes referred to as middleware) may provide a higher-level common infrastructure that may be utilized by the applications 620 or other software components / modules. For example, the frameworks 618 may provide various graphical user interface functions, high-level resource management, high-level location services, and so forth. The frameworks 618 may provide a broad spectrum of other APIs that may be utilized by the applications 620 and / or other software components / modules, some of which may be specific to a particular operating system or platform.
[0094] The applications 620 include built-in applications 640 and / or third-party applications 642. Examples of representative built-in applications 640 may include, but are not limited to, a home application, a contacts application, a browser application, a book reader application, a location application, a media application, a messaging application, or a game application.
[0095] The third-party applications 642 may include any of the built-in applications 640, as well as a broad assortment of other applications. In a specific example, the third-party applications 642 (e.g., an application developed using the AndroidTM or iOSTM software development kit (SDK) by an entity other than the vendor of the particular platform) may be mobile software running on a mobile operating system such as iOSTM, AndroidTM, or other mobile operating systems. In this example, the third-party applications 642 may invoke the API calls 624 provided by the mobile operating system such as the operating system 614 to facilitate functionality described herein.
[0096] The applications 620 may utilize built-in operating system functions (e.g., kernel 628, services 630, or drivers 632) , libraries (e.g., system libraries 634, API libraries 636, and other libraries 638) , or frameworks / middleware 618 to create user interfaces to interact with users of the system. Alternatively, or additionally, in some systems, interactions with a user may occur through a presentation layer, such as the presentation layer 644. In these systems, the application / module “logic” can be separated from the aspects of the application / module that interact with the user.
[0097] Some software architectures utilize virtual machines. In the example of FIG. 6, this is illustrated by a virtual machine 648. The virtual machine 648 creates a software environment where applications / modules can execute as if they were executing on a hardware machine (e.g., the machine 700 of FIG. 7) . The virtual machine 648 is hosted by a host operating system (e.g., the operating system 614) and typically, although not always, has a virtual machine monitor 646, which manages the operation of the virtual machine 648 as well as the interface with the host operating system (e.g., the operating system 614) . A software architecture executes within the virtual machine 648, such as an operating system 650, libraries 652, frameworks 654, applications 656, or a presentation layer 658. These layers of software architecture executing within the virtual machine 648 can be the same as corresponding layers previously described or may be different.
[0098] FIG. 7 illustrates a diagrammatic representation of a machine 700 in the form of a computer system within which a set of instructions may be executed for causing the machine 700 to perform any one or more of the methodologies discussed herein, according to an embodiment. Specifically, FIG. 7 shows a diagrammatic representation of the machine 700 in the example form of a computer system, within which instructions 716 (e.g., software, a program, an application, an applet, an app, or other executable code) for causing the machine 700 to perform any one or more of the methodologies discussed herein may be executed. For example, the instructions 716 may cause the machine 700 to execute the method 300 described above with respect to FIG. 3, the method 400 described above with respect to FIG. 4, and the method 500 described above with respect to FIG. 5. The instructions 716 transform the general, non-programmed machine 700 into a particular machine 700 programmed to carry out the described and illustrated functions in the manner described. In alternative embodiments, the machine 700 operates as a standalone device or may be coupled (e.g., networked) to other machines. In a networked deployment, the machine 700 may operate in the capacity of a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 700 may comprise, but not be limited to, a server computer, a client computer, a personal computer (PC) , a tablet computer, a laptop computer, a netbook, a personal digital assistant (PDA) , an entertainment media system, a cellular telephone, a smart phone, a mobile device, or any machine capable of executing the instructions 716, sequentially or otherwise, that specify actions to be taken by the machine 700. Further, while only a single machine 700 is illustrated, the term “machine” shall also be taken to include a collection of machines 700 that individually or jointly execute the instructions 716 to perform any one or more of the methodologies discussed herein.
[0099] The machine 700 may include processors 710, memory 730, and I / O components 750, which may be configured to communicate with each other such as via a bus 702. In an embodiment, the processors 710 (e.g., a hardware processor, such as a central processing unit (CPU) , a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU) , a digital signal processor (DSP) , an application-specific integrated circuit (ASIC) , a radio-frequency integrated circuit (RFIC) , another processor, or any suitable combination thereof) may include, for example, a processor 712 and a processor 714 that may execute the instructions 716. The term “processor” is intended to include multi-core processors that may comprise two or more independent processors (sometimes referred to as “cores” ) that may execute instructions contemporaneously. Although FIG. 7 shows multiple processors 710, the machine 700 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor) , multiple processors with a single core, multiple processors with multiples cores, or any combination thereof.
[0100] The memory 730 may include a main memory 732, a static memory 734, and a storage unit 736 including machine-readable medium 738, each accessible to the processors 710 such as via the bus 702. The main memory 732, the static memory 734, and the storage unit 736 store the instructions 716 embodying any one or more of the methodologies or functions described herein. The instructions 716 may also reside, completely or partially, within the main memory 732, within the static memory 734, within the storage unit 736, within at least one of the processors 710 (e.g., within the processor’s cache memory) , or any suitable combination thereof, during execution thereof by the machine 700.
[0101] The I / O components 750 may include a wide variety of components to receive input, provide output, produce output, transmit information, exchange information, capture measurements, and so on. The specific I / O components 750 that are included in a particular machine will depend on the type of machine. For example, portable machines such as mobile phones will likely include a touch input device or other such input mechanisms, while a headless server machine will likely not include such a touch input device. It will be appreciated that the I / O components 750 may include many other components that are not shown in FIG. 7. The I / O components 750 are grouped according to functionality merely for simplifying the following discussion, and the grouping is in no way limiting. In some examples, the I / O components 750 may include output components 752 and input components 754. The output components 752 may include visual components (e.g., a display such as a plasma display panel (PDP) , a light-emitting diode (LED) display, a liquid crystal display (LCD) , a projector, or a cathode ray tube (CRT) ) , acoustic components (e.g., speakers) , haptic components (e.g., a vibratory motor, resistance mechanisms) , other signal generators, and so forth. The input components 754 may include alphanumeric input components (e.g., a keyboard, a touch screen configured to receive alphanumeric input, a photo-optical keyboard, or other alphanumeric input components) , point-based input components (e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or another pointing instrument) , tactile input components (e.g., a physical button, a touch screen that provides location and / or force of touches or touch gestures, or other tactile input components) , audio input components (e.g., a microphone) , and the like.
[0102] In further embodiments, the I / O components 750 may include biometric components 756, motion components 758, environmental components 760, or position components 762, among a wide array of other components. The motion components 758 may include acceleration sensor components (e.g., accelerometer) , gravitation sensor components, rotation sensor components (e.g., gyroscope) , and so forth. The environmental components 760 may include, for example, illumination sensor components (e.g., photometer) , temperature sensor components (e.g., one or more thermometers that detect ambient temperature) , humidity sensor components, pressure sensor components (e.g., barometer) , acoustic sensor components (e.g., one or more microphones that detect background noise) , proximity sensor components (e.g., infrared sensors that detect nearby objects) , gas sensors (e.g., gas detection sensors to detect concentrations of hazardous gases for safety or to measure pollutants in the atmosphere) , or other components that may provide indications, measurements, or signals corresponding to a surrounding physical environment. The position components 762 may include location sensor components (e.g., a Global Positioning System (GPS) receiver component) , altitude sensor components (e.g., altimeters or barometers that detect air pressure from which altitude may be derived) , orientation sensor components (e.g., magnetometers) , and the like.
[0103] Communication may be implemented using a wide variety of technologies. The I / O components 750 may include communication components 764 operable to couple the machine 700 to a network 780 or devices 770 via a coupling 782 and a coupling 772, respectively. For example, the communication components 764 may include a network interface component or another suitable device to interface with the network 780. In further examples, the communication components 764 may include wired communication components, wireless communication components, cellular communication components, near field communication (NFC) components, components (e.g., Low Energy) , components, and other communication components to provide communication via other modalities. The devices 770 may be another machine or any of a wide variety of peripheral devices (e.g., a peripheral device coupled via a USB) .
[0104] Moreover, the communication components 764 may detect identifiers or include components operable to detect identifiers. For example, the communication components 764 may include radio frequency identification (RFID) tag reader components, NFC smart tag detection components, optical reader components (e.g., an optical sensor to detect one-dimensional bar codes such as Universal Product Code (UPC) bar code, multi-dimensional bar codes such as Quick Response (QR) code, Aztec code, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D bar code, and other optical codes) , or acoustic detection components (e.g., microphones to identify tagged audio signals) . In addition, a variety of information may be derived via the communication components 764, such as location via Internet Protocol (IP) geolocation, location via signal triangulation, location via detecting an NFC beacon signal that may indicate a particular location, and so forth.
[0105] Certain embodiments are described herein as including logic or a number of components, components, elements, or mechanisms. Such components can constitute either software components (e.g., code embodied on a machine-readable medium or in a transmission signal) or hardware components. A “hardware component” is a tangible unit capable of performing certain operations and can be configured or arranged in a certain physical manner. In various example embodiments, one or more computer systems (e.g., a standalone computer system, a client computer system, or a server computer system) or one or more hardware components of a computer system (e.g., a processor or a group of processors) are configured by software (e.g., an application or application portion) as a hardware component that operates to perform certain operations as described herein.
[0106] In some examples, a hardware component is implemented mechanically, electronically, or any suitable combination thereof. For example, a hardware component can include dedicated circuitry or logic that is permanently configured to perform certain operations. For example, a hardware component can be a special-purpose processor, such as a field-programmable gate array (FPGA) or an ASIC. A hardware component may also include programmable logic or circuitry that is temporarily configured by software to perform certain operations. For example, a hardware component can include software encompassed within a general-purpose processor or other programmable processor. It will be appreciated that the decision to implement a hardware component mechanically, in dedicated and permanently configured circuitry, or in temporarily configured circuitry (e.g., configured by software) can be driven by cost and time considerations.
[0107] Accordingly, the phrase “component” should be understood to encompass a tangible entity, be that an entity that is physically constructed, permanently configured (e.g., hardwired) , or temporarily configured (e.g., programmed) to operate in a certain manner or to perform certain operations described herein. Considering embodiments in which hardware components are temporarily configured (e.g., programmed) , each of the hardware components need not be configured or instantiated at any one instance in time. For example, where a hardware component comprises a general-purpose processor configured by software to become a special-purpose processor, the general-purpose processor may be configured as respectively different special-purpose processors (e.g., comprising different hardware components) at different times. Software can accordingly configure a particular processor or processors, for example, to constitute a particular hardware component at one instance of time and to constitute a different hardware component at a different instance of time.
[0108] Hardware components can provide information to, and receive information from, other hardware components. Accordingly, the described hardware components can be regarded as being communicatively coupled. Where multiple hardware components exist contemporaneously, communications can be achieved through signal transmission (e.g., over appropriate circuits and buses) between or among two or more of the hardware components. In embodiments in which multiple hardware components are configured or instantiated at different times, communications between or among such hardware components may be achieved, for example, through the storage and retrieval of information in memory structures to which the multiple hardware components have access. For example, one hardware component performs an operation and stores the output of that operation in a memory device to which it is communicatively coupled. A further hardware component can then, at a later time, access the memory device to retrieve and process the stored output. Hardware components can also initiate communications with input or output devices, and can operate on a resource (e.g., a collection of information) .
[0109] The various operations of example methods described herein can be performed, at least partially, by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors constitute processor-implemented components that operate to perform one or more operations or functions described herein. As used herein, “processor-implemented component” refers to a hardware component implemented using one or more processors.
[0110] Similarly, the methods described herein can be at least partially processor-implemented, with a particular processor or processors being an example of hardware. For example, at least some of the operations of a method can be performed by one or more processors or processor-implemented components. Moreover, the one or more processors may also operate to support performance of the relevant operations in a “cloud computing” environment or as a “software as a service” (SaaS) . For example, at least some of the operations may be performed by a group of computers (as examples of machines 700 including processors 710) , with these operations being accessible via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., an API) . In certain embodiments, for example, a client device may relay or operate in communication with cloud computing systems and may access circuit design information in a cloud environment.
[0111] The performance of certain of the operations may be distributed among the processors, not only residing within a single machine 700, but deployed across a number of machines 700. In some example embodiments, the processors 710 or processor-implemented components are located in a single geographic location (e.g., within a home environment, an office environment, or a server farm) . In other example embodiments, the processors or processor-implemented components are distributed across a number of geographic locations.
[0112] The various memories (i.e., 730, 732, 734, and / or the memory of the processor (s) 710) and / or the storage unit 736 may store one or more sets of instructions 716 and data structures (e.g., software) embodying or utilized by any one or more of the methodologies or functions described herein. These instructions (e.g., the instructions 716) , when executed by the processor (s) 710, cause various operations to implement the disclosed embodiments.
[0113] As used herein, the terms “machine-storage medium, ” “device-storage medium, ” and “computer-storage medium” mean the same thing and may be used interchangeably. The terms refer to a single or multiple storage devices and / or media (e.g., a centralized or distributed database, and / or associated caches and servers) that store executable instructions 716 and / or data. The terms shall accordingly be taken to include, but not be limited to, solid-state memories, and optical and magnetic media, including memory internal or external to processors. Specific examples of machine-storage media, computer-storage media and / or device-storage media include non-volatile memory, including by way of example semiconductor memory devices, e.g., erasable programmable read-only memory (EPROM) , electrically erasable programmable read-only memory (EEPROM) , FPGA, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The terms “machine-storage media, ” “computer-storage media, ” and “device-storage media” specifically exclude carrier waves, modulated data signals, and other such media, at least some of which are covered under the term “signal medium” discussed below.
[0114] In some examples, one or more portions of the network 780 may be an ad hoc network, an intranet, an extranet, a virtual private network (VPN) , a LAN, a wireless LAN (WLAN) , a WAN, a wireless WAN (WWAN) , a metropolitan-area network (MAN) , the Internet, a portion of the Internet, a portion of the public switched telephone network (PSTN) , a plain old telephone service (POTS) network, a cellular telephone network, a wireless network, a network, another type of network, or a combination of two or more such networks. For example, the network 780 or a portion of the network 780 may include a wireless or cellular network, and the coupling 782 may be a Code Division Multiple Access (CDMA) connection, a Global System for Mobile communications (GSM) connection, or another type of cellular or wireless coupling. In this example, the coupling 782 may implement any of a variety of types of data transfer technology, such as Single Carrier Radio Transmission Technology (1xRTT) , Evolution-Data Optimized (EVDO) technology, General Packet Radio Service (GPRS) technology, Enhanced Data rates for GSM Evolution (EDGE) technology, third Generation Partnership Project (3GPP) including 3G, fourth generation wireless (4G) networks, Universal Mobile Telecommunications System (UMTS) , High-Speed Packet Access (HSPA) , Worldwide Interoperability for Microwave Access (WiMAX) , Long-Term Evolution (LTE) standard, others defined by various standard-setting organizations, other long-range protocols, or other data transfer technology.
[0115] The instructions may be transmitted or received over the network using a transmission medium via a network interface device (e.g., a network interface component included in the communication components) and utilizing any one of a number of well-known transfer protocols (e.g., hypertext transfer protocol (HTTP) ) . Similarly, the instructions may be transmitted or received using a transmission medium via the coupling (e.g., a peer-to-peer coupling) to the devices 770. The terms “transmission medium” and “signal medium” mean the same thing and may be used interchangeably in this disclosure. The terms “transmission medium” and “signal medium” shall be taken to include any intangible medium that is capable of storing, encoding, or carrying the instructions for execution by the machine, and include digital or analog communications signals or other intangible media to facilitate communication of such software. Hence, the terms “transmission medium” and “signal medium” shall be taken to include any form of modulated data signal, carrier wave, and so forth. The term “modulated data signal” means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal.
[0116] The terms “machine-readable medium, ” “computer-readable medium, ” and “device-readable medium” mean the same thing and may be used interchangeably in this disclosure. The terms are defined to include both machine- storage media and transmission media. Thus, the terms include both storage devices / media and carrier waves / modulated data signals. For instance, an embodiment described herein can be implemented using a non-transitory medium (e.g., a non-transitory computer-readable medium) .
[0117] Throughout this specification, plural instances may implement resources, components, operations, or structures described as a single instance. Although individual operations of one or more methods are illustrated and described as separate operations, one or more of the individual operations may be performed concurrently, and nothing requires that the operations be performed in the order illustrated. Structures and functionality presented as separate components in example configurations may be implemented as a combined structure or component. Similarly, structures and functionality presented as a single component may be implemented as separate components.
[0118] As used herein, the term “or” may be construed in either an inclusive or exclusive sense. The terms “a” or “an” should be read as meaning “at least one, ” “one or more, ” or the like. The presence of broadening words and phrases such as “one or more, ” “at least, ” “but not limited to, ” or other like phrases in some instances shall not be read to mean that the narrower case is intended or required in instances where such broadening phrases may be absent. Additionally, boundaries between various resources, operations, components, engines, and data stores are somewhat arbitrary, and particular operations are illustrated in a context of specific illustrative configurations. Other allocations of functionality are envisioned and may fall within a scope of various embodiments of the present disclosure. The specification and drawings are, accordingly, to be regarded in an illustrative rather than a restrictive sense.
[0119] It will be understood that changes and modifications may be made to the disclosed embodiments without departing from the scope of the present disclosure. These and other changes or modifications are intended to be included within the scope of the present disclosure.
Claims
1.A system comprising:one or more hardware processors; andat least one machine-storage medium for storing instructions that, when executed by the one or more hardware processors, cause the one or more hardware processors to perform operations comprising:identifying a behavioral feature associated with a first user;generating, using a first machine learning (ML) model, an embedding vector of the first user based on the behavioral feature;identifying, using a cosine similarity metric, a second user based on the embedding vector of the first user;generating one or more augmented features representing the first user based on one or more features of the second user; andtraining a second ML model based on the one or more augmented features.2.The system of claim 1, wherein the operations comprise:performing, using the trained second ML model, a fraud risk prediction of the first user.3.The system of claim 1, wherein the operations comprise:identifying, using the cosine similarity metric, a plurality of users based on the embedding vector of the first user, the plurality of users including the second user;generating the one or more augmented features that represent the first user based on a plurality of features of the plurality of users; andtraining the second ML model based on the one or more augmented features representing the first user.4.The system of claim 3, wherein the plurality of users comprises a plurality of existing users with historical data, and wherein the first user corresponds to a new user without historical data.5.The system of claim 3, wherein the operations comprise:determining a suspension count based on historical data of the plurality of users, the suspension count representing a number of users from the plurality of users with a suspended account;calculating a suspension rate based on the suspension count; andidentifying a plurality of model scores based on the historical data of the plurality of users, each model score representing a fraud risk determination of a user from the plurality of users.6.The system of claim 5, wherein the operations comprise:assigning one or more of the suspension count, the suspension rate, and the plurality of model scores as the one or more augmented features to the first user; andtraining the second ML model based on one or more of the suspension count, the suspension rate, and the plurality of model scores.7.The system of claim 1, wherein the behavioral feature comprises a page view sequence feature, the page view sequence feature corresponding to a plurality of webpages, each webpage being associated with a dwell time; and wherein the dwell time represents an amount of time taken by the first user to view a webpage before moving on to another webpage.8.The system of claim 1, wherein the operations comprise:retrieving a plurality of embedding vectors of a plurality of users, the plurality of users including the second user, each user in the plurality of users being associated with historical data; andidentifying, using the cosine similarity metric, the second user based on the embedding vector of the first user and the plurality of embedding vectors of the plurality of users.9.The system of claim 1, wherein the second ML model comprises a tree-based ML model, and wherein the tree-based ML model corresponds to an application of eXtreme Gradient Boosting (XGBoost) algorithm.10.The system of claim 1, wherein the first ML model comprises a deep learning ML model.11.A method comprising:identifying, by at least one hardware processor, a behavioral feature associated with a first user;generating, using a first machine learning (ML) model, an embedding vector of the first user based on the behavioral feature;identifying, using a cosine similarity metric, a second user based on the embedding vector of the first user;generating one or more augmented features representing the first user based on one or more features of the second user; andtraining a second ML model based on the one or more augmented features.12.The method of claim 11, comprising:performing, using the trained second ML model, a fraud risk prediction of the first user.13.The method of claim 11, comprising:identifying, using the cosine similarity metric, a plurality of users based on the embedding vector of the first user, the plurality of users including the second user;generating the one or more augmented features that represent the first user based on a plurality of features of the plurality of users; andtraining the second ML model based on the one or more augmented features representing the first user.14.The method of claim 13, wherein the plurality of users comprises a plurality of existing users with historical data, and wherein the first user corresponds to a new user without historical data.15.The method of claim 13, comprising:determining a suspension count based on historical data of the plurality of users, the suspension count representing a number of users from the plurality of users with a suspended account;calculating a suspension rate based on the suspension count; andidentifying a plurality of model scores based on the historical data of the plurality of users, each model score representing a fraud risk determination of a user from the plurality of users.16.The method of claim 15, comprising:assigning one or more of the suspension count, the suspension rate, and the plurality of model scores as the one or more augmented features to the first user; andtraining the second ML model based on one or more of the suspension count, the suspension rate, and the plurality of model scores.17.The method of claim 11, wherein the behavioral feature comprises a page view sequence feature, the page view sequence feature corresponding to a plurality of webpages, each webpage being associated with a dwell time; and wherein the dwell time represents an amount of time taken by the first user to view a webpage before moving on to another webpage.18.The method of claim 11, comprising:retrieving a plurality of embedding vectors of a plurality of users, the plurality of users including the second user, each user in the plurality of users being associated with historical data; andidentifying, using the cosine similarity metric, the second user based on the embedding vector of the first user and the plurality of embedding vectors of the plurality of users.19.The method of claim 11, wherein the second ML model comprises a tree-based ML model, wherein the tree-based ML model corresponds to an application of eXtreme Gradient Boosting (XGBoost) algorithm and wherein the first ML model comprises a deep learning ML model.20.A machine-storage medium for storing instructions that, when executed by one or more hardware processors, cause the one or more hardware processors to perform operations comprising:identifying a behavioral feature associated with a first user;generating, using a first machine learning (ML) model, an embedding vector of the first user based on the behavioral feature;identifying, using a cosine similarity metric, a second user based on the embedding vector of the first user;generating one or more augmented features representing the first user based on one or more features of the second user; andtraining a second ML model based on the one or more augmented features.
Citation Information
Patent Citations
Method and device for identifying fraudulent user, server and storage medium
CN111125658A
Method and device for generating model and method and device for generating information
CN113780607A
Personalized model training method, information display method and equipment
CN114491342A
Behavior fingerprint data enhanced identity authentication method and system
CN115828209A
Method and apparatus for training model based on relation network, and method and apparatus for determining representation based on relation network
US20240185069A1