Method and system for adaptive fraud detection using ai self-learning engine
Patent Information
- Application Number
- IN202441079562
- Authority / Receiving Office
- IN · IN
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-10-19
- Publication Date
- 2026-08-12
- Estimated Expiration
- 2044-10-19
AI Technical Summary
Traditional fraud detection systems in online transactions face challenges such as reliance on manual feature engineering, static models, imbalanced datasets, inability to integrate external data, inadequate performance metrics, and inefficiencies in real-time detection, leading to suboptimal performance and increased financial losses.
The use of an AI Self-Learning Engine and large language models (LLMs) to analyze historical transaction data, generate additional features, and iteratively refine machine learning models based on performance metrics, enabling real-time processing and adaptation to evolving fraud patterns by incorporating external data sources like credit scores.
This approach enhances fraud detection accuracy, adaptability, and reduces manual effort, enabling timely alerts and minimizing financial losses by continuously optimizing model performance and feature relevance.
Abstract
Description
FIELD OF THE INVENTION
[0001] The present invention relates to fraud detection techniques in online transac-tions. More particularly, it relates to a system and method for adaptive fraud detection using AI Self-Learning Engine to analyze transaction data, generate additional fea-tures, and iteratively optimize model performance to improve fraud detection accura-cy.BACKGROUND OF THE INVENTION
[0002] In recent years, financial institutions have faced an increasing challenge in detecting and preventing fraudulent transactions. The growth of online banking, digi-tal payments, and other financial technologies has significantly expanded the scope for fraudulent activities, making it difficult for traditional fraud detection systems to keep pace. Detecting fraud in such a dynamic and rapidly evolving landscape requires highly accurate and adaptable solutions, yet many of the currently available methods fall short of this requirement.
[0003] One of the primary issues with traditional fraud detection systems is their reliance on manually engineered features. This approach is time-consuming and la-bor-intensive, requiring domain experts to manually extract patterns from the data. Furthermore, manual feature engineering often fails to capture the full complexity of transaction data, particularly when dealing with large-scale datasets that contain di-verse patterns and behaviors. As a result, these systems may overlook subtle indica-tors of fraud, leading to false negatives.
[0004] Another significant problem in the field is that many fraud detection models are static in nature. They rely on a predefined set of features and rules that may not be sufficient to detect new or evolving fraud patterns. Fraudsters constantly adapt their tactics, exploiting new vulnerabilities in financial systems. As a result, static models become less effective over time, as they struggle to generalize to new types of fraudu-lent behavior that were not accounted for during their initial training.
[0005] Additionally, traditional models often suffer from imbalanced datasets, where the number of fraudulent transactions is significantly lower than the number of legitimate transactions. This imbalance can skew the model's performance, leading to a higher rate of false positives or false negatives. Fraud detection models must be able to effectively handle such imbalances to avoid flagging legitimate transactions as fraudulent, which can cause unnecessary friction for customers and harm the institu-tion's reputation.
[0006] Moreover, many fraud detection systems lack the ability to integrate exter-nal data sources, such as credit scores or other financial indicators, which could pro-vide valuable context for identifying suspicious activities. By failing to leverage ex-ternal data, these systems miss opportunities to improve detection accuracy and en-hance their ability to assess the risk associated with specific transactions.
[0007] In addition to data-related challenges, performance metrics used in fraud detection can often be inadequate. While accuracy is a common metric, it does not necessarily reflect the model's effectiveness in identifying fraud. Precision, recall, and other more targeted metrics like ROC-AUC are crucial for assessing the true per-formance of a fraud detection system. However, many systems do not optimize for these metrics, leading to suboptimal performance in real-world scenarios.
[0008] Furthermore, real-time detection is a critical requirement in modern fraud prevention systems, yet many existing models are unable to process transactions quickly enough to provide timely alerts. Delays in detecting fraudulent activities can result in significant financial losses, as institutions may fail to act before further fraudulent transactions occur.
[0009] Lastly, the cost and complexity of continuously updating fraud detection systems pose a significant burden on financial institutions. As new fraud patterns emerge, systems must be retrained or updated, often requiring substantial time and computational resources. This need for constant manual intervention and model ad-justment leads to inefficiencies and slows down the overall fraud detection process.SUMMARY OF THE INVENTION
[0010] To address the foregoing problems, in whole or in part, and / or other problems that may have been observed by persons skilled in the art, the present disclosure provides system and methods as described by way of example as set forth below.
[0011] The principal object of the present invention is to improve the accuracy of detecting fraudulent transactions by utilizing AI Self-Learning Engine, machine learning models and large language models (LLMs) to analyze historical transaction data and generate additional features.
[0012] Another object of the invention is to automate the extraction and generation of relevant features from transaction data using LLMs, reducing the reliance on manual feature engineering and improving the system's adaptability to changing fraud patterns.
[0013] Another object of the invention is to develop a fraud detection system that can adapt to new and evolving fraudulent behaviors through iterative model retraining and refinement based on performance metrics.
[0014] Another object of the invention is to enhance detection capabilities by incorporating external data, such as credit scores and transaction types, to provide additional context and improve risk assessment.
[0015] Another object of the invention is to design the system to enable real-time processing of transactions to provide timely alerts for potential fraudulent activities, minimizing financial losses and enhancing customer experience.
[0016] In view of the foregoing, the present invention provides a method for adaptive fraud, comprising the steps of receiving a dataset that includes historical transaction data, training a machine learning model on this dataset to detect fraudulent transactions, and analyzing feature importance using a large language model (LLM) to identify key features that influence fraud detection. The method further includes generating additional features from the transaction data based on the identified patterns using the LLM, augmenting the dataset with these newly generatedfeatures, and retraining the machine learning model using the augmented dataset.Subsequently, the performance of the retrained machine learning model is evaluated based on predefined fraud detection metrics. Finally, the method involves iteratively refining the machine learning model by adding or removing features based on changes in the performance metrics.
[0017] In another aspect of the present invention, the machine learning model is a LightGBM model.
[0018] In another aspect of the present invention, the pre-processing of the dataset includes removing missing data and normalizing numerical values.
[0019] In another aspect of the present invention, the predefined fraud detection metrics include accuracy, precision, recall, and ROC-AUC score.
[0020] In another aspect of the present invention, the additional features generated by the large language model are based on textual data related to transaction descrip-tions.
[0021] In another aspect of the present invention, the dataset is augmented with ex-ternal data, the external data comprising credit scores and transaction types.
[0022] In another aspect of the present invention, the invention discloses a system for adaptive fraud detection in a financial system is disclosed, comprising a data stor-age unit configured to store a dataset that includes historical transaction data, and a machine learning model configured to be trained on this dataset to detect fraudulent transactions. The system includes a large language model (LLM) module configured to analyze feature importance and identify key features that influence fraud detection. Additionally, a feature generation module is provided, configured to generate addi-tional features from the transaction data based on the identified patterns using the LLM. A data augmentation module is included, configured to augment the dataset with the newly generated features. The system further comprises a model retraining unit configured to retrain the machine learning model using the augmented dataset, and a performance evaluation module configured to evaluate the performance of the retrained machine learning model based on predefined fraud detection metrics. Fur-ther, the system discloses an iterative refinement module is included, configured to iteratively refine the machine learning model by adding or removing features based on changes in the performance metrics.
[0023] Additional features of the invention will be or will become apparent to one with skill in the art upon examination of the following figures and detailed descrip-tion. It is intended that all such additional features and advantages be included within this description, be within the scope of the invention, and be protected by the accom-panying claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Having thus described the subject matter of the present invention in general terms,reference will now be made to the accompanying drawings, which are not necessarily drawn toscale, and wherein:
[0025] Figure 1 illustrates a framework for fraud detection in financial transactions, in accordance with an embodiment of the present invention;
[0026] Figure 2 illustrates a flowchart of the system for adaptive fraud detection in a financial system, in accordance with an embodiment of the present invention
[0027] Skilled artisans will appreciate that elements in the drawings are illustrated for simplicity and may not have necessarily been drawn to scale. Furthermore, in terms of the construction of the device, one or more components of the device may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the embodiments of the present invention so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein.DETAILED DESCRIPTION OF THE INVENTION
[0028] The subject matter of the present invention now will be described more fully hereinafter with reference to the accompanying drawings, in which some, but not all embodiments of the subject matter of the present invention are shown. Like numbers refer to like elements throughout. The subject matter of the present invention may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will satisfy applicable legal requirements. Indeed, many modifications and other embodiments of the subject matter of the present invention set forth herein will come to mind to one skilled in the art to which the subject matter of the present invention pertains having the benefit of the teachings presented in the foregoing descriptions and the associated drawings. All illustrations of the drawings are for the purpose of describing selected versions of the present invention and are not intended to limit the scope of the present invention. Therefore, it is to be understood that the subject matter of the present invention is not to be limited to the specific embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the appended claims.
[0029] As a preliminary matter, it will readily be understood by one having ordinary skill in the relevant art that the present disclosure has broad utility and application. As should be understood, any embodiment may incorporate only one or a plurality of the above-disclosed aspects of the disclosure and may further incorporate only one or a plurality of the above-disclosed features. Furthermore, any embodiment discussed and identified as being "preferred" is considered to be part of a best mode contemplated for carrying out the embodiments of the present disclosure. Other embodiments also may be discussed for additional illustrative purposes in providing a full and enabling disclosure. Moreover, many embodiments, such as adaptations, variations, modifications, and equivalent arrangements, will be implicitly disclosed by the embodiments described herein and fall within the scope of the present disclosure.
[0030] Accordingly, while embodiments are described herein in detail in relation to one or more embodiments, it is to be understood that this disclosure is illustrative and example of the present disclosure and are made merely for the purposes of providing a full and enabling disclosure. The detailed disclosure herein of one or more embodiments is not intended, nor is to be construed, to limit the scope of patent protection afforded in any claim of a patent issuing here from, which scope is to be defined by the claims and the equivalents thereof. It is not intended that the scope of patent protection be defined by reading into any claim a limitation found herein that does not explicitly appear in the claim itself.
[0031] Thus, for example, any sequence(s) and / or temporal order of steps of various processes or methods that are described herein are illustrative and not restrictive. Accordingly, it should be understood that, although steps of various processes or methods may be shown and described as being in a sequence or temporal order, the steps of any such processes or methods are not limited to being carried out in any particular sequence or order, absent an indication otherwise. Indeed, the steps in such processes or methods generally may be carried out in various different sequences and orders while still falling within the scope of the present invention. Accordingly, it is intended that the scope of patent protection is to be defined by the issued claim(s) rather than the description set forth herein.
[0032] Additionally, it is important to note that each term used herein refers to that which an ordinary artisan would understand such term to mean based on the contextual use of such term herein. To the extent that the meaning of a term used herein - as understood by the ordinary artisan based on the contextual use of such term - differs in any way from any particular dictionary definition of such term, it is intended that the meaning of the term as understood by the ordinary artisan should prevail.
[0033] Furthermore, it is important to note that, as used herein, "a" and "an" each generally denotes "at least one", but does not exclude a plurality unless the contextual use dictates otherwise. When used herein to join a list of items, "or" denotes "at leastone of the items", but does not exclude a plurality of items of the list. Finally, whenused herein to join a list of items, "and" denotes "all of the items of the list".
[0034] The present invention discloses a system and method for adaptive fraud detection in online transactions, addressing the growing challenges associated with identifying fraudulent activities in an increasingly digital landscape. With the rise of online banking, digital payments, and complex financial transactions, traditional fraud detection methods have become inadequate. They often rely on static models and manual feature engineering, which struggle to keep pace with the rapidly evolving tactics employed by fraudsters. This invention aims to provide a more dynamic and effective solution by integrating machine learning and large language models (LLMs) to enhance fraud detection capabilities.
[0035] The invention discloses a method that involves receiving a dataset comprising historical transaction data. This dataset serves as the foundation for training a machine learning model, which is designed to detect fraudulent transactions. By utilizing advanced algorithms, the system can analyze vast amounts of transaction data, allowing it to identify patterns and anomalies indicative of fraud. This training phase is crucial, as it enables the model to learn from past transactions and refine its ability to distinguish between legitimate and fraudulent activities.
[0036] The invention uses LLMs to analyze feature importance within the dataset. These models are capable of extracting complex relationships and identifying key features that influence fraud detection. This analysis allows the system to gain insights into which factors are most relevant in predicting fraudulent behavior. By understanding these key features, the system can generate additional, contextually relevant features from the transaction data, further enhancing its detection capabilities.
[0037] The invention further discloses augmenting the dataset with newly generated features, which enrich the existing data and provide the machine learning model with more informative inputs. This augmentation is complemented by an iterative refinement process, where the machine learning model is retrained using theenhanced dataset. Performance evaluation is a critical aspect of this process, as thesystem assesses the model's accuracy using predefined fraud detection metrics such as precision, recall, and ROC-AUC score. This continuous evaluation ensures that the model adapts to changing fraud patterns and improves over time.
[0038] Furthermore, the disclosed system is designed to be proactive in its approach to fraud detection. By providing an alert mechanism that triggers notifications when shifts in transaction patterns are detected, the invention enables financial institutions to respond promptly to potential threats. This capability is vital in minimizing losses and maintaining trust with customers. The integration of external data sources, such as credit scores and transaction types, further enhances the system's ability to assess risk and improve detection accuracy.
[0039] In accordance with an embodiment of the present invention, Figure 1 illustrates a framework for fraud detection in financial transactions. The flowchart shown in this figure shows an approach that combines machine learning techniques with large language models (LLMs) to automatically extract and generate features necessary for effective task detection. The method aims to enhance the accuracy and generalizability of fraud detection models, addressing the limitations of traditional systems that often rely on static, manually engineered features. By using advanced algorithms and data-driven insights, the framework is designed to adapt dynamically to the evolving landscape of financial fraud.
[0040] The process starts with data preparation (A), where raw transaction data is preprocessed and cleaned to ensure its suitability for machine learning algorithms. This crucial step involves removing any inconsistencies, handling missing values, and normalizing the data to facilitate better learning outcomes. Once the data is prepared, it is divided into three distinct sets: training, validation, and testing. This partitioning is essential for training the machine learning model effectively while allowing for unbiased evaluation of its performance later in the process.
[0041] Following data preparation, the framework moves to the training of the LightGBM model (B). The model is trained on the training dataset, during which itlearns to identify patterns indicative of fraudulent transactions. Hyperparameters ofthe model are optimized using the validation set to ensure the best performance. After training, the model generates initial task detection results (C) by applying its learned parameters to the testing dataset. The performance of the model is evaluated using various metrics, including accuracy, precision, and recall, providing a baseline measure of effectiveness in detecting fraud.
[0042] Once the initial performance metrics are established, the framework proceeds to calculate model metrics and feature importance (D) to identify which features are most influential in detecting fraud. This information serves as a foundation for further improvements to the model. The LLM then analyzes feature importance (E) to uncover additional insights about the data. By identifying the key features that contribute to effective fraud detection, the system can guide future feature engineering efforts, ensuring that the model focuses on the most impactful variables.
[0043] The system then utilizes the LLM to generate new features (F) from the existing data. This step is crucial because the LLM can extract complex patterns and relationships that traditional feature engineering methods may overlook. The newly generated features are subsequently incorporated into the existing dataset (G), effectively augmenting it with richer, more informative data. This expanded dataset enhances the training process and provides the LightGBM model with more relevant information for detecting fraudulent transactions. After expanding the dataset, the LightGBM model is retrained (H) using the newly enriched data. This retraining allows the model to leverage the additional features, improving its detection capabilities. The framework then generates new task detection results (I) by applying the retrained model to the testing data and assesses its performance against the original model. New metrics are computed (J) to evaluate whether there has been a significant improvement compared to the baseline established earlier. A critical step in this iterative process is the evaluation of metrics improvement (K). If the retrained model's performance has improved, the LLM-generated new features are added to thefeature library (L). Conversely, if no significant improvement is observed, the newfeatures are removed (M) to streamline the model.
[0044] Finally, the framework emphasizes the iterative nature of the process (N), where the analysis of feature importance, feature generation, dataset expansion, model retraining, and metrics evaluation are continuously repeated. This iterative approach aims to identify the optimal set of features that achieve the highest possible task detection metrics. By continuously refining the model based on empirical data and evolving patterns, the framework holds the potential to significantly improve the metrics and generalizability of fraud detection models. This comprehensive methodology not only enhances detection accuracy but also allows financial institutions to respond proactively to fraudulent activities, thereby safeguarding their operations and maintaining customer trust.
[0045] In accordance with an embodiment of the present invention, Figure 2 illustrates a flowchart of the system for adaptive fraud detection in a online transactions (200). Data Storage Unit (202) is responsible for storing the dataset comprising historical transaction data. This component acts as the repository of all transaction records, both legitimate and fraudulent, that will be used for training the machine learning model. The stored dataset provides the basis for detecting patterns of fraudulent behavior in financial transactions. Next, the Machine Learning Model (204) is trained on the dataset stored in the data storage unit. The model is designed to detect fraudulent transactions by learning from historical transaction data. It uses algorithms to distinguish between normal and fraudulent activities based on transaction patterns, making it the core engine for identifying suspicious behaviors in financial operations.
[0046] The Large Language Model (LLM) Module 206 analyzes the trained machine learning model to identify the feature importance, highlighting the key features that influence fraud detection. The LLM enables the system to automatically identify critical aspects of the data, such as patterns and relationships that are mostindicative of fraudulent behavior. By analyzing the model's output, the LLMprovides insights into the features that significantly impact fraud detection.
[0047] Following this, the Feature Generation Module 208 utilizes the LLM to generate additional features from the transaction data. These additional features are derived based on the patterns identified by the LLM. The generated features, which may not have been obvious through traditional methods, provide richer data inputs for the system, further improving the model's ability to detect fraudulent activities.
[0048] Once the new features are generated, the Data Augmentation Module 210 takes over. This component augments the existing dataset with the newly generated features, thereby enriching the original dataset with more informative data. The enriched dataset is crucial for improving the accuracy and performance of the machine learning model by providing it with additional, more relevant data points.
[0049] Subsequently, the Model Retraining Unit 212 retrains the machine learning model using the augmented dataset. By incorporating the newly generated features into the training process, the model is continuously updated to reflect the latest patterns in fraudulent behavior. This ensures that the system remains up to date and effective in detecting evolving fraud tactics.
[0050] After retraining the model, the Performance Evaluation Module 214 evaluates the performance of the retrained machine learning model based on predefined fraud detection metrics such as accuracy, precision, recall, and ROC-AUC. This step determines whether the retrained model demonstrates improvements in fraud detection compared to the previous version. The evaluation process provides a quantitative assessment of the model's effectiveness.
[0051] Lastly, the Iterative Refinement Module 216 is responsible for iteratively refining the machine learning model. Based on changes in performance metrics, this component adds or removes features from the dataset, ensuring that the model continues to improve over time. This iterative refinement ensures the continuous optimization of the system for maximum fraud detection accuracy and adaptability.
[0052] The claimed invention offers several significant advantages that enhance the effectiveness of fraud detection in financial systems. One of the primary benefits is the reduction of time and effort required for feature extraction. By automating this process using machine learning algorithms and large language models (LLMs), the invention eliminates the labor-intensive manual feature engineering typically associated with traditional fraud detection systems. This not only saves valuable time but also reduces the reliance on human expertise, making the system more user-friendly and efficient. Additionally, the automation allows for the identification of more relevant and informative features, which leads to improved performance in detecting fraudulent transactions. The ability to learn from vast datasets enables the system to adapt to evolving fraud tactics more effectively.
[0053] Moreover, the invention addresses critical issues such as low accuracy of task detection and lack of generalizability commonly faced in existing systems. By employing a more effective feature extraction method, the invention significantly enhances the accuracy of fraud detection, ensuring that relevant patterns are captured and utilized. This approach also provides increased adaptability, allowing the system to easily adjust to new tasks and changing data distributions through model retraining. Consequently, the claimed invention proves to be cost-effective, as it reduces the overall expenses associated with feature engineering while simultaneously improving accuracy. Furthermore, the system's increased efficiency in data processing allows financial institutions to save both time and resources, while its scalability ensures that it can handle growing volumes of transactions without compromising performance.
[0054] Further, in addition to the above, the invention provides several beneficial effects:
[0055] Some of the non-limiting advantages of the present invention are:-In-depth Feature Analysis: Analyzing feature importance using the trained LightGBM model helps identify the most influential featuresfor task detection. This insight guides further feature engineeringefforts, ensuring that the most relevant features are retained and enhanced.-Innovative Feature Generation: Utilizing an LLM to generate new features from the data introduces advanced and complex patterns that traditional feature engineering methods might miss. This step leverages the sophisticated understanding of LLMs to enrich the dataset with more informative features.-Streamlined Feature Engineering: The LLM automates the feature generation process, reducing the need for manual feature engineering efforts. This saves time and resources while potentially uncovering valuable features that might have been overlooked.-Dataset Augmentation: Incorporating the newly generated features into the existing dataset augments the data, providing a richer and more comprehensive foundation for model training. This expanded dataset enhances the model's ability to detect tasks accurately.-Iterative Model Improvement: The iterative process of LLM-based feature analysis and evaluation allows for the identification of the most impactful features for task detection. Irrelevant or redundant features can be removed, leading to a more efficient and effective model. Retraining the LightGBM model with the expanded dataset and generating new task detection results ensures continuous improvement. This iterative process allows for regular updates and refinements to the model, maintaining its relevance and accuracy over time. -Feature Library Optimization: The process of adding beneficial LLM-generated features to the feature library and removing non-beneficial ones ensures that the feature set remains optimized for performance.This dynamic adjustment contributes to sustained improvements intask detection accuracy.-Continuous Metrics Evaluation: By consistently calculating and comparing model metrics after each iteration, the process ensures that any improvements or declines in performance are promptly identified. This facilitates informed decision-making regarding feature retention or removal.-Continuous Improvement Potential: The iterative nature of the method allows for continuous improvement of the task detection model. New features and insights can be readily incorporated, leading to ongoing enhancements in performance.-Scalable and Adaptive: The described approach is scalable and adaptable, allowing for continuous iterations of feature generation, dataset expansion, model retraining, and evaluation. This flexibility ensures that the model can evolve with changing data patterns and detection requirements.
[0056] Terms and phrases used in this document, and variations thereof, unless oth-erwise expressly stated, should be construed as open-ended as opposed to limiting. As examples of the foregoing: the term "including" should be read as mean "including, without limitation" or the like; the term "example" is used to provide exemplary in-stances of the item in discussion, not an exhaustive or limiting list thereof; and adjec-tives such as "conventional," "traditional," "standard," "known" and terms of similar meaning should not be construed as limiting the item described to a given time period or to an item available as of a given time, but instead should be read to encompass conventional, traditional, normal, or standard technologies that may be available or known now or at any time in the future. Likewise, a group of items linked with the conjunction "and" should not be read as requiring that each and every one of thoseitems be present in the grouping, but rather should be read as "and / or" unless express-ly stated otherwise. Similarly, a group of items linked with the conjunction "or" should not be read as requiring mutual exclusivity among that group, but rather should also be read as "and / or" unless expressly stated otherwise. Furthermore, alt-hough item, elements or components of the disclosure may be described or claimed in the singular, the plural is contemplated to be within the scope thereof unless limita-tion to the singular is explicitly stated. The presence of broadening words and phrases such as "one or more," "at least," "but not limited to" or other like phrases in some instances shall not be read to mean that the narrower case is intended or required in instances where such broadening phrases may be absent.
[0057] For the purposes of this specification and appended claims, unless otherwise indicated, all numbers expressing amounts, sizes, dimensions, proportions, shapes, formulations, parameters, percentages, quantities, characteristics, and other numerical values used in the specification and claims, are to be understood as being modified in all instances by the term "about" even though the term "about" may not expressly appear with the value, amount, or range. Accordingly, unless indicated to the contrary, the numerical parameters set forth in the following specification and attached claims are not and need not be exact, but may be approximate and / or larger or smaller as desired, reflecting tolerances, conversion factors, rounding off, measurement error and the like, and other factors known to those of skill in the art depending on the desired properties sought to be obtained by the subject matter of the present invention. For example, the term "about," when referring to a value can be meant to encompass variations of, in some embodiments ± 100%, in some embodiments ± 50%, in some embodiments ± 20%, in some embodiments ± 10%, in some embodiments ± 5%, in some embodiments ± 1%, in some embodiments ± 0.5%, and in some embodiments ± 0.1% from the specified amount, as such variations are appropriate to perform the disclosed methods or employ the disclosed compositions.
[0058] Further, the term "about" when used in connection with one or more numbers or numerical ranges, should be understood to refer to all such numbers,including all numbers in a range and modifies that range by extending the boundariesabove and below the numerical values set forth. The recitation of numerical ranges by endpoints includes all numbers, e.g., whole integers, including fractions thereof, subsumed within that range (for example, the recitation of 1 to 5 includes 1, 2, 3, 4, and 5, as well as fractions thereof, e.g., 1.5, 2.25, 3.75, 4.1, and the like) and any range within that range.
[0059] All publications, patent applications, patents, and other references mentioned in the specification are indicative of the level of those skilled in the art to which the presently disclosed subject matter pertains. All publications, patent applications, patents, and other references are herein incorporated by reference to the same extent as if each individual publication, patent application, patent, and other reference was specifically and individually indicated to be incorporated by reference. It will be understood that, although a number of patent applications, patents, and other references are referred to herein, such reference does not constitute an admission that any of these documents forms part of the common general knowledge in the art. Although the foregoing subject matter has been described in some detail by way of illustration and example for purposes of clarity of understanding, it will be understood by those skilled in the art that certain changes and modifications can be practiced within the scope of the appended claims.
Claims
1. A method for adaptive fraud detection in online transactions (200) using AI self-learning engine, the method comprising: receiving a dataset comprising historical transaction data; training a machine learning model on the dataset to detect fraudulent transactions; analyzing, using a large language model (LLM), feature importance to identify key features that influence fraud detection; generating, using the LLM, additional features from the transaction data based on the identified patterns; augmenting the dataset with the newly generated features; retraining the machine learning model using the augmented dataset; evaluating the performance of the retrained machine learning model based on predefined fraud detection metrics; and iteratively refining the machine learning model by adding or removing features based on changes in the performance metrics.
2. The method as claimed in claim 1, wherein the machine learning model is a LightGBM model.
3. The method as claimed in claim 1, wherein the pre-processing of the dataset includes removing missing data and normalizing numerical values.
4. The method as claimed in claim 1, wherein the predefined fraud detection metrics include accuracy, precision, recall, and ROC-AUC score.
5. The method as claimed in claim 1, wherein the additional features generated by the large language model are based on textual data related to transaction descrip-tions.
6. The method as claimed in claim 1, wherein the dataset is augmented with ex-ternal data, the external data comprising credit scores and transaction types.
7. The method as claimed in claim 1, further comprising applying dimensionali-ty reduction to the dataset before training the machine learning model.
8. The method as claimed in claim 1, further comprising providing an alert sys-tem that triggers an alert when shifts in transaction patterns are detected.
9. The method as claimed in claim 1, wherein the additional features generated by the large language model include temporal patterns in customer behaviour.
10. The method as claimed in claim 1, wherein the iterative refinement process stops when the performance improvement of the machine learning model falls below a predefined threshold.
11. A system for adaptive fraud detection in online transactions (200), using AI self-learning engine, the system comprising: a data storage unit (202) configured to store a dataset comprising historical transaction data; a machine learning model (204) configured to be trained on the dataset to detect fraudulent transactions; a large language model (LLM) module (206) configured to analyze feature importance to identify key features that influence fraud detection; a feature generation module (208) configured to generate additional features from the transaction data based on the identified patterns using the LLM; a data augmentation module (210) configured to augment the dataset with the newly generated features; a model retraining unit (212) configured to retrain the machine learning model using the augmented dataset; a performance evaluation module (214) configured to evaluate the performance of the retrained machine learning model based on predefined fraud detection metrics; and an iterative refinement module (216) configured to iteratively refine the machine learning model by adding or removing features based on changes in the performance metrics.