Method and device for event processing in an Anti-fraud system
The method enhances the accuracy and speed of event processing in anti-fraud systems by enabling real-time data enrichment and analysis, addressing the limitations of existing systems in detecting fraudulent transactions.
Patent Information
- Application Number
- PCT/RU2023/000373
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-21
- Filing Date
- 2023-11-29
- Publication Date
- 2025-05-30
AI Technical Summary
Existing anti-fraud systems face challenges in achieving sufficient accuracy and speed in processing events due to the inability to update parameters for data enrichment in real-time.
A computer-implemented method for processing events in an anti-fraud system, which involves loading a configuration representing parameters for data enrichment, parsing and formatting incoming events, enriching data in real-time, analyzing and processing the enriched data, and making decisions on fraudulent transactions based on the processed results.
The solution enhances the accuracy and speed of event processing in anti-fraud systems by enabling real-time data enrichment and analysis, thereby improving the detection of fraudulent transactions.
Smart Images

Figure RU2023000373_30052025_PF_FP_ABST
Abstract
Description
METHOD AND DEVICE FOR PROCESSING EVENTS IN AN ANTI-FRAUD SYSTEM AREA OF TECHNOLOGY
[0001] The declared solution relates to the field of information security, in particular, to anti-fraud systems. LEVEL OF TECHNOLOGY
[0002] Financial institutions such as banks, fintech companies, payment services, etc. provide a wide range of services to their customers. Customers using financial institution services may encounter fraudulent transactions by third parties regarding their accounts, cards, credit limits, and other financial instruments. To prevent fraud, financial institutions use anti-fraud systems.
[0003] For the effective operation of anti-fraud systems, it is not enough to have data for analysis from current operations; it is also necessary to obtain data from external sources, such as: • historical data (internal information of the financial institution), • reference data, • data from companies providing analytical information (data on clients / companies from credit bureaus, etc.), • data from government institutions and agencies associated with them (FinCERT of the Bank of Russia).
[0004] It is worth emphasizing that financial institutions work with a large volume of data, i.e. they use a highly loaded system with a large number of events. Also, within the framework of the laws of the Russian Federation [1], the requirements of payment networks [2], [3], payment providers, payment services, the regulations of work [4] are strictly defined, including the response time to an event. Thus, a financial institution imposes strict requirements on the load and time of event processing on the side of anti-fraud systems.
[0005] One of the important tasks associated with analysis in the anti-fraud system is enrichment with additional data to improve the quality of analysis from other sources.
[0006] The state of the art includes a system for combating fraudulent transactions from JSC Cross Technologies and MKB [5], designed to improve fraud detection efficiency and reduction of false positives, as well as a fraud prevention system, which is part of the SmartVista product family from BPS [6].
[0007] In addition, patent application US 2015 / 0026027 Al (GUARDIAN ANALYTICS, INC., 01 / 22 / 2015) discloses a fraud prevention system and method for use in preventing account fraud and identity theft that supports an end-to-end online risk management process using behavior-based modeling and advanced analytics.
[0008] Patent application CN 112862505 A- (INDUSTRIAL AND COMMERCIAL BANK OF CHINA CO., LTD., May 28, 2021) discloses a method and device for exchanging anti-fraud information based on blockchain.
[0009] Patent application CN 109242107 A- (BEIJING XINDUN TIMES TECHNOLOGY CO., LTD., 18.01.2019) discloses a system and method for countering fraud based on migration training in e-banking.
[0010] Patent application CN 113837885A- (SHANGHAI CINTEL INTELLIGENCE TELECOM SYSTEM CO., LTD. et al., 12 / 24 / 2021) discloses a financial anti-fraud service system that provides big data processing and analysis to combat financial fraud. [UN] The disadvantage of these solutions is the insufficient accuracy and speed of event processing in the anti-fraud system due to the lack of the ability to update parameters for data enrichment in real time. ESSENCE OF THE INVENTION
[0012] The proposed solution allows us to solve a technical problem in terms of increasing the accuracy and speed of detecting a fraudulent transaction by increasing the accuracy and speed of processing events in the anti-fraud system.
[0013] The technical result is an increase in the accuracy and speed of event processing in the anti-fraud system.
[0014] The claimed technical result is achieved by executing a computer-implemented method for processing events in an anti-fraud system, executed by at least one processor, and containing the steps of: a) loading at least one configuration representing a set of parameters for enriching data, b) receiving at least one incoming event, c) parsing at least one received event and converting it into a given format, d) enriching data taking into account the configuration and adding the enriched data to the event formatted in step c), wherein the configuration is updated in real time, e) analyzing and processing data from the event and / or data marts required for the model block, and adding the analyzed and processed data to the event, f) sending at least one event for verification to the model block, which returns the final value of the processing result, wherein the final value is added to the event, g) generating a response message based on the data stored in the event, h) making a decision to recognize the operation as fraudulent based on the generated response message.
[0015] In one particular implementation example, the incoming event is a message containing data in the form of a set of key-value pairs.
[0016] In another particular implementation example, data enrichment occurs from memory containing data marts.
[0017] In another particular implementation example, the configuration comprises a unique identifier by means of which data is loaded from data marts, wherein the unique identifier corresponds to at least one data mart, and the value of the unique identifier is obtained from an incoming event.
[0018] In another particular example of implementation, predictors are additionally calculated that increase the predictive power of the model, based on the analyzed and processed data from the data marts, with at least the following acting as predictors: • aggregate variables, • demographic data, • reference information.
[0019] In another particular example of implementation, step f) additionally includes comparing the final value of the processing result with a threshold range of values, whereby if the final value of the processing result corresponds to the threshold range of values, then additional enrichment of the data is performed.
[0020] In another particular example of implementation, additional data enrichment is carried out taking into account the configuration and the parameter correspondence table.
[0021] In another particular implementation example, the response message is additionally stored in the data mart.
[0022] In another particular example of implementation, the method additionally comprises the step of restricting access of a user performing at least one operation recognized as fraudulent.
[0023] The claimed technical result is also achieved by implementing a device for processing events in an anti-fraud system, containing at least one processor, at least one memory associated with the processor and containing machine-readable instructions, which, when executed by at least one processor, ensure the execution of a method for processing events in an anti-fraud system. BRIEF DESCRIPTION OF DRAWINGS
[0024] Fig. 1 shows a block diagram of a computer-implemented method for processing events in an anti-fraud system.
[0025] Fig. 2 shows the general diagram of the computing device. IMPLEMENTATION OF THE INVENTION
[0026] Below, concepts and terms necessary for understanding the present invention will be described.
[0027] Anti-fraud system (also known as anti-fraud system) is a system designed to evaluate financial and non-financial events (card transactions, insurance claims and insurance payments, user actions in the remote banking system, transactions with loyalty points, etc.) for suspiciousness from the point of view of fraud and offering recommendations for their further processing.
[0028] Data enrichment is the process of adding new information to data to increase its value for analysis.
[0029] Event - an action performed by a user (for example, payment, transfer, change of account data, login to online banking).
[0030] A fraudulent transaction is at least one event that is not initiated or confirmed by the user (e.g., carried out with using the user's device or an account belonging to the user, but without the user's knowledge).
[0031] Parsing (from English, parsing) is the automated collection and systematization of information.
[0032] A data warehouse is a data management system that supports business data analysis and analytics for the entire organization. Data warehouses often contain large amounts of data, including historical data. A data warehouse stores structured data for specific purposes [7].
[0033] A data lake is a place where structured and unstructured data is stored, and a method for organizing large volumes of very different data coming from different sources. Data lakes allow data to be brought in in its original format without modification [7], [8].
[0034] A data mart is a simple form of data warehouse that is focused on a specific topic or area of activity, such as sales, finance, or marketing [7].
[0035] The final value of the processing result (also known as the score) is a metric of the performance of the machine learning model.
[0036] A predictor [in machine learning] is a predictive variable or characteristic that is used to forecast or predict an outcome.
[0037] A message broker is a software program that connects applications, systems, and services to help them exchange information with each other by translating messages from one formal messaging protocol to another [9].
[0038] Fig. 1 shows a computer-implemented method (100) for processing events in an anti-fraud system. The method (100) is executed using at least one processor. In a particular example of implementation, the method (100) is executed using a device for processing events in an anti-fraud system, which can be implemented on the basis of a computing device modified in the software and hardware part in such a way as to perform the functions of a device for processing events in an anti-fraud system. A more detailed description of the computing device is disclosed below with reference to Fig. 2.
[0039] In the first step (101), at least one configuration is loaded, representing a set of parameters for data enrichment. The configuration contains a unique identifier used to load data from data marts, where the unique identifier corresponds to at least one data mart, and the value of the unique identifier is obtained from an incoming event. The unique identifier allows for optimization (reduction) of the query time to the data marts.
[0040] The configuration may further comprise, but is not limited to, at least one of: • configuration of information loading, • model filters, • settings of the parameter lookup table used for additional data enrichment.
[0041] By configuring the data mart loading configuration, you can add a new data mart at any time without restarting or stopping the anti-fraud system.
[0042] Example of implementation of configuration of loading information: , "topic "view userid avg sum last l O days ", "ignite, cache "userld avg sum last l O days ", "key": "userid" where the topic parameter specifies the message broker section from which the data is loaded. The ignite, cache parameter specifies the name of the new data mart into which the loading occurs. The key parameter specifies a unique identifier with which the necessary information is obtained from the data mart.
[0043] The example demonstrates obtaining information from a data mart containing data on the average amount of customer spending over the last 10 days (userld avg sum last 10 days'), by a unique identifier userid.
[0044] The data in the message broker comes from the data lake, after going through a normalization process. Using a message broker in an anti-fraud system allows for real-time processing of streaming data with high throughput and low latency.
[0045] The presence of a unique identifier and information loading configuration in the loaded configuration automatically allows the use of information from data marts after loading, thus reducing integration time and significantly reducing the time required for data enrichment.
[0046] Model filters allow you to set up conditions for checking an incoming event. If an incoming event does not match the conditions, the corresponding model will not be activated. If an incoming event matches the conditions, data is passed to the model block for the corresponding model, which may include data from the event and / or enriched data and / or data from data marts. Model filters can be applied either selectively to a specific incoming event or automatically to every incoming event.
[0047] Filter example (Example #1): "esot Jilter" : "inList(f event, subchannel', ['esot', 'UFS. WEBAPF , 'ESA. WEBAPF]) AND isNotEmpty(fevent.user _id) " { avg sumjast l O days, avg sum current day, cf('event. deviceRequest. atm. merchantld)
[0048] The example checks two fields: 1) The event, subchannel type field must be one of the following values: C esot', ' UFS. WEB AP G , 'ESA. WEBAPF), if the operation is performed via: • Internet ('esot') • via channel (' UFS. WEB AP G) • via partner channel ('ESA. WEBAPF) 2) The event. user id field must not be empty (i.e. the event must be associated with a specific identifiable user).
[0049] If the two conditions are met, then the variables avg sumJast lO days and avg sum current day are calculated / received for the incoming event. After calculation, the incoming event is sent for analysis to the model block, and the model name есот ^filter is also added to the event data.
[0050] Model filters (since they are part of the configuration) can be updated (e.g. changed, added, deleted) in real time, which reduces the time required to process events in the anti-fraud system.
[0051] Model filters allow you to optimize the data set required for a model block, i.e. to transfer not all data from the incoming event and / or enriched data and / or calculated / analyzed data, but only the data set that affects the quality of the final value of the processing result.
[0052] Additionally, model filters allow you to send only specific incoming events, which can improve the quality of the final result value. processing and not sending incoming events, the analysis of which does not allow to determine whether the operation is fraudulent or not. The above-mentioned capabilities of the model filters help to increase the accuracy and speed of event processing in the anti-fraud system.
[0053] At step (102), at least one incoming event is received. In a particular example of implementation, the incoming event is received in a specified interaction format, where the specified interaction format of the received incoming event is the JSON (JavaScript Object Notation) format.
[0054] In one alternative embodiment of the claimed solution, the incoming event is a message containing data in the form of a set of key-value pairs.
[0055] Example of received incoming event in JSON format (Example #2): {"event" : { "deviceRequest" : { "atm" : { "merchant Id" : "101000015560”, "MCCGroup": ”1”, "MCC": "5200", "Terminalld": "20156357", "AcquiringllC" : "111295", "terminal" : { "terminalclass": "8" }} , "user id" : "10001" , "channel": "ISSUER" , "operation amount" : 10256700 , "operation geo data" : "35.456;47.987" , "eventTime" : "2023-08-17 14:16:54" , "subchannel" : "ecom"}} , "ext":[ "name " : "Card number ", "value": "12253" "name" : "Geodata of the previous operation", "value": "99.555;100.321" ]} \0
[0056] Each event is separated by a separator, which is a combination of specified characters. In the above example, the separator is a combination of characters - \0. Using a separator allows you to reduce the time it takes to process an event.
[0057] In the known level of technology, a normal byte message over the TCP / IP protocol begins with a header, which is a block of information about the incoming event, which contains the length of the message. At the beginning, it is necessary to read the parameter from the header, then calculate the length, and only then does the message size determination occur.
[0058] In the stated solution, messages are separated by a separator, and therefore no additional transformations are required, which helps speed up event processing.
[0059] Description of the syntax for the first event from example #2:
[0060] At step (103), at least one received event is parsed and converted into a specified format. The specified format may be, but is not limited to, a modified Java HashMap. As part of the stated solution, the Java HashMap format was modified in such a way as to provide direct access to the fields. The original Java HashMap format does not provide such functionality, since the field is accessed through an array.
[0061] An example of the format into which the event is converted after parsing (Example #3): event\deviceRequest\atm\merchant!d= ”101000015560" event\deviceRequest\atm\MCCGroup= "1 " event\deviceRequest\atm\MCC= "5200" event\deviceRequest\atm\TerminalId= "20156357" event\deviceRequest\atm\AcquiringIIC= "111295 " event\deviceRequest\atm \terminal\terminalClass = ”8 " event\user _id= "10001 " event\channel="ISSUER" event\operation_amount =10256700 event\operation_geo data = "35.456; 47.98 " event\eventTime="2023-08-17 14:16:54" event\subchannel="ecom" "Card number" = ”12253" "Previous operation geodata" = "99.555, 100.321"
[0062] The specified format allows access to data directly, which reduces the time for event processing and analysis.
[0063] At step (104), the data is enriched taking into account the configuration and the enriched data is added to the event formatted at the previous step (103).
[0064] The data is enriched by means of a unique identifier stored in the configuration. The unique identifier can correspond to at least one data mart. Using the unique identifier, we obtain data from the corresponding data mart.
[0065] Let's look at an example of enrichment for the above event (see Example #3). Using the unique identifier userid, we access the corresponding data mart userId_avg_sum_last_10_days. The value of the unique identifier userid is 10001, this value is extracted from the formatted incoming event. In the specified mart, we find the row corresponding to the value of the unique identifier userid'. 10001 [variable from datajnart avg last l O days = 2340051 ]
[0066] From the above line, we extract the variable_from_data_mart_avg_last_10_days parameter, which characterizes the average amount of customer spending over the last 10 days, and its value. Then, the extracted parameter and its value (i.e., enriched data) are added to the formatted incoming event.
[0067] An enriched event (i.e. an incoming event transformed into a specified format and enriched with data (the added field is shown in bold) from at least one data mart) is shown below in Example #4: event\deviceRequest\atm\merchantld= "101000015560" event\deviceRequest\atm\MCCGroup= "1 " event\deviceRequest\atm\MCC="5200" event\deviceRequest\atm\TerminalId= "20156357" event\deviceRequest\atm\AcquiringIIC= "111295 " event\deviceRequest\atm\terminal\terminalClass="8" event\user _id= "10001 " event\channel= "ISSUER " event\operation amount =10256700 event operation geo data = "35.456; 47.98 " event\eventTime ="2023-08-17 14:16:54" event\subchannel= "ecom" "Card number" = "12253" "Geodata of previous operation" = "99.555; 100.321" variable from data mart avg last 10 days=2340051
[0068] The process of enriching data from data marts allows to reduce the processing time of an incoming event, since there is no need to transform parameters from the data mart (for example, saving parameters to fields supported by the anti-fraud system, or transforming the data type supported by the anti-fraud system), all received information is used as is, i.e. in its original form and without additional transformation.
[0069] In a preferred embodiment, the configuration is updated in real time, i.e., updated configuration parameters (e.g., unique identifiers) are continuously loaded, which increases the volume of data being processed and, as a result, increases the variety of enrichment options (due to updating or adding new data marts). The ability to enrich data from multiple constantly updated / added data marts helps to improve the accuracy of event processing in the anti-fraud system.
[0070] Since unique identifiers can be updated continuously (in real time) as part of a configuration update, there is no need to send update requests, which helps to improve the speed of event processing in the anti-fraud system.
[0071] In a particular implementation, data enrichment occurs from memory that contains data marts. The memory may be, but is not limited to, RAM or cache memory. In one of the particular implementation examples, an In-memory data grid (IMDG) is used to store data - an object-oriented key-value storage distributed in RAM with support for ACID transactions. Using IMDG allows for faster real-time operation with large memory structures.
[0072] In one of the alternative embodiments of the claimed solution, the data enrichment stage additionally contains sending a request for data enrichment, wherein the parameters of the enrichment request are taken from the configuration. A unique identifier, with the help of which data is loaded from the data marts, may serve as a request parameter, but is not limited to the specified example.
[0073] At step (105), the data from the event and / or data marts required for the model block are analyzed and processed, and the analyzed and processed data is added to the event.
[0074] The stage (105) of data analysis and processing includes checking the incoming event using the model filter (see Example No. 1).
[0075] For the previously presented incoming event (see Example #4), let's look at an example of analyzing and processing the data contained in this event. We check the conditions of the model filter: "esot Jilter" : "inList(f event. subchannel', ['ecom', 'UFS. WEBAPI' , 'ESA. WEBAPI']) AND isNotEmpty(f event. user id')” { avg sum last l O days, avg sum current day, cf ('event. deviceRequest. atm. merchantld)}
[0076] An analysis of the enriched event fields (see Example #4) showed that the event\subchannel field contains the value eсот and the event\user id field contains the value 10001 , i.e. it is not empty. Then two variables are calculated: • avg_sum last 10 days, • avg sum current day.
[0077] The variable avg_sum_last_10_days, which characterizes the average amount of customer spending over the last 10 days, is assigned the value cf ('variable Jrom data mart avg last 10 days') from the previously enriched data.
[0078] To calculate avg sum current _day=cardsTransactions. byUser (). avg() we will need data from the card transactions data mart cardsTransactions, which stores events only for the last 24 hours. We use the byUser() method, which selects all transactions by user (user Jd=" 10001 ') from the incoming event.
[0079] Example of data from the cardsTransactions data mart (Example No. 5 / • {clientTransactionld=000000517100277999230201120710, amount=500000, userId=10001, merchantName=MEGAMARKET, eventTime=1675253230, terminalClass=2, sender AccountNumber=null, beneficiarAccountNumber=null, subChannel=ISSUER_A CQUIRER, eventType =CHECK, subType=POS CARD VERIFICATION, tokenNumber=null, deviceId=20226242, seid=null, mcc=5200, terminalId^20226242. cvv2Data=null} • {clientTransactionId=5fd9d075baa740b8bb8bbl4523cd5ceb, amount=300000, userId=10001, merchantName=EAPTEKA, eventTime= 1675778231, terminalClass=null, senderAccountNumber=30233810000001170234, beneficiarAccountNumber=101000019772, subChannel=WEBACQUIRER, eventType = DEPOSIT, subType =SBP _C2B RETURN, tokenNumber=null, deviceId=20190264, 5ez7 / =null, mcc=1740, terminal!d=20190264, cvv2Data=null}
[0080] All data refers to the client from the incoming event: event\user_id=" 10001". The avg() method calculates the average amount of the client's transactions for two events retrieved from the cardsTransactions data mart. The result of the calculation for the avg sum current day variable is 400000.
[0081] The variable avg sum current day is added to the incoming event and can be accessed at any stage of the declared method, which allows reducing the time of analysis and additional recalculations.
[0082] Example #6 shows the event after adding the analyzed and processed data: event\deviceRequest\atm\merchantld= "101000015560" event\deviceRequest\atm\MCCGroup= "1" event\deviceRequest\atm\MCC=”5200" event\deviceRequest\atm\TerminalId= "20156357" event\deviceRequest\atm\AcquiringIIC="l 11295" event\deviceRequest\atm\terminal\terminalClass="8" event\user _id= "10001" event\channel= "ISSUER ” event\operation_amount =10256700 event\operation_geo_data = "35.456; 47.98" event\eventTime ="2023-08-1714:16:54” event\subchannel= "ecom" "Card number" = "12253" "Previous operation geodata” = ”99.555;100.321” variable Jrom data mart avg last _10_days=2340051 avg_sum_current_day=400000 avg_sum_last_l 0_days=2340051
[0083] In a particular example of implementation, predictors are additionally calculated that increase the predictive power of the model, based on the analyzed and processed data from the data marts, with at least the following acting as predictors: • aggregate variables (e.g. average / median customer spending per week / month / quarter / year), • demographic information (e.g. age, income, place of work / residence, credit rating), • reference information (black / gray lists; lists of invalid documents).
[0084] Predictors may be calculated in advance and loaded from data marts for subsequent use in analyzing and processing data from data marts or calculating other predictors (step 105), or calculated directly at step 105.
[0085] The use of predictors that increase the predictive power of the model helps to improve the quality of the final value of the processing result calculated with their help, and, as a consequence, to increase the accuracy of event processing in the anti-fraud system.
[0086] Let's consider the use of predictors using specific models from the block of models as an example.
[0087] Model of analysis of non-payment transactions in Internet banking.
[0088] Pre-calculated predictors are loaded from data marts, for example: • The client’s age and how long he has been a client of the bank. The client's age is characterized by several parameters, the older the person, the more he has: o income, o accumulated funds, o available credit limit, these factors increase the likelihood of fraud. • How many times was the Internet banking application installed. In a normal situation, the application changes quite rarely (for example, when buying a new device). • Ratio of the number of devices and the number of applications: o if the device changes, the client installs a new application - this is a standard scenario, o abnormal behavior - they install the application and immediately carry out operations to withdraw / transfer funds, etc. o if the client uses a push-button phone - and suddenly bought a new phone, installed the application and began to carry out operations - the likelihood of fraud increases.
[0089] Using data from data marts, transactions over the last 24 hours are analyzed and predictors for the last 24 hours are generated. • How many different clients accessed the Internet bank from the same device, predictors for time intervals: day, three hours and one hour. Usually the client uses 1-2 devices, anything more than this number most likely indicates the fact that the fraudster has gained access to the client's credentials. Basically, the credentials are obtained using social engineering - the client himself transmits the login / password to connect to the Internet bank. Accordingly, the use of one device by different clients increases the likelihood of fraud. • Type of transaction (there are transactions that do not involve fraud, and separate transactions that are highly likely to be fraudulent). The sequence of operations per day is also taken into account; usually the client performs the same actions, for example: o go to the Internet bank and check the balance, o go to the Internet, check the balance and top up the account. In case of fraudulent actions, the picture changes dramatically: o the transaction amount increases, o the frequency of transfers, o the same type of transfer operations.
[0090] The calculated predictors and defined parameters from the operation are transferred to the model for subsequent analysis and calculation of the final value of the processing result.
[0091] Model of “customer operations at a bank branch” (opening / closing an account; payments; payment for services; issuing loans, etc.).
[0092] Pre-calculated predictors are loaded from data marts, for example: Sleeping client indicator: • Age of the client. Older clients have more money than younger clients. Also, this age group is approved for a much larger amount of credit, which fraudsters can ask the client to take out. Thus, for older ages, the likelihood of fraud increases. • The period of communication between the client and the bank branch. The client usually makes transactions in the same bank branch. In case of fraud, the client usually visits new branches.
[0093] Using data from data marts, transactions over the last 24 hours are analyzed and predictors for the last 24 hours are generated. • Number of loan applications over the past 24 hours. A customer may submit multiple loan applications under the influence of social engineering. • Number of unique IP addresses (client devices) over the last 24 hours. Before visiting a branch, the client or fraudster who has obtained the data can log into the Internet bank and perform a number of operations, for example, top up the account. This activity increases the likelihood of fraud. • Number of fraudulent calls per day. This indicator increases the likelihood of fraud. • Amount of incoming transfers to the client per day The client can make a transfer from his other accounts or the fraudsters can transfer funds to him for withdrawal (cashing out funds obtained by criminal means). • Number of calls - from the bank to the client or how many times the bank called the client per day. Used as a downward indicator of fraud, i.e. the client was informed of the fact of fraud or he reported the fact of fraud to the security service.
[0094] The calculated predictors and defined parameters from the operation are transferred to the model for subsequent analysis and calculation of the final value of the processing result.
[0095] At step (106), at least one event is sent for verification to the model block, which returns the final value of the processing result. Then, the final value of the processing result is added to the event.
[0096] In one particular example of implementation, the block of models contains at least one model that calculates the final value of the processing result.
[0097] In a preferred embodiment, the machine learning method used in the model block includes, but is not limited to, gradient boosting.
[0098] As an example, consider sending an event (see Example #6) to a model block, where the Jilter model calculates the final value of the processing result (score) taking into account all the parameter values contained in the event.
[0099] After processing, the model returns the final value of the processing result: есот ^filter model score _1 = 740
[0100] We add the resulting final value of the processing result, as well as the name of the model with which this value was calculated, to the event (Example No. 7): event\deviceRequest\atm\merchantId= "101000015560" event\deviceRequest\atm\MCCGroup= "1 ” event\deviceRequest\atm\MCC="5200" ecom filter model score l = 740 model_list = [“ecom_filter”]
[0101] Adding the specified parameters and their values to the message allows to reduce the event processing time, since the parameter can now be accessed directly without conversion.
[0102] In a particular example of implementation, step (106) additionally includes a comparison of the final value of the processing result with a threshold range of values, wherein, if the final value of the processing result corresponds to the threshold range of values, then additional enrichment of the data is performed.
[0103] The threshold range of values contains values that indicate suspicion of fraud (for example, the value of the model score l parameter falls in the range [400-749]), which necessitates additional enrichment of the data to confirm the fact of fraud.
[0104] If the final value of the processing result exceeds the threshold range of values (for example, the value of the model_score_l parameter falls within the range [750-1000]), then this indicates a high (critical) probability of recognizing the operation as fraudulent. Therefore, additional data enrichment is not required, since re-checking (by additional data enrichment and processing them with the model) will only confirm the conclusion made, but will increase the time for event processing.
[0105] In an alternative embodiment of the claimed solution, a re-check may be performed despite exceeding the threshold range of values for some cases, for example, for VIP client transactions or certain types of payments (mortgage, loan repayment).
[0106] If the final value of the processing result is less than the threshold range of values (for example, the value of the model_score_l parameter falls within the range [0-399]), then a conclusion is made about the absence of a fraudulent operation and additional data enrichment is not required.
[0107] In one of the particular examples of implementation, additional enrichment of data is carried out taking into account the configuration and the parameter correspondence table. In this case, in the preferred embodiment, the parameter correspondence table is included in the configuration (i.e. is part of the configuration).
[0108] The parameter mapping table contains information about: a. the model name, b. the data from the incoming event, c. unique identifiers to the data marts for additional enrichment, d. the weighting factor characterizing the relationship between the data "a" and "b" for the model. The higher the value of the weight coefficient, the higher the priority the data has for use in models and the more accurately the final value of the processing result is calculated.
[0109] Example of a parameter correspondence table (Example No. 8): / additional score / table congruence { "geo model table congruence" : [ "ecom Jilter": [ ["user id", "merchantld", 0, 7] , ["user id", "geo user id", 0.5] , ["terminalClass", "geo HardwarelD", 0,4] , ["terminalclass", "geo card number", 0,3] ]} [IT] In the example:
[0111] In a particular embodiment, a parameter correspondence table can be created using a model that has been pre-trained on historical data to identify relationships between parameters stored in data marts. Based on these relationships, a parameter correspondence table is built, and the table stores information on the strength of the relationship between the parameters (weight coefficient). Subsequently, the model has the ability to retrain when new data marts appear. If additional enrichment is required, the parameter correspondence table is accessed and the parameter that best matches the parameter from the event is found based on the largest weight coefficient (i.e. the strongest relationship between the parameters). Based on the found parameter, enrichment is performed from the data mart corresponding to the found parameter.
[0112] The parameter mapping table usage condition check filter allows you to configure the parameter mapping table usage condition and re-check the incoming event by the model.
[0113] Example of checking the condition of using the parameter mapping table:
[0114] If the final value of the processing result (esot Jilter model score ), obtained during the initial verification of the model, is less than 750, then it is necessary to use the correspondence of the parameters of the geo model table congruence.
[0115] At the previous stage (see Example #7), the model returned the final value of the processing result of Jilter model score 1 = 740, which is less than 750.
[0116] Next, the parameter correspondence table is checked (see Example No. 8) and the additional enrichment parameters are determined, which showed that the Jilter model is present in the parameter correspondence table.
[0117] Then the parameters are analyzed: 1) First entry from the parameter correspondence table: ["user id", "merchantld", 0, 7] Two parameters are present in the incoming event and have already been checked by the model block in the previous step, therefore these parameters cannot be used for additional enrichment. 2) Next, the following entry from the parameter correspondence table is analyzed: , ["user id", "geo user id", 0.5] the first parameter userjd is present in the incoming event, and the second geo user id is missing. 3) Then the configuration of the unique identifier 'geo user id' is checked. {"geo user id": "cf ('event. user jd')" , "geo HardwarelD "getFromJsonStringfcf event. Terminalld), 'HardwarelD ') " , "geo card number" : "cf ('Card number')"}
[0118] For the unique identifier geo user id there is data in the incoming message (cf('event.user_id)” = 10001). According to the value of the unique identifier field, additional enrichment is performed from the data mart geo data last _2 _days table : 10001 [userld_geo_data_last_2 days = ["90.00; 110.2", "98.5, 105.3", "80.5;90.3"]]
[0119] The userld_geo_data_last_2_days parameter and its value are saved in the event. After additional enrichment, the event is sent for re-checking to the model block with a new parameter obtained using the parameter mapping table.
[0120] Based on the results of the re-check, the model returns a new calculated final value of the processing result (equot ilter model score congruent _1 = 810), which is added to the event.
[0121] At step (107), a response message is generated based on the data stored in the event and / or data marts. In addition, when generating the response message, the necessary data is added to it, such as: • enriched data, • any data from the event, • calculated, transformed and analyzed data (e.g. predictors), • final values of the model processing results.
[0122] In a particular embodiment of the claimed invention, information from a data mart storing information on the historical final value of the processing result (historical score) can be added to the response message, for example, the last 10 final values of the processing result for a period of time (day / week). Based on the specified data, the average value and / or maximum coefficient are calculated. deviations from the average value, for subsequent comparison of the calculated deviation with the final value of the processing result from the current operation. The calculation data and comparison results can also be added to the response message.
[0123] In one of the particular implementation examples, the response message is additionally saved to the data mart. Saving the response message to the data mart allows: • update the data in the data marts used for enrichment, i.e. the accuracy of event processing is increased by enriching the data marts themselves, • increase the speed of updating data in the showcases, as a result, the declared solution will receive up-to-date information from the data showcases faster, thereby increasing the speed of event processing.
[0124] In one of the particular examples of the implementation of the declared solution, during the event processing, all data on this event is in the cache memory. After copying the event (separately or as part of the response message) to the data mart, the cache memory cleaning mechanism removes unused events from the cache memory. Using the specified cache memory cleaning mechanism allows preventing memory leaks.
[0125] In a specific implementation example, the response message is pre-stored in the data lake, from where it is sent to the data mart via a message broker.
[0126] At step (108), a decision is made to recognize the transaction as fraudulent based on the generated response message, and more specifically based on the data contained in the generated response message.
[0127] Checking an event to determine whether it is a fraudulent transaction may involve checking whether certain conditions are met, such as:
[0128] By checking the conditions, a decision is made about the presence or absence of a fraudulent transaction. In the example above: • if the calculated final values of the result of processing by the model ecom filter model score _1 or ecom ^filter model score congruent _1 are greater than 800, • if the time of the incoming event {event, event Time) is greater than 13:00:00, • if the amount of the incoming event (event. operation amount) is greater than 10000000 (in the format 100000.00, where the integer part characterizes the amount of the operation in rubles, and the fractional part is kopecks), • if the geodata (event. operation geo data) from the incoming event is not present in the historical information about geodata (userld_geo data_last_2_days) for the last two days, then a decision is made that the operation is fraudulent and the verification result code rule l = 1 is assigned, • if the incoming event does not meet the described conditions, then a decision is made to recognize the operation as not fraudulent and the verification result code rule l = 0 is assigned.
[0129] We receive parameters from the response message generated based on the data saved in the event (see Example No. 7): • eсот filter model score _1 = 740 does not meet the rule condition, • esot filter model score congruent _1 = 810 matches the conditions of the rule, since one of the two parameters matches the conditions of the rule, then the condition is satisfied, • cf('event.eventTime ) = "2023-08-17 14:16:54" matches the rule condition, the operation was performed after 13:00:00, • cf( 'event. operation amount) = 10256700 meets the condition, since the amount is greater than 10000000, • cf('event.operation_geo_data ) = "35.456;47.98" and userld_geo_data_last_2_days = ["90.00;110.2","98.5;105.3","80.5;90.3"] the client's current location does not match the geodata values for the last two days - meets the condition.
[0130] Since all conditions are met, the operation is considered fraudulent and the verification result code rule l = 1 is assigned.
[0131] In an alternative embodiment of the claimed solution, after a decision is made on the presence or absence of a fraudulent transaction, the verification result code is added to the response message. Thus, the response message may have the following form: event\deviceRequest\atm\merchantld= "101000015560” event\deviceRequest\atm\MCCGroup="l " event\deviceRequest\atm\MCC="5200" event\deviceRequest\atm \TerminalId= ”20156357" event\deviceRequest\atm\AcquiringIIC=" 111295" event\deviceRequest\atm\terminal\terminalClass=”8" event\user id= "10001 " event\channel= "ISSUER " event\operation_amount =10256700 event\operation_geo_data = "35.456; 47.98" event\eventTime ="2023-08-17 14:16:54" event\subchannel= "ecom" "Card number" = "12253" "Previous operation geodata" = "99.555;100.321" variable Jr from data mart avg last l 0 days =2340051 avg sum current day =400000 avg sum last 10 _days=2340051 model list = ["ecom ilter"] ecom Jilter model score 1 = 740 ecom Jilter model score congruent 1 = 810 userid geo data last _2 days = ["90.00; 110.2 ", "98.5; 105.3 ", "80.5; 90.3 "] rule l = "1"
[0132] In one of the particular variants of implementing the claimed solution, the attribute composition of the response message can be changed in real time, for example, using a response message generation filter.
[0133] Example of a response message generation filter: { "isNotEmpty(cf('event.user id')) AND isNotEmpty(cf('event.deviceRequest.atm.MerchantId')) AND (inList(f event. deviceRequest. atm.terminal.terminalClass', ['8', '2']) "avg sum" : "avg_sum_current_day", "risk score" : "model score 1", "resolution": "rule J", "user Id_geo data" : "user Id_geo data last _2 days" , "transaction eventTime " : "cf(' event. eventTime ') "}}
[0134] In this example, the filter checks three conditions that must be met simultaneously: • the 'event, user id' field from the incoming event must not be empty, • the 'event. deviceRequest. atm. Merchantld' field from the incoming event must not be empty, • the field 'event. deviceRequest. atm.terminal. terminalclass' from the incoming event, which is responsible for the device type (ecom, ATM, POS, etc.), must take one of the listed values: 8 or 2.
[0135] If all three conditions are met, the following fields are sent in the response message: • avg sum - the calculated value avg sum current day (average amount of customer spending for the current day) is assigned, • risk score - the final value of the processing result is assigned, recalculated by the model block есот ^filter model score congruent _1, • resolution - a value is assigned about the decision whether the operation is fraudulent or not rule l, • user Id_geo data - the calculated value userld_geo_data_last_2_days is assigned (information about the client’s geolocation during operations over the last 2 days), • transaction eventTime - the time of the transaction is assigned from the incoming event cf(' event. eventTime') .
[0136] Since all filter conditions are met, the response message is generated in the following form: { "avg sum" : 400000, "risk score" : "810", "resolution": "1", "user Id_geo data" : ["90.00; 110.2", "98.5; 105.3", "80.5; 90.3"], "transaction eventTime" : "2023-08-1714:16:54"}
[0137] In one of the particular implementation examples, the claimed solution additionally contains a stage of restricting access to a user performing at least one operation recognized as fraudulent. The following may serve as a restriction of user access, but are not limited to the specified examples: restriction / prohibition of user access to the target system; restriction of specific access rights; suspension and / or cancellation of operations performed by the user; adding the user to the blacklist; restriction / prohibition of client access to the Internet bank, or a combination thereof.
[0138] Despite the fact that the presented examples of implementation relate to the financial sphere, the claimed invention is not limited to application only in the specified area. The claimed invention can be applied in any field of technology where there is a need to protect sensitive data that may be subject to fraudulent influence, for example, health care, the social sphere, etc.
[0139] Fig. 2 shows a general view of a computing device (200), on the basis of which a device for processing events in an anti-fraud system can be implemented, ensuring the implementation of a method for processing events in an anti-fraud system.
[0140] In general, the computing device (200) comprises one or more processors (201) connected by a common information exchange bus, memory means such as RAM (202) and ROM (203), input / output interfaces (204), input / output devices (205), and means for network interaction (206).
[0141] The processor (201) (or several processors, multi-core processor) can be selected from a range of devices that are widely used at present, such as those from Intel™, AMD™, Apple™, Samsung Exynos™, MediaTEK™, Qualcomm Snapdragon™, etc. A graphics processor can also be used as the processor (701), such as those from Nvidia, AMD, Graphcore, etc.
[0142] RAM (202) is a random access memory and is intended for storing machine-readable instructions executed by the processor (201) to perform the necessary operations for logical data processing. RAM (202), as a rule, contains executable instructions of the operating system and the corresponding software components (applications, software modules, etc.).
[0143] ROM (203) is one or more permanent data storage devices, such as a hard disk drive (HDD), a solid-state drive (SSD), flash memory (EEPROM, NAND, etc.), optical storage media (CD-R / RW, DVD-R / RW, Blu-Ray Disc, MD), etc.
[0144] To organize the operation of the device components (200) and to organize the operation of external connected devices, various types of I / O interfaces (204) are used. The choice of the corresponding interfaces depends on the specific design of the computing device, which may include, but are not limited to: PCI, AGP, PS / 2, IrDa, FireWire, LPT, COM, SATA, IDE, Lightning, USB (2.0, 3.0, 3.1, micro, mini, type C), TRS / Audio jack (2.5, 3.5, 6.35), HDMI, DVI, VGA, Display Port, RJ45, RS232, etc.
[0145] To ensure user interaction with the computing device (700), various I / O information devices (205) are used, such as a keyboard, display (monitor), touch display, touchpad, joystick, mouse, light pen, stylus, touch panel, trackball, speakers, microphone, augmented reality tools, optical sensors, tablet, light indicators, projector, camera, biometric identification tools (retina scanner, fingerprint scanner, voice recognition module), etc.
[0146] The network interaction means (206) ensures the transmission of data by the device (200) via an internal or external computer network, for example, Intranet, Internet, LAN, etc. One or more means (206) may be, but are not limited to: Ethernet card, GSM modem, GPRS modem, LTE modem, 5G modem, satellite communication module, NFC module, Bluetooth and / or BLE module, Wi-Fi module, etc.
[0147] Additionally, the device (200) may also include satellite navigation tools, such as GPS, GLONASS, BeiDou, Galileo.
[0148] The submitted application materials disclose preferred examples of the implementation of the technical solution and should not be interpreted as limiting other, particular examples of its implementation that do not go beyond the scope of the requested legal protection, which are obvious to specialists in the relevant field of technology.
[0149] Sources of information: [1] Federal Law of 27.06.2018 No. 167-FZ “On Amendments to Certain Legislative Acts of the Russian Federation in Terms of Combating the Theft of Funds” [2] Federal Law of 27.06.2011 No. 161-FZ “On the National Payment System” [3] Federal Law of July 24, 2023 No. 369-FZ “On Amendments to the Federal Law “On the National Payment System” [4] GOST R 57580.1-2017. Security of financial (banking) transactions. Protection of information of financial organizations. Basic composition of organizational and technical measures [5] Creation of a system to combat fraudulent transactions, https: / / crosstech.su / proj ects / sistemy-protivodeystviya-moshennicheskim-tranzaktsiyam / ?ysclid=lo4abk3je652194739 [6] Fraud prevention system. https: / / cMapTBHCTa^ / fraud-management / [7] What is a data mart? https: / / www.oracle.com / cis / autonomous-database / what-is-data-mart / [8] What is a data lake? https: / / www.oracle.com / cis / big-data / data-lake / what-is-data-lake / [9] What are Message Brokers? https: / / www.ibm.com / topics / message-brokers
Claims
FORMULA 1. A computer-implemented method for processing events in an anti-fraud system, executed by at least one processor, comprising the steps of: a) loading at least one configuration representing a set of parameters for data enrichment, b) receiving at least one incoming event, c) parsing at least one received event and converting it into a given format, d) enriching data taking into account the configuration and adding the enriched data to the event formatted in step c), wherein the configuration is updated in real time, e) analyzing and processing data from the event and / or data marts necessary for the model block, and adding the analyzed and processed data to the event, f) sending at least one event for verification to the model block, which returns the final value of the processing result, wherein the final value is added to the event, g) generating a response message based on the data,stored in the event, h) make a decision to recognize the transaction as fraudulent based on the generated response message, 2. The method according to paragraph 1, characterized in that the incoming event is a message containing data in the form of a set of key-value pairs.
3. The method according to item 1, characterized in that the data enrichment occurs from memory containing data marts.
4. The method according to item 1, characterized in that the configuration contains a unique identifier, with the help of which data is loaded from data marts, wherein the unique identifier corresponds to at least one data mart, and the value of the unique identifier is obtained from an incoming event.
5. The method according to paragraph 1, characterized in that predictors are additionally calculated that increase the predictive power of the model, based on analyzed and processed data from data marts, wherein at least the following act as predictors: • aggregate variables, • demographic data, • reference information.
6. The method according to paragraph 1, characterized in that step f) additionally includes a comparison of the final value of the processing result with a threshold range of values, wherein, if the final value of the processing result corresponds to the threshold range of values, then additional enrichment of the data is performed.
7. The method according to item 6, characterized in that additional enrichment of data is carried out taking into account the configuration and the parameter correspondence table.
8. The method according to paragraph 1, characterized in that the response message is additionally protected in the data mart.
9. The method according to paragraph 1, characterized in that it additionally contains the step of restricting access of a user performing at least one operation recognized as fraudulent.
10. A device for processing events in an anti-fraud system, comprising at least one processor, at least one memory associated with the processor and containing machine-readable instructions which, when executed by at least one processor, ensure the execution of the method according to any of paragraphs 1-9.
Citation Information
Patent Citations
System and method for fraud detection using event driven architecture
US11803854B1
Multi-stage filtering for fraud detection
US20130024375A1
Real-time enrichment of raw merchant data from iso transactions on data communication networks for preventing false declines in fraud prevention systems
US20190147450A1
Enriching transaction request data for maintaining location privacy while improving fraud prevention systems on a data communication network with user controls injected to back-end transaction approval requests in real-time with transactions
US20210166238A1