Intelligent anomaly management system with real-time risk assessment

The multi-stage anomaly prevention system in e-commerce platforms uses machine learning and identity linking to detect and prevent return fraud and abuse, achieving significant loss reduction with efficient latency and resource usage.

WO2026161554A1PCT designated stage Publication Date: 2026-07-30CHEWY INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
CHEWY INC
Filing Date
2026-01-22
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Conventional e-commerce platforms struggle with detecting complex anomalous activities such as return abuse and fraud, particularly due to the reliance on static rules and limited data, leading to false positives and missed fraudulent activities, including policy abuse and counterfeit returns, which significantly impact retailers' profitability.

Method used

A multi-stage anomaly prevention system utilizing machine learning, historical transaction data, and identity linking with graph-based representation to detect and deter anomalous activities, incorporating real-time customer behavior analysis and enhanced identity linking to identify and prevent return fraud and abuse.

Benefits of technology

The system effectively reduces anomalous activity losses by over 90% with latency under 100 milliseconds for return initiation and 150 milliseconds for refund decisions, maintaining CPU usage under 3% and memory utilization under 25%, ensuring customer satisfaction while minimizing exposure to fraud.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2026012142_30072026_PF_FP_ABST
    Figure US2026012142_30072026_PF_FP_ABST
Patent Text Reader

Abstract

An anomalous activity detection system for e-commerce transactions can leverage distributed computing architecture for real-time anomaly prevention. The system may process and store relationship data in a graph database, maintaining consistency between real-time scoring and batch operations through a shared data layer. An identity linking module can analyze connections between transactions using graph-based algorithms, including breadth- first search and temporal filtering, to identify potential correlated anomalous activities. Natural language processing capabilities may evaluate customer communications through TF-IDF analysis and sequence mapping to generate risk indicators. Machine learning models can process these signals in real time at key transaction points, including return initiation and logistics scanning, to generate anomaly risk scores. A decision engine may evaluate these scores against predetermined thresholds, determining appropriate action for each return request. By combining advanced graph analytics, natural language understanding, and machine learning, the system can adaptively identify and prevent anomalous activities while maintaining legitimate customer experience.
Need to check novelty before this filing date? Find Prior Art

Description

Intelligent Anomaly Management System with Real-Time Risk AssessmentCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 748,814 filed January 23, 2025 entitled “Intelligent Return Abuse Management System with Real-Time Risk Assessment”, which is incorporated by reference herein in its entirety.TECHNICAL FIELD

[0002] The disclosed embodiments relate to e-commerce security and more specifically to systems and methods for e-commerce anomalous activity (e.g., return abuse) management.BACKGROUND

[0003] E-commerce has revolutionized the way people shop, offering convenience and accessibility to consumers worldwide. However, this digital transformation has also opened new avenues for anomalous activities such as, but not limited to, return abuse and other fraudulent activities. As e-commerce platforms evolve, so do the tactics employed by fraudsters, making it increasingly challenging for online retailers to protect their businesses and customers. Conventional e-commerce platforms face significant challenges from bad actors exploiting return policies. Traditional anomaly detection methods (e.g., abuse detection methods, fraud detection methods) rely heavily on static rules and limited data, often resulting in false positives or missed abuse activities. A common type of anomaly is policy abuse, where bad actors exploit lenient return policies by identifying scenarios where online retailers do not require sending back certain product types or items below a specific price threshold, causing significant revenue loss. Another prevalent anomalous activity is return fraud, where bad actors send back counterfeit products, empty boxes, or items from other retailers to claim refunds, further impacting retailers’ profitability. Conventional anomaly prevention methods fail to handle complex correlated anomalous activities (e.g., organized / coordinated fraud rings) and to identity -masking tactics employed by such bad actors.124489-5054-WO 1SUMMARY

[0004] Accordingly, there is a need for systems and methods that address at least some of the problems described above. The techniques described herein address at least some of the challenges described above by combining machine learning, historical transaction data, and / or identity linking for mitigating e-commerce policy and / or anomalous activities (e.g., return fraud abuse). Aspects of the system disclosed herein prevent anomalous activities such as e-commerce abuse, which may include, but is not limited to, policy abuse and / or counterfeit, empty box return abuse. Online retail systems can use these techniques to identify and / or deter bad actors exploiting return policies. Some embodiments may include a multi-stage anomaly prevention system that utilizes rules and / or machine learning models to detect and / or deter anomalous activities such as, but not limited to, policy and / or return abuse in e-commerce transactions at return initiation and / or refund processing points by assessing the risk level of each return request. Some embodiments may use historical data, real-time customer behavior, and / or enhanced identity linking with graph-based representation to detect suspicious behavior patterns and / or potential fraud rings. The techniques described herein allow for targeted friction, ensuring customer satisfaction for legitimate returns while reducing exposure to anomalous activity (e.g., return abuse activity). In some embodiments, the latency for return initiation decisions may be less than 100 milliseconds, and / or latency for refund decisions may be less than 150 milliseconds. In contrast, in conventional systems and methods, the estimated latencies can range from several seconds to minutes. Additionally, the techniques described herein may be implemented efficiently (e.g., overall CPU usage may be less than 3% and memory utilization may be under 25%). Some implementations of the techniques described herein showed over a 90% decrease in losses associated with anomalous activities such as, for example, return fraud and abuse losses.

[0005] Anomalies and / or anomalous activities as discussed herein may include, but are not limited to, data points, events, and / or patterns that deviate, within a desired threshold (e.g., significantly), from expected, normal, and / or historical behavior within a given system, population or context. Anomalies and / or anomalous activities as discussed herein may include fraud, misuse, abuse, and / or non-compliance with rules or policies. Anomalies and / or anomalous activities as discussed herein may include activities such as, but not limited to, an124489-5054-WO 2action, sequence of actions and / or pattern of actions that deviates from established norms and / or expected behavior suggesting potentially abusive, fraudulent, or otherwise irregular behavior. For example, an anomalous activity may be an activity or series of activities associated with return abuse or another fraudulent activity. In some embodiments, the terms “anomaly” and “anomalous activity” may be used interchangeably throughout the present application. In some embodiments, the anomalies and / or anomalous activities discussed herein are made in reference to an e-commerce platform. In some embodiments, anomalies and / or anomalous activities as discussed herein include, but are not limited to, return fraud abuse, also referred to as return fraud, return abuse, or refund theft. For the sake of brevity, aspects of the present disclosure are described in reference to return abuse and the detection and prevention thereof. However, it should be understood that aspects of the present disclosure may be used to detect and / or prevent other anomalous activities at, for example, an e-commerce platform.

[0006] In accordance with some embodiments, a computer-implemented system may be provided for detecting and / or preventing anomalies in e-commerce transactions. The system may include a computing architecture. The computing architecture may include one or more servers, a graph database, and / or a shared data layer. The one or more processors may be configured to perform real-time scoring. The graph database may be configured to store graph traversal paths. The shared data layer may be configured to maintain consistency between realtime and / or batch processing. The system may also include a data storage module, which may be configured to store and / or maintain: (i) a graph database schema defining node types including return origin identification, login network identification, physical address of origin, authorization token, and / or device identification for return transactions; (ii) relationship attributes including timestamps and / or anomaly labels for the return transactions; and / or (iii) feature vectors related to the return transactions. The system may also include an identity linking module, which may be configured to identify correlated anomalous activities (e.g., fraud rings) for obtaining graph-based linking features for the plurality of feature vectors by performing: (i) breadth-first search, based on the graph database schema, with predetermined maximum depth and / or configurable neighbor expansion; (ii) temporal filtering of network address-based relationships within predetermined time windows; and / or (iii) community detection algorithms. The system may also include a natural language processing module configured to generate communication-based features for the plurality of feature vectors using124489-5054-WO 3(i) Term Frequency-Inverse Document Frequency (TF-IDF) analysis, (ii) sequence mapping, and / or (iii) probability scoring, based on unstructured communications. The system may also include a real-time scoring module, which may be configured to generate anomalous activity risk scores by processing return transactions at initiation and / or real-time scans at logistics scan points, using a machine learning pipeline, based on the graph-based linking features and / or the communication-based features of the plurality of feature vectors. The system may also include a decision module, which may be configured to compare the generated anomalous activity risk scores to a predetermined threshold, and / or determine whether to process the return transactions.

[0007] In some embodiments, the real-time scoring module may include a machine learning pipeline, which may be configured to combine scores from a batch-scored anomalous activity model and / or a real-time anomaly detection model.

[0008] In some embodiments, the batch-scored anomalous activity model may be configured to maintain a 90:10 ratio of non-anomaly labeled data to anomaly labeled data through under-sampling during training.

[0009] In some embodiments, the real-time scoring module may further include a feature normalization module, which may be configured to maintain consistent scoring across the models.

[0010] In some embodiments, the shared data layer may include a centralized profile database that serves as a source of truth by storing real-time scoring updates.

[0011] In some embodiments, the shared data layer may further include a versioning system that manages updates to the database by checking for newer real-time updates within a predetermined time window before applying any batch changes.

[0012] In some embodiments, the shared data layer may further include a batch processing component, which may be configured to process incremental changes through the versioning system to update the centralized database without overwriting decisions.

[0013] In some embodiments, the identity linking module may include a graph processing engine, which may be configured to perform breadth-first search with a depth limit of a predetermined number of hops.124489-5054-WO 4

[0014] In some embodiments, the identity linking module may include a neighbor limiting component, which may be configured to restrict expansion of the search to a predetermined number of neighbors for high-volume nodes including shared IP addresses and / or shipping addresses.

[0015] In some embodiments, the identity linking module may include a temporal filtering component, which may be configured to only traverse IP node relationships where the timestamp difference between incoming and / or outgoing relationships is within a predetermined time.

[0016] In some embodiments, the identity linking module may include a graph structure that connects nodes representing order IDs, login IPs, shipping addresses, payment tokens, login devices, and / or email recipients through relationship attributes including payment, login, shipping, and / or gift card information.

[0017] In some embodiments, the identity linking module may be configured to identify correlated anomalous activities (e.g., organized fraud rings) through breadth-first search limited to 3 hops, temporal edge filtering within 60-day windows, weighted relationship scoring based on attribute overlap, and / or community detection.

[0018] In some embodiments, the NLP module may be configured to perform (i) text vectorization using TF-IDF computation, (ii) sequential pattern mining with configurable thresholds, (iii) anomalous activity indicator extraction using domain vocabularies, and / or (iv) real-time probability scoring with feature normalization.

[0019] In some embodiments, the system may provide monitoring of latency metrics across processing stages, measurement of graph traversal performance, feature importance analysis (e.g., using SHAP), and / or accuracy metrics with false positive / negative tracking.

[0020] In some embodiments, the correlated anomalous activity identification may include cycle detection algorithms, pruning of high-centrality nodes based on temporal patterns, community detection with configurable density thresholds, and / or graph updates with atomic edge modification.124489-5054-WO 5

[0021] In some embodiments, the system may provide optimization through caching with invalidation, feature computation pipelines, batched database operations, and / or query optimization for graph traversals.

[0022] In some embodiments, the decision module may be configured to apply predefined rules in combination with the generated anomalous activity risk scores and / or detected correlated anomalous activity, and / or determine the appropriate level of friction to apply to each return transaction.

[0023] In some embodiments, the real-time scoring module may be further configured to transmit a generated anomalous activity risk score for a return transaction to a computing device for review, and / or receive a decision from the computing device on whether to process or hold the return transaction.

[0024] In some embodiments, the system further includes a refund hold mechanism configured to activate when a returned product reaches a logistics provider for a return; apply additional risk assessment based on the pickup location, account history, and / or return details; and / or determine, based on the additional risk assessment, whether to hold the refund for further review.

[0025] In some embodiments, the data storage module may be configured to store features including at least one of: a total ratio of processed refunds, replacements and / or concessions; an expected count of refunds, replacements and / or concessions to be received; a maximum ratio of category refunds, replacements and / or concessions value over 365 days; an international chat IP address indicator; and / or a count of chat IP addresses linked to a network identifier within 7 days.

[0026] In some embodiments, the identity linking module may be configured to connect customer accounts by analyzing the shipping address, payment information, login IP, login device, and / or email attributes in the graph-based representation.

[0027] In some embodiments, the real-time scoring module may be configured to evaluate distances between shipping addresses and / or chat IP locations, and / or monitor international IP usage patterns.124489-5054-WO 6

[0028] In some embodiments, the real-time scoring module may be configured to analyze return patterns for specific product categories and / or value thresholds using data from the data storage module.

[0029] In another aspect, a computer-implemented system is provided for detecting and / or preventing anomalous activity in e-commerce transactions, according to some embodiments. The system may include a data storage module, which may be configured to store historical transaction data, customer behavior data, and / or identity information.

[0030] The system may also include an identity linking module, which may be configured to analyze the stored data to identify relationships between customer accounts and / or entities, generate a graph-based representation of the identified relationships, and / or detect potential correlated anomalous activities based on the graph-based representation. The system may also include a real-time scoring module, which may be configured to receive a return request from a customer, access the stored data to evaluate the risk of anomalies associated with the return request, and / or generate an anomalous activity risk score for the return request based on the evaluation. In some embodiments, the system may include a decision module, which may be configured to compare the generated anomalous activity risk score to a predetermined threshold, and / or determine, based on the comparison, whether to process the return request or hold the request for further review. The system may be configured to use the generated anomalous activity risk scores, the detected correlated anomalous activities, and / or predefined rules to selectively apply friction to return requests, ensuring a balance between response times and / or anomaly prevention.

[0031] In some embodiments, the real-time scoring module may be further configured to access a batch-scored anomalous activity model and / or an anomaly detection model, and / or combine the scores from the two models to generate the anomalous activity risk score for the return request. The batch-scored anomalous activity model and / or anomaly detection model may be trained on historical transaction data using decision tree algorithms.

[0032] In some embodiments, the real-time scoring module may be further configured to provide the generated anomalous activity risk score to a customer service representative for review, and / or receive the customer service representative’s decision on whether to process or hold the return request.124489-5054-WO 7

[0033] In some embodiments, the refund hold mechanism may be configured to activate when the returned product reaches a logistics provider for the return process, apply additional risk assessment based on the pickup location, account history, and / or return details, and / or determine, based on the additional risk assessment, whether to hold the refund for further review.

[0034] In some embodiments, the identity linking module may be further configured to use attributes including the shipping address, payment information, login IP, login device, and / or email to connect customer accounts, use graph algorithms, including breadth-first search and / or cycle detection, to identify potential correlated anomalous activities, and / or represent the identified relationships in a graph-based data structure.

[0035] In some embodiments, the system may further include a natural language processing (NLP) module configured to analyze customer communications, including chat, email, and / or call interactions, identify patterns and / or language indicative of potential anomalous activity, generate a probability score predicting the likelihood of anomalous activity, and / or provide the probability score as an input feature to the real-time scoring module.

[0036] In some embodiments, the decision module may be further configured to apply predefined rules in combination with the generated anomalous activity risk scores and / or detected correlated anomalous activities, and / or determine the appropriate level of friction to apply to each return request, balancing customer satisfaction and / or anomaly prevention.

[0037] In some embodiments, the data storage module may be configured to store a plurality of features related to customer transactions, device fingerprinting, geographic data, and / or interaction patterns with customer service.

[0038] In some embodiments, the data consistency between the real-time scoring module and / or batch processing components may be maintained through a shared data layer with a centralized customer profile database, incremental updates to the batch jobs to incorporate real-time changes, and / or a versioning system to resolve any data conflicts.

[0039] In some embodiments, the decision tree algorithm used in the machine learning models may be configured to handle class imbalance inherent in anomaly detection by undersampling non-anomaly labeled data to achieve a 90: 10 ratio between non-anomaly to anomaly.124489-5054-WO 8

[0040] In some embodiments, the graph-based identity linking module may be configured to store node types including order ID, login IP, shipping address, payment token, login device, and / or gift card email recipient, create edges between nodes with relationship attributes including paid with, login IP, login device, ship to, and / or email, and / or apply limited traversal depth and / or neighbor expansion to optimize performance for real-time anomaly detection.

[0041] In some embodiments, the graph-based identity linking module may be configured to use community detection algorithms to identify correlated anomalous activity.

[0042] In some embodiments, the graph-based identity linking module may be configured to limit traversals from IP nodes based on timestamp difference between incoming and / or outgoing relationships to eliminate noise from mobile IPs.

[0043] In some embodiments, the decision module may be configured to apply a combination of the generated risk scores, detected correlated anomalous activity, and / or predefined rules to determine the appropriate level of friction for each return request, and / or ensure that legitimate customers receive a satisfactory experience while minimizing exposure to anomalous activity.

[0044] In another aspect, an electronic device includes one or more processors, memory, a display, and / or one or more programs stored in the memory. The programs are configured for execution by the one or more processors and / or are configured to perform any of the methods described herein, according to some embodiments.

[0045] In another aspect, a non-transitory computer-readable storage medium stores one or more programs configured for execution by a computing device having one or more processors, memory, and / or a display. The one or more programs are configured to perform any of the methods described herein, according to some embodiments.

[0046] Thus, methods, systems, and / or interfaces are disclosed for detecting and / or preventing anomalous activity in e-commerce transactions.

[0047] Both the foregoing general description and the following detailed description are exemplary and explanatory and are intended to provide further explanation of the invention as claimed.124489-5054-WO 9BRIEF DESCRIPTION OF THE DRAWINGS

[0048] For a better understanding of the aforementioned systems, methods, and graphical user interfaces, as well as additional systems, methods, and graphical user interfaces that provide data visualization analytics, reference should be made to the Description of Embodiments below, in conjunction with the following drawings in which like reference numerals refer to corresponding parts throughout the figures.

[0049] Figure 1 is a block diagram of an example system for detecting and / or preventing anomalous activity in e-commerce transactions, according to some embodiments.

[0050] Figure 2 illustrates an example anomaly identified by a cycle detection algorithm, according to some embodiments.

[0051] Figure 3 shows an example end-to-end anomalous activity prevention process flow, according to some embodiments.

[0052] Figure 4 is a schematic diagram of an example anomalous activity model realtime scoring process, according to some embodiments.

[0053] Figure 5 is a schematic diagram of an example anomaly detection system, according to some embodiments.

[0054] Figure 6 illustrates an example use case for identifying and / or preventing anomalous activities, according to some embodiments.

[0055] Figure 7 illustrates an example decision tree, according to some embodiments.

[0056] Figure 8 is a schematic diagram of an example architecture for integration of real-time scoring model with batch processing, according to some embodiments.

[0057] Figure 9 is a flowchart of an example method for detecting and / or preventing anomalous activity in e-commerce transactions, according to some embodiments.

[0058] Figure 10 is a flowchart of another example method for detecting and / or preventing anomalous activity in e-commerce transactions, according to some embodiments.

[0059] Reference will now be made to embodiments, examples of which are illustrated in the accompanying drawings. In the following description, numerous specific details are set124489-5054-WO 10forth to provide a thorough understanding of the present invention. However, it will be apparent to one of ordinary skill in the art that the present invention may be practiced without requiring one or more of these specific details.DESCRIPTION OF EMBODIMENTS

[0060] Figure 1 is a block diagram of an example system 100 for detecting and / or preventing anomalous activity in e-commerce transactions, according to some embodiments. The system 100 may be a distributed computing architecture, which may include application servers 106 and / or database servers 118, and / or a refund hold mechanism 130. The application servers 106 may include a data storage module 108, an identity linking module 110, an optional natural language processing module 112, a real-time scoring module 114 (also referred to herein as real-time analysis module 114), and / or a decision module 116. The application servers 106 may perform real-time processing with latencies between 50-150 milliseconds, for example. The database servers 118 may include a graph database server 120, a shared data layer server 122, a centralized profile database 124, a batch processing module 126, and / or an optional version control module 128.

[0061] In some embodiments, the refund hold mechanism 130 may activate when returned products reach logistics providers for returns, apply additional risk assessment based on the pickup location, account history, and / or return details, and / or determine whether to hold refunds for further review. The additional risk assessment may include, for example, determining if the return shipping state (the state in which the return is being shipped back from) is different than the state which the package is shipped to. The additional risk assessment may also include if the return shipping state is different from a bordering state of the state which the package is shipped to. The additional risk assessment may include, for example, determining if the return shipping package has more than one label on it (often used to confuse a courier company used to ship the package), which may delay the package from arriving at one of the fulfillment centers. At return scan, for example, an integration point may activate when the product reaches a return point (e.g., FedEx) for the return process. The system may apply additional risk assessment based on the pickup location, account history, and / or specific return details, such as product and / or amount. A set of heuristics may be combined with modelbased scoring to determine if a refund should be held for further review, providing an added124489-5054-WO 11layer of anomaly prevention. In some embodiments, at return initiation, the system 100 may evaluate if a return request warrants sending the product back before a refund is processed. Two or more models may provide scores: a batch-scored anomalous activity model (e.g., a batch-scored return abuse model) and / or an anomaly detection model (e.g., an order fraud anomaly detection model). These models may use decision trees trained on historical anomalous patterns (e.g., historical fraud and / or fraudulent activity patterns) and / or use a plurality of features (e.g., 100 or more features), including IP address, device type, customer transaction history, and / or interaction history. The decision may include a customer service representative (CSR) review, adding a human-in-the loop component to mitigate any potential false positives.

[0062] The data storage module 108 may store and / or maintain: (i) a graph database schema defining node types including, for example, return origin identification, login network identification, physical address of origin, authorization token, and / or device identification for return transactions; (ii) relationship attributes including timestamps and / or anomaly labels for return transactions; and / or (iii) feature vectors related to return transactions. In some embodiments, the data storage module 108 may store features including the total ratio of processed refunds, replacements and / or concessions, expected count of refunds to be received, maximum ratio of category refunds over 365 days, international chat IP indicators, and / or count of chat IP addresses linked to network identifiers within 7 days. In some embodiments, the data storage module 108 may store historical transaction data, customer behavior data, identity information, and / or features related to customer transactions, device fingerprinting, geographic data, and / or interaction patterns with customer service.

[0063] The identity linking module 110 may perform breadth-first search, community detection, and / or temporal filtering to identify correlated anomalous activities (e.g., organized fraud rings) and / or obtain graph-based linking features. The identity linking module 110 may analyze stored data to identify relationships between customer accounts and / or entities, generate graph-based representations of identified relationships, and / or detect potential correlated anomalies based on the graph-based representations. In some embodiments, the identity linking module 110 may perform breadth-first search with predetermined maximum depth (e.g., 2 hops or 3 hops) and / or configurable neighbor expansion limited to a124489-5054-WO 12predetermined number of neighbors (e.g., 10 neighbors). The predetermined number of hops may be determined based on an average number of linked customers for a given node, which is typically lower than 10. The number of hops (e.g., 2 hops or 3 hops) may depend on the starting node. If a customer is a starting node, then 2 hops may be sufficient to get linked customers. If a login device is a starting node, then the identity linking module 110 may use 3 hops, for example. In some embodiments, the identity linking module 110 may perform temporal filtering of IP node relationships where timestamp differences between incoming and / or outgoing relationships are within 2 months.

[0064] In some embodiments, the identity linking module 110 may connect customer accounts by analyzing the shipping address, payment information, login IP, login device, and / or email attributes in the graph-based representation. In some embodiments, the identity linking module 110 may include a graph structure that connects nodes representing order IDs, login IPs, shipping addresses, payment tokens, login devices, and / or email recipients through relationship attributes including payment, login, shipping, and / or gift card information. In some embodiments, the identity linking module 110 may identify correlated anomalies (e.g., coordinated and / or organized fraudulent activities such as fraud rings) through cycle detection algorithms, pruning of high-centrality nodes based on temporal patterns, community detection with configurable density thresholds, and / or graph updates with atomic edge modification. In some embodiments, the identity linking module 110 may use community detection algorithms, such as Louvain algorithm variants, to identify correlated anomalies. In some embodiments, the identity linking module 110 may limit traversals from IP nodes based on timestamp difference between incoming and / or outgoing relationships to eliminate noise from mobile IPs. In some embodiments, the system 100 may use identity linking to detect correlated anomalies such as, but not limited to, fraud rings by creating graph representations of interconnected accounts. This graph-based approach may be used to detect patterns of collusion among multiple accounts by identifying shared attributes, such as IP addresses or payment methods.

[0065] The natural language processing module 112 may generate communicationbased features using (i) Term Frequency-Inverse Document Frequency (TF-IDF) analysis, (ii) sequence mapping, and / or (iii) probability scoring, based on unstructured communications. In some embodiments, the natural language processing module 112 may analyze customer124489-5054-WO 13communications including chat, email, and / or call interactions; identify patterns and / or language indicative of potential anomalous activity; generate probability scores predicting the likelihood of anomalous activity; and / or provide these scores as input features to the real-time scoring module. In some embodiments, the natural language processing module 112 may perform text vectorization, sequential pattern mining with configurable thresholds, anomalous activity indicator extraction using domain vocabularies, and / or real-time probability scoring with feature normalization. Some embodiments may use a vocabulary, which may include fraud, scam, reseller, unauthorized, fake, stolen, account takeover, did not order, did not authorize, did not place, hacked, and / or suspicious. Some embodiments may identify return reasons and similarity in reasons that would not allow the fraudster / abuser to send the product back, for example: damaged, did not receive, does not want, and / or pet does not like. Some embodiments may detect patterns (e.g., identical language used for returns) in bad actor terminology, often with misspellings and bad grammar, for example: “My pet not like the product, please return.” Some embodiments may use these techniques to help identify anomalies such as, but not limited to, return fraud and / or abuse.

[0066] The real-time scoring module 114 may generate anomalous activity risk scores by processing return transactions at initiation and / or real-time scans at logistics scan points, using a machine learning pipeline, based on graph-based linking features and / or communication-based features. The real-time scoring module 114 may receive return requests from customers, access stored data to evaluate anomaly risk, and / or generate anomalous activity risk scores based on the evaluation. In some embodiments, the real-time scoring module 114 may combine scores from a batch-scored anomalous activity model and / or a realtime anomaly detection model, where these models may be trained on historical transaction data using decision tree algorithms. Some embodiments include boosted decision tree models trained on historical data (e.g., one year of data), using a plurality of features (e.g., over 100 features) that capture anomaly patterns. These models can identify, for example, anomalies based on transaction data, device fingerprinting, geographic data, and / or interaction patterns with customer service.

[0067] In some embodiments, the real-time scoring module 114 may include a feature normalization module to maintain consistent scoring across models. In some embodiments, the124489-5054-WO 14real-time scoring module 114 may evaluate distances between shipping addresses and / or chat IP locations and / or monitor international IP usage patterns. In some embodiments, the realtime scoring module 114 may analyze return patterns for specific product categories and / or value thresholds. In some embodiments, the real-time scoring module 114 may transmit generated risk scores to a computing device for review, provide scores to customer service representatives for review, and / or receive decisions on whether to process or hold return transactions.

[0068] The decision module 116 may compare generated anomalous activity risk scores to predetermined thresholds and / or determine whether to process return transactions or hold requests for further review. In some embodiments, the decision module 116 may apply predefined rules in combination with the generated anomalous activity risk scores and / or detected correlated anomalous activities to determine appropriate levels of friction for return transactions. In some embodiments, the decision module 116 may balance customer satisfaction and / or anomaly prevention, ensuring legitimate customers receive a satisfactory experience while minimizing exposure to anomalous activity (e.g., abuse activity). In some embodiments, the decision module 116 may use the generated anomalous activity risk scores, detected correlated anomalous activities, and / or predefined rules to selectively apply friction to return requests, ensuring balance between response times and / or fraud prevention. Some embodiments may apply one or more forms of frictions within the returns process. A first form of friction may include requiring the package to be returned to the fulfillment center(s). A second form of friction may include holding the return payment until the package is received by the fulfillment center(s) and contents of the package are validated. Validation may include, for example, determining if the product was the intended product to receive back, if the quantity is correct, if an item is a counterfeit item, and / or if the item is heavily used.

[0069] The graph database server 120 may store graph traversal paths and / or maintain node types including, for example, order ID, login IP, shipping address, payment token, login device, and / or gift card email recipient. The graph database server 120 may store and / or manage identity and / or relationship data traversal paths.

[0070] The shared data layer server 122 may maintain consistency between real-time and / or batch processing through a centralized profile database. In some embodiments, the124489-5054-WO 15shared data layer server 122 may include a versioning system that manages updates by checking for newer real-time updates within predetermined time windows (e.g., 60-day lookback period) before applying batch changes. In some embodiments, the shared data layer server 122 may maintain data consistency between real-time scoring and / or batch processing components through incremental updates to batch jobs and / or conflict resolution.

[0071] The centralized profile database 124 may serve as a source of truth by storing real-time scoring updates.

[0072] The batch processing module 126 may process incremental changes through the versioning system to update the centralized database without overwriting decisions. In some embodiments, the batch-scored models within the system may maintain a 90:10 ratio of nonanomaly labeled data to anomaly labeled data (e.g., non-fraud labeled data to fraud labeled data) through under-sampling during training. The non-anomaly data may be defined as orders deposited and turned out to be non-anomalous, not an unauthorized transaction, chargeback, and / or declined / canceled as an anomalous order. Some embodiments may use random sampling with substitution for under-sampling.

[0073] The optional version control module 128 may manage updates and / or resolve conflicts between real-time and / or batch processing components. In some embodiments, the system 100 may provide optimization through feature computation pipelines, batched database operations, and / or query optimization for graph traversals.

[0074] The system 100 may provide monitoring of latency metrics across processing stages, measurement of graph traversal performance, feature importance analysis, and / or accuracy metrics with false positive / negative tracking. The system 100 may provide monitoring of latency metrics across processing stages, measurement of graph traversal performance, feature importance analysis (using SHAP), and / or accuracy metrics with false positive / negative tracking. Some embodiments may include multiple components and / or workflows, including real-time model scoring, batch processing, and / or identity linking.Example Natural Language Processing (NLP)-Driven Workflow for Anomalous Activity Detection124489-5054-WO 16

[0075] Some embodiments integrate a Natural Language Processing (NLP) model to analyze customer communications and / or predict potential anomalous activity. The NLP processing module 112 may include the NLP model. This model may process data from customer interactions and / or provide essential insights that feed into the overall anomalous activity detection system. Some embodiments may include data collection, which may gather text data from customer contacts across chat, email, and / or call interactions. Some embodiments may include text processing, which may apply TF-IDF (Term Frequency-Inverse Document Frequency) to identify important terms and / or patterns in customer communication. Some embodiments may include sequence mapping, which may analyze sentence structures and / or sequences to detect suspicious language or behavior indicative of potential anomalous activity. Some embodiments may include probability scoring, which may generate a probability score predicting the likelihood that the customer is a bad actor committing an anomalous activity (e.g., the customer is engaged in return abuse) based on communication patterns. Some embodiments integrate the NLP model output with the ML model. For example, NLP scores from the NLP model may be used as features for the machine learning model, enhancing the ability of the ML models to detect and / or prevent anomalous activity. The workflow described herein may incorporate customer communication data into predictive models, increasing accuracy in identifying anomalous behavior.Example Identity Linking and Correlated Anomalous Activity Detection

[0076] Some embodiments may use an identity linking mechanism using graph representations of customer data. The identity linking mechanism may be implemented as part of the identity linking module 110. The identity linking mechanism may connect accounts sharing similar attributes, allowing the system to uncover potential correlated anomalous activities (e.g., fraud rings) that may otherwise remain undetected. Some embodiments may use attributes, such as the shipping address, payment information, login IP, login device and / or email information to link identities to each other. The identity linking mechanism may use graph algorithms, such as breadth-first search with limited hops and / or cycle detection to detect correlated anomalous activity.124489-5054-WO 17

[0077] Figure 2 illustrates an example anomalous activity (e.g., a fraud and / or abuse ring) 200 identified by a cycle detection algorithm, according to some embodiments. The cycle detection algorithm may be implemented in the identity linking module 110. The diagram shows interconnected components forming a circular pattern of related activities. The connection flow begins with a payment identifier 202 (“Payment xxx61”) which connects to customer 204 (“Customer xxxxl49”, marked as return fraud) through an order payment relationship. This customer connects to an address 206 (“Address - xxx NJ”) through an order address relationship. The address then links to another customer 208 (“Customer xx63”, also marked as return fraud) through another order address relationship. Customer 208 connects through a login success event to a login IP 210 (“Login IP, xx211”). This IP address connects through another login success to customer 212 (“Customer xx014”, marked as experienced ATO). This pattern continues through login IP 214 (“Login IP, xxl76”) to customer 216 (“Customer xx743”, also marked as experienced ATO), and / or finally to login IP 218 (“Login IP, ex. 101”). Some of the nodes are coded by line type (e.g., as illustrated in the key in Figure 2) to indicate their roles: solid line nodes (204, 208) indicate return fraud customers, shorter dashed line nodes (212, 216) represent customers who experienced account takeover (ATO), and longer dash line nodes (206, 210, 214, 218) represent login activities prior to ATO. This visualization demonstrates how the system’s cycle detection algorithm can identify complex correlated anomalies (e.g., fraud rings) by connecting disparate activities through shared attributes like addresses, login IPs, and / or payment information.

[0078] In the example illustrated in Figure 2, observations may be found indicating potential anomalous activities. For example, return fraud and / or ATO orders may be observed in Figure 2. In Figure 2, the shipping address and login IP address may be based in a shared geographical location (e.g., in the same street, city, county state, country). In Figure 2, it may also be observed that login timestamps are within 2 months of each other and occur around (e.g., within a matter of minutes) when the order is placed. In Figure 2, based on at least the observations outlined above, it may be determined that the IP address is part of a correlated fraudulent activity (e.g., a potential credential stuffing pattern).

[0079] The correlated anomalous activity example shown in Figure 2 illustrates cycle detection, according to some embodiments. In some embodiments, what is shown in Figure 2124489-5054-WO 18illustrates an organized fraud ring orchestrated by a group of users and / or user accounts coordinating efforts and / or resources to commit fraudulent activity. Other common graph patterns beyond cycles may indicate correlated anomalous activities such as, but not limited to, fraud rings. The other graph patterns include breadth-first search traversal to identify if an anomalous activity (e.g., return abuse, fraud link) exists for a given customer or related attributes. Graph community detection may be used to identify relationships and / or nodes that form a distinct group within the graph and / or have characteristics of anomalous behavior such as, but not limited to, those associated with fraud or abuse, within the community. The graph structure improves upon traditional relational database approaches for identity linking. For example, graph representation can identify relationships between identities at a deeper level efficiently compared to relational databases. Relational databases may be used until 2 hops, but beyond that the performance can severely degrade as it is not meant for the graph structure. Graph-based traversals and / or algorithms provide a better approach to connect identities that otherwise would not be identified as connected with relational databases.

[0080] The identity linking and / or data processing techniques can be used for the downstream machine learning analysis in several ways. For example, the graph-based identity linking, temporal filtering, and community detection algorithms create high-quality engineered features that capture complex relationships between entities. These rich features may allow downstream ML models to better identify sophisticated anomaly patterns that would be missed by looking at individual transactions in isolation. The approach may also help address one of the challenges in anomaly detection, which is the inherent class imbalance where legitimate transactions vastly outnumber anomalous (e.g., fraudulent) ones. By identifying correlated anomalous activity and their associated characteristics through graph analytics, the system can ensure the anomaly examples in the training data capture actual correlated anomalous activity (e.g., coordinated fraud / abuse) rather than just random anomalies. The 90:10 ratio of nonanomaly to anomaly data, for example, becomes more meaningful when the anomaly cases represent these genuine patterns of anomalous activity. Additionally, the temporal filtering of relationships (especially for IP addresses) and neighbor expansion limits may help reduce noise in the input data. This improved signal-to-noise ratio can lead to more robust ML models that are less likely to learn spurious patterns or overfit to temporary anomalies. The optimized graph traversal methods, with their limited depth and configurable neighbor expansion, can help124489-5054-WO 19ensure that these complex relational features can be generated quickly enough for real-time ML inference while still maintaining their predictive power.

[0081] With identity linking and ring detection, the system described herein can connect identities across deeper links with any combination of attributes, which is difficult to perform with relational structure. The ring detection may provide the ability to provide additional evidence to investigators (manual or automatic) on the potential links between several types of anomalies that they need to be aware of. Both identity linking and ring detection representations may be used to calculate features for the return model to improve the anomalous activity detection capabilities. With conventional systems, the anomalous actors (e.g., abusers) likely know that controls are in place with traditional linking, so the abusers could try to get around the controls. Graph representation for identity linking may be used to stop and / or detect such abusers.Example End-to-End Anomalous Activity Prevention Process Flow

[0082] Figure 3 shows an example end-to-end anomalous activity prevention process flow 300, according to some embodiments. The real-time scoring module 114 may include an anomalous activity model for real-time scoring illustrated in Figure 3. The model may include the capability for assessing customer behavior dynamically, and / or updating the risk profile at specific touchpoints. The process flow 300 may begin when customers 308 initiate a return, generating a return request 310. The real-time scoring module 114 and / or the decision module 116 may perform fraud system pre-return analysis 312 through the anomaly system 306. Based on this analysis, a decision 314 may determine if a product return is required. If yes, the system may generate a return shipping label 316. After the customer returns the package, a return shipping scan 318 may trigger another touchpoint where the real-time scoring module 114 may perform post-return analysis 320. The decision module 116 may then evaluate whether to hold the refund at decision point 322. If a hold is determined necessary, the system may route the return for fulfillment center review 324. The fulfillment center may check if good product was received at decision point 326. Based on this verification: if no, a refund may be cancelled or partial refund issued 328; if yes, or if no hold was required from decision point 322, a refund may be issued 330. Throughout this process, the natural language processing module 112 may124489-5054-WO 20analyze customer communications, while the identity linking module 110 may evaluate connections to known anomaly patterns. The data storage module 108 may store all transaction details and / or outcomes, while the database servers 118 maintain data consistency between real-time decisions and / or batch processing.Example Anomalous Activity Model Real-Time Scoring

[0083] Figure 4 is a schematic diagram of an example anomalous activity model realtime scoring process 400, according to some embodiments. The real-time scoring module 114 may perform the real-time scoring process 400. Figure 4 illustrates capability of one or more models in the real-time scoring module 114 in assessing customer behavior dynamically, updating the risk profile at specific touchpoints. The system 100 may monitor and / or assess various customer touchpoints and / or interaction types to dynamically update risk profiles. The system 100 may process multiple input events including order completion events 404, return creation events 406, login completion events 408, and / or chat session creation events 410. These events may feed into an anomalous activity detection model 412.

[0084] The anomalous activity detection model 412 may process these inputs through several feature sets. Customer interaction features 414 may analyze customer service interactions, while connected linked features 416 may incorporate graph-based identity linking results. Overall anomaly features 418 may evaluate general anomaly patterns, and / or customer-related NLP score 420 may process communication patterns. The model pipeline continues through an XGBoost model 422 that processes these features, followed by model scoring 424 that generates risk scores. Database integration 426 maintains scoring history, while the anomaly management system 430 makes decisions based on the generated scores. Additional touchpoints may include FedEx scan events 432 that may trigger new risk assessments and / or return initiation events 428 that may start the process. The real-time scoring module 114 may perform these assessments within 50-150 milliseconds at each touchpoint, updating risk profiles dynamically as new information becomes available. Throughout this process, each component may interact with the others to maintain real-time risk assessment capabilities while ensuring consistent data processing and / or storage. This allows the system 100 to adapt quickly to new information while maintaining its accuracy in anomaly detection.124489-5054-WO 21Example Use Cases of Anomaly Prevention System

[0085] Figure 5 is a schematic diagram of an example anomaly detection system 500, according to some embodiments. The anomaly detection system 500 may be an example implementation of the system 100. The system 100 may identify and / or prevent anomalous activities throughout a customer lifecycle, as illustrated in Figure 5. The system 100 may process input from multiple services, for example: an order service 504 may handle order creation events, a return service 506 may process return requests, and / or a customer service UI 508 may manage customer interactions. These services shown in Figure 5 may be part of the application servers 106. The anomaly detection system 500 shown in Figure 5 may include several components that may correspond to modules from Figure 1. The anomalous activity customer profile tagging 512 may be part of the data storage module 108. The anomaly detection system 500 may process various rules through components including linking rules, empty -box received tracking, and / or return-to-order ratio analysis 514, which may be executed by the decision module 116.

[0086] The refund hold component 516 may operate with the real-time scoring module 114 to determine when to pause refunds for review. The anomaly engine 518 and / or anomalous activity detection model 520 may be part of the real-time scoring module 114, processing events and / or generating risk scores in real time. The system 100 may flag anomalous activity through component 522, which may update customer profile data serving 524. These components may interact with the centralized profile database 124 to maintain customer risk profiles and / or transaction histories. Throughout the process, the system 100 may evaluate order events, return submissions, received items, and / or customer interactions to dynamically assess risk and / or prevent anomalous activities. The identity linking module 110 may analyze connections between these events to detect patterns of anomalous activity, while the natural language processing module 112 may analyze customer communications at various touchpoints.

[0087] Figure 6 illustrates another example use case 600 showing how the system 100 may identify and / or prevent anomalous activities throughout the customer lifecycle, according to some embodiments. The system 100 may collect and / or process multiple data streams related124489-5054-WO 22to customer orders: order shipping address stream 602, order payment stream 604, order gift card email stream 606, order login Internet Protocol (IP) stream 608, and / or order login device stream 610. These data streams may feed into a graph data store 612, which may correspond to the graph database server 120. The graph data store 612 may support the anomaly detection capabilities of the identity linking module 110 through breadth-first search and / or community detection. The anomaly detection models (e.g., fraud and / or returns abuse detection models) with features and / or inference component 614 may represent the combined functionality of the real-time scoring module 114 and / or natural language processing module 112.

[0088] The anomalous activity service 616 may operate with the decision module 116 to determine appropriate actions based on detected patterns. The order management service 618 may process ongoing transactions while monitoring for suspicious patterns. The returns service 620 may handle return requests while integrating with the real-time scoring module 114 for risk assessment. These services may interact continuously, allowing the system 100 to maintain real-time anomaly detection throughout the customer lifecycle. The graph data store 612 may maintain relationships between entities, enabling the system to detect complex anomaly patterns across multiple accounts and / or transactions.Example Decision Tree for Anomalous Activity Detection

[0089] Figure 7 illustrates an example decision tree 700, according to some embodiments. The decision tree may be used in the real-time scoring module 114 for anomalous activity detection. The tree may use XGBoost decision tree methodology to classify return requests based on features related to refunds, replacements, and / or concessions (RRC). The decision tree may begin with an initial split based on total_rrc_processed_ratio < 0.456768394. This ratio may represent a customer’s overall engagement with returns or concessions relative to their total orders. When the ratio is less than 0.456768394, the tree may evaluate should receive rrc count < 1.5, which may examine the number of RRCs expected at the fulfillment center. Alternatively, when the ratio is greater than or equal to 0.456768394, the tree may evaluate max_vulture_score_180d < 0.257102817, which may consider the maximum anomaly risk score from the past 180 days.124489-5054-WO 23

[0090] The decision tree 700 shows these branches continuing to split further, with nodes represented by ovals connected by solid and dashed lines indicating different decision paths. The leaf nodes at various depths may represent raw classification scores that the realtime scoring module 114 may use to generate risk assessments. The solid and dashed lines in the visualization may represent different decision paths, with the tree structure becoming more granular at deeper levels. The model may use these paths to classify return requests and / or determine their risk levels based on the combination of features including total_rrc_processed_ratio, should receive rrc count, and / or max_vulture_score_180d.

[0091] Some embodiments may use an XGBoost decision tree, which may use features related to refunds, replacements, and / or concessions (RRC) to classify data points, with leaf values representing raw scores for classification, example breakdown of the splits and their implications are described below, according to some embodiments. Suppose the initial split is total_rrc_processed_ratio < 0.456768394. A first decision point may use the ratio of processed refunds, replacements, or concessions to total orders. This feature may indicate the overall customer engagement with returns or concessions. Data points may be divided into 2 branches: (i) Yes (Ratio < 0.456768394): further split based on shouldreceive rrc count < 1.5, and / or (ii) No (Ratio > 0.456768394): directs to another split based on max_vulture_score_180d < 0.257102817. Subsequent splits may include branch 1 (Yes: shouldreceive_rrc_count < 1.5): this checks the number of RRCs expected to be received at the fulfillment center. Lower counts may signify lower involvement in anomalous activities or benign behavior. Branch 2 (No or Missing) may include max_vulture_score_180d < 0.257102817. This branch may use the maximum anomaly risk score from the vulture model in the past 180 days. A lower score may suggest a lower likelihood of anomalous activity, leading to further splits or classifications.

[0092] Features that contribute to the model may include total_rrc_processed_ratio (an indicator of a customer’s engagement with refunds and / or concessions, which may serve as a feature for an initial decision), shouldreceive rrc count (tracks the expected returns to be received at the fulfillment center, which may highlight suspicious patterns in returns handling), and / or max_vulture_score_180d (represents the historical anomaly risk score, which may provide a direct risk assessment for a customer based on past behavior).124489-5054-WO 24

[0093] Some embodiments may under sample the non-anomaly labeled data to obtain a 90:10 ratio of non-anomaly to anomaly, before sending the training data to XGBoost. XGBoost parameters for class imbalance may be adjusted.Example Signals, Features, Signal Analysis and Feature Importance

[0094] In addition to, or instead of, the features described above (total_rrc_processed_ratio, shouldreceive rrc count, max_vulture_score_180d), some embodiments may use other signals that may provide unique value for anomaly detection. For example, chat ip is intemational max indicates whether the chat IP address is international. Anomalous activities often involve international actors attempting to exploit return policies. This signal helps detect anomalies based on geographic inconsistencies. Another example is max_rrc_value_product_category_365d_ratio, which represents the maximum ratio of refunds, replacements, or concessions value to the product category in a predetermined time (e.g., the past 365 days). Anomalous actors (e.g., fraudulent actors, bad actors) tend to target high-value items disproportionally within specific categories. count_chat_ip_linked_cid_7d may track the number of chat IPs linked to a single customer ID in a predetermined time (e.g., the last 7 days). An unusually high count may indicate account sharing or anomalous activity exploiting customer service.

[0095] Some embodiments combine complementary signals to create a holistic anomaly risk profile. An example combination is total_rrc_processed_ratio + max_vulture_score_180d + max_rrc_value_product_category_365d_ratio. This combination may provide a layered view of customer behavior, balancing overall refund trends (total_rrc_processed_ratio), historical anomaly scores (max_vulture_score_180d), and / or disproportionate refund claims by category (max_rrc_value_product_category_365d_ratio). Another example combination is chat ip is intemational max + dist ship chat ip max + count_chat_ip_linked_cid_7d. These signals together may uncover inconsistencies between geographic locations, unusual distances between shipping and / or communication IPs, and / or abnormal account activities for identifying anomalous networks. Using one or more of these combinations, the system may create a multi-dimensional risk assessment that increases precision in detecting anomalies while minimizing false positives.124489-5054-WO 25

[0096] In practice, the combination of machine learning (ML) models and graph linking may detect anomaly patterns that may be missed by traditional rule-based systems The combination of ML models and graph linking offers a powerful approach to detect anomaly patterns that traditional rule-based systems alone would miss. The ML model identified an anomalous activity scenario where a new account may be created with a number and / or word combination, a suspicious email and / or an order placed through phone call with an agent. Traditional rules may not catch this subtle pattern if values hover just below hard thresholds. The model with the combination of features count_chat_ip_linked_cid_7d and / or csr_placed_flag was able to detect this subtle pattern. The machine learning model and / or the identify linking mechanism complement each other to improve overall detection accuracy. For example, graph linking reveals connections among accounts sharing the same IP address and / or payment methods. While individual accounts may appear legitimate, their collective behavior may indicate coordinated, correlated anomalous activities that rules and / or standalone ML models might overlook.

[0097] Return initiation, refund processing, and / or NLP analysis each may use several different features. Return initiation uses total_rrc_processed_ratio that indicates overall customer engagement with returns or concessions, which may predict suspicious behavior. Another feature shouldreceive rrc count reflects the expected number of returns to be received at a fulfillment center, which may flag patterns of potential anomalous activity. max_vulture_score_365d indicates a historical anomaly risk score that may indicate customer risk based on past behavior. These features may effectively capture key aspects of return behavior while maintaining a streamlined model structure for efficient real-time scoring. For refund processing, the features emphasize post-return activities and / or contextual signals, enabling the model to identify anomalies at the refund stage.

[0098] Such features may include (i) validation wait rrc count that tracks returns waiting for validation, which may indicate risky transactions; (ii) max_rrc_value_product_category_365d_ratio that highlights disproportionate refunds within certain product categories, a common anomaly tactic; and / or (iii) dist ship chat ip max that measures discrepancies between shipping address and / or chat IP location, flagging geographic inconsistencies. These features allow the model to focus on returns that deviate significantly124489-5054-WO 26from typical patterns, improving anomaly detection accuracy. The NLP model may utilize customer interactions via chat, email, and / or calls, converting unstructured data into meaningful inputs for anomaly detection. NLP features may include (i) TF-IDF Scores, which may indicate key terms and / or phrases associated with anomalous activity behavior, extracted using Term Frequency-Inverse Document Frequency analysis, and / or (ii) sequence analysis, which may include sentence structures and / or conversational patterns indicative of anomalous intent (e.g., fraudulent). Probability scores generated by the NLP model may be passed as inputs to the broader ML framework for holistic anomaly assessment. These features may help uncover behavioral cues from customer communications that traditional numerical features might miss.

[0099] For anomaly detection, total_rrc_processed_ratio feature measures the ratio of processed refunds, replacements, or concessions (RRC) to total orders. High values, such as >0.7, may be strong indicators of potential anomalous activity, as they reflect frequent engagement with return processes. SHAP values for this feature may rise sharply with increasing ratios, emphasizing its importance in flagging repeat offenders, shouldreceive rrc count feature tracks the number of returns expected at the fulfillment center. SHAP analysis highlights that counts exceeding 5 often indicate suspicious behavior, particularly when returns fail to match historical norms or are not physically received. This feature may identify anomalies in return patterns. max_vulture_score_180d feature derived from historical anomaly risk scores (e.g., scores over the past six months) may capture a customer’ s anomaly likelihood based on previous behavior. SHAP values spike for scores >0.7, correlating with customers exhibiting consistent anomalous activity patterns. This feature may provide a temporal view of risk and / or strengthen anomalous activity detection, dist ship chat ip max feature assesses geographic discrepancies between shipping addresses and / or chat IP locations. SHAP analysis shows a significant contribution, for example when distances exceed 500 miles, flagging potential location-related anomalies often linked to anomalous activity. It is particularly effective when combined with international IP indicators. max_rrc_value_product_category_365d_ratio feature calculates the refund-to-product-category ratio over a time (e.g., over a year). SHAP values highlight its significance when the ratio exceeds a certain threshold value that it learns from training, signaling targeted anomalous activity of specific high-value product categories.124489-5054-WO 27

[0100] Some embodiments may use the feature is_csr_placed_order, which indicates orders placed by customer service representatives. A non-obvious combination of this feature is with the feature CNT_NULL_REASON_365d, which is indicative of the count of returns that are received at a fulfillment center with product mismatch combined with is_csr_placed_order. Anomalous actors may exploit customer service representative channels to bypass automated checks, resulting in returns with incomplete justifications. These unexpected correlations may enhance the system’s ability to detect sophisticated anomaly patterns by leveraging interactions between unrelated signals.

[0101] Anomaly patterns can evolve over time and / or any feature that contributes to the system’s overall effectiveness may be included. Based on periodic SHAP analysis (e.g., daily analysis), which captures trends and / or variations across periods, one or more features may be added, updated, or removed, based on performance measurements, for example. Each feature provides unique insights that complement others, and / or the dynamic nature of anomalies such as, for example, fraud and / or abuse may require a comprehensive feature set to adapt to emerging patterns and / or tactics. While some features may have higher importance in specific scenarios, maintaining a diverse feature set may help ensure the system to remain resilient and / or effective against a broad range of anomalous activity strategies.

[0102] The multi-faceted approach, combining machine learning models, rule-based systems, and / or advanced data signals, helps address complex and / or evolving anomaly patterns. Specific behaviors that necessitate this approach include correlated anomalous activities such as, for example, fraud rings. Anomalous actors such as fraudsters operate in groups, creating multiple accounts with shared attributes (e.g., IP addresses, payment methods). These behaviors are difficult to detect with standalone rules but can be identified through machine learning and / or graph-based analysis. Additionally, such anomalous actors adapt to policy changes, targeting specific thresholds or product categories (e.g., high-value items with lenient return policies). Machine learning models identify these shifting patterns by analyzing historical data, while rules provide quick action for clear violations.

[0103] Example geographic and / or behavioral anomalies that may be handled by the system 100 may include large distances between shipping and / or chat IPs or frequent124489-5054-WO 28international IP usage. These require a combination of geographic analysis, behavior-based rules, and / or real-time ML scoring to flag inconsistencies effectively.

[0104] Suppose that a group of anomalous actors creates multiple accounts linked by shared IP addresses and / or payment methods. Individually, the accounts appear legitimate, and / or no single rule flags them as anomalous. To detect such correlated anomalous activities, the ML model may flag suspicious behaviors like excessive returns and / or high-value refunds across accounts. Graph analysis may connect the accounts through shared attributes, uncovering correlated anomalous activities. Rules may apply final restrictions, such as blocking additional returns. In this way, the correlated anomalous activity may be identified and / or stopped, preventing further anomalous activity.

[0105] Machine learning models may be combined with specific rules to stop targeted policy exploitation. For example, suppose that fraudsters identify certain types of SKUs because of their nature. The SKUs are rarely required to be sent back for returns. Further to the preceding example, the fraudsters may exploit this by repeatedly claiming refunds for such items. The ML model may detect a pattern of frequent returns with high refund-to-order ratios for low-value items. Rules may enforce stricter checks for specific product categories or SKUs. Combined, the ML models and the rules flag this behavior and restrict further refunds for the account. In this way, the anomaly pattern is caught early, protecting transactions without affecting legitimate customers.Example Query Optimization

[0106] Some embodiments of the identity linking module 110 may use limited traversal depth to optimize the performance for real time. Some embodiments may use breadth-first search depth of 2 or 3 depending on the starting node. Some embodiments may limit the neighbor’s expansion to 10, thereby improving performance for nodes that tend to be common, such as shared IP addresses or shipping addresses (e.g., apartment building). Some embodiments may limit traversal from an IP node only if the incoming relationship and / or outgoing relationship timestamp difference is within 2 months. This may help eliminate noise from mobile IP that get reused often. For community detection, some embodiments may use graph database in-built performant online algorithm for real-time processing. While doing124489-5054-WO 29graph traversals, the performance of the traversal can depend on neighbors attached to the node that is being queried upon. There may be nodes that have an unrealistic number of relationships, which may be worked around by placing limits and / or whitelisting, as the nodes create noise. IP nodes, for example, may create several relationships that may be redundant due to reusability of IPs in corporate or privacy ISPs. For mobility IPs, relationships that span over a longer period may be redundant, so some embodiments may add time difference as a factor to eliminate certain relationships that would otherwise create noise.

[0107] For storage or memory optimization, some embodiments may retain only the nodes that have some relationship with another customer (linked to another entity, for example). If none of the neighbors of a node connect to a different identity, then those neighbors need not be maintained in the graph database. The neighbors may be kept in a relational database until the system detects a link.Example of Ring Used to Identify Returns Anomalous Activity and Account Takeover

[0108] The example described herein corresponds to situations that would otherwise stay undetected with traditional methods due to deeper links. Some embodiments may use a list of edges, for example:[((‘4859xxxxxxxx6161’, 916790149), (916790149, ‘37 xxxxx 08096’)), ((916790149, ‘37 xxxxx 08096’), (‘37 xxxxx 08096’, 298628863)), ((‘37 xxxxx 08096’, 298628863), (298628863, ‘71.58.49.211’)), ((298628863, ‘71.58.49.211’), (‘71.58.49.211’, 112026014)), ((‘71.58.49.211’, 112026014), (112026014, ‘76.131.131.178’)), ((112026014, ‘76.131.131.178’), (‘76.131.131.178’, 1695749)), ((‘76.131.131.178’, 1695749), (1695749, ‘71.168.180.101’))]

[0109] Each edge may include 2 nodes. For an edge, one node may correspond to payment or Login IP, and another node may correspond to customer, for example. Node can be one of customer ID, address, payment, gift card recipient email, login IP, and device identifier. Node attributes can include, for example, node type, anomaly account flag, anomaly account marked timestamp, and / or anomalous activity flag. Edge attributes may include, for124489-5054-WO 30example: timestamp, anomaly label timestamp, return created timestamp, anomaly label, and / or edge link information (relationship detail value).

[0110] For the example list of edges shown above, customer identifier (CID) -916790149 - Returns anomalous activity customer (Empty box) may be linked to CID -298628863, which also may be marked for Returns Abuse / Fraud trend (via Shipping Address). The CID - 916790149 may also be linked to CID - 112026014, which may have an ATO order marked 1489682067 (via the same login IP as 298628863 customer login IP). The CID -916790149 may also be linked to CID - 1695749, which has an ATO order marked 1493652134 (via the same login IP as 112026014 customer and ordered the same SKU as the CID - 112026014 order).Example Architecture for Real-Time Scoring Model with Batch Processing Integration

[0111] The typical latency range for the real-time scoring component at the return initiation and scan checkpoints may include, for example: (i) return initiation checkpoint: The real-time scoring component typically processes return requests within 50-150 milliseconds. This may help ensure minimal disruption to the user experience while performing comprehensive risk assessments; (ii) return package scan checkpoint: The latency here ranges from 100-250 milliseconds. This checkpoint may include additional data points, such as the pickup location and / or account history. This checkpoint integrates with external scan data.

[0112] Figure 8 is a schematic diagram of an example architecture 800 for integration of real-time scoring model with batch processing, according to some embodiments. A real-time processing component 802 may process order events 804 using services 806, which may use synchronous and / or asynchronous mechanisms. Decisions and / or feature outputs from the services 806 may be input to a database 808. Data from the databases 808 may be replicated in data warehouses 814 in a batch processing component 810. A return real-time processing component 828 may process return events 830, which may include refund requests, and / or return initiations from one or more customers 832, using services 834, which may include synchronous and / or asynchronous mechanisms. Model decisions and / or rules from the services 834 may be written to (and / or read from) a database 836. Data from the database 836 may also be replicated into the data warehouses 814.124489-5054-WO 31

[0113] Turning next to the batch processing component 810, a machine learning model 826 in the batch processing component 810 may evaluate customer entities, identities, and / or identifiers in the data warehouses 814 and / or write back output into the data warehouses 814. Real-time services 838, which may include system internal and / or third-party services, may write account activity data 818, such as login, password changes, customer interaction data 820, such as interaction data obtained via applications and / or websites, customer return activity 822, and / or customer return shipment data 824, which may in turn be processed by extract, transform and / or load (ETL) jobs 816. The ETL jobs 816 may evaluate customers and / or store feature results into the data warehouses 814.

[0114] In some embodiments, the real-time processing component 802 may be triggered by event changes, such as updates in an order’s status. The component 802 may evaluate customer profiles, interaction data, and / or other dynamic features to calculate decisions in real time. These decisions and / or computed feature outputs may be stored in a centralized data warehouse (e.g., the data warehouses 814). For the batch processing component 810, the system 100 may aggregate critical data points, such as login activity, chat history, device identifiers, and / or other behavioral data, from multiple internal sources. Periodic data processing jobs may be scheduled to process historical data, compute additional features, and / or store results in the data warehouses 814 for comprehensive analysis. Decision models in the batch processing component 810 may utilize these precomputed features to evaluate customer behavior and / or write outputs back to the data warehouses 814. These decisions may then be replicated to the transactional database 836 for integration with the realtime layer.

[0115] In some embodiments, the integration points may include a shared data layer. For example, real-time scoring results are replicated to the data warehouses 814 in near real time, ensuring that batch processes may access the latest updates. The incremental ETL jobs 816 may be batch processing jobs that periodically incorporate real-time updates, enabling synchronized feature calculations and / or model refinement.Example Techniques for Low Latency Scoring While Maintaining Accuracy124489-5054-WO 32

[0116] Some embodiments may include optimized infrastructure. For example, the system 100 may include event-driven compute mechanisms and / or periodic data processing jobs, which may enable scalability and / or low-latency inference while maintaining high accuracy. Some embodiments may include real-time model optimization. For example, models may be designed for rapid inference using pre-computed features and / or lightweight architectures to balance speed and / or precision. Event pipelines may leverage asynchronous and / or synchronous mechanisms to optimize the processing of real-time transactions. For data synchronization, some embodiments may include a continuous data synchronization mechanism between the transactional database and / or the data warehouse, for ensuring that real-time results are immediately available for batch processing, eliminating the need for redundant full data transfers. Some embodiments may include incremental updates in batch processing. For example, the batch layer (sometimes referred to as the batch processing component) may process only new or changed data since the last update, reducing processing time and / or computational resources while ensuring accurate feature calculations.Example Optimizations to Reduce Processing Overhead While Maintaining System Reliability

[0117] Some embodiments of the system 100 may include event-driven real-time processing, which may include the use of lightweight, event-driven compute mechanisms that minimize system overhead compared to conventional server-based architectures. For caching and / or pre-computed features, some embodiments may cache frequently used features to reduce redundant computations and / or improve the speed of real-time decisions. Some embodiments may include streamlined data processing pipelines. For example, batch processing may include optimized data pipelines that aggregate and / or process data from multiple sources, reducing resource consumption and / or improving system efficiency. Some embodiments may include redundant data elimination. For example, by integrating real-time outputs into the shared data layer and / or leveraging incremental processing techniques, the system 100 may eliminate unnecessary replication and / or duplication of data across layers.Example Advantages over Traditional Separate Real-Time / Batch Processing Systems124489-5054-WO 33

[0118] For centralized data management, the shared data layer server 122 may ensure that all components operate on consistent, up-to-date information, addressing discrepancies inherent in separate systems. For improved decision accuracy, real-time updates may be incorporated into the batch processing 126. Batch evaluations may validate and / or enhance real-time decisions, creating a feedback loop for improved reliability. For scalability, the system 100 may include event-driven and / or auto scaling techniques, which dynamically adjust to workload demands, avoiding the inefficiencies and / or fixed capacity of traditional systems. For enhanced flexibility, the integration of real-time and / or batch processing enables the system 100 to make immediate decisions while periodically refining features and / or models, enhancing adaptability. For cost and / or resource efficiency, shared infrastructure, elimination of redundant data replication, and / or optimized processing pipelines may significantly reduce operational costs and / or computational overhead compared to traditional siloed architectures.Example Methods for Detecting and / or Preventing Anomalous Activities

[0119] Figure 9 is a flowchart of an example method 900 for detecting and / or preventing anomalous activity in e-commerce transactions, according to some embodiments. The method may be performed by one or more modules and / or components of the system 100. The system 100 may include a computing architecture, such as the one shown in Figure 1. The computing architecture may include one or more servers (e.g., the servers 106, 118), a graph database (e.g., the graph database 120), and / or a shared data layer (e.g., the shared data layer server 122). One or more processors may be configured to perform real-time scoring (e.g., perform operations of the application servers 106 and / or the modules therein). The graph database may be configured to store graph traversal paths. The shared data layer may be configured to maintain consistency between real-time and batch processing.

[0120] The data storage module 108 may store and / or maintain (902): (i) a graph database schema defining node types including return origin identification, login network identification, physical address of origin, authorization token, and / or device identification, for return transactions, (ii) relationship attributes including timestamps and / or anomaly labels for the return transactions, and / or (iii) a plurality of feature vectors related to the return transactions. In some embodiments, the data storage module 108 may store features including124489-5054-WO 34at least one of: a total ratio of processed refunds, replacements and / or concessions; an expected count of refunds, replacements and / or concessions to be received; a maximum ratio of category refunds, replacements and / or concessions value over 365 days; an international chat IP address indicator, and / or a count of chat IP addresses linked to a network identifier within 7 days.

[0121] The identity linking module 110 may identify (904) correlated anomalous activity for obtaining graph-based linking features for the plurality of feature vectors by performing: (i) breadth-first search, based on the graph database schema, with predetermined maximum depth and / or configurable neighbor expansion, (ii) temporal filtering of network address-based relationships within predetermined time windows, and / or (iii) community detection algorithms. In some embodiments, the identity linking module may include a graph processing engine, which may be configured to perform breadth-first search with a depth limit of a predetermined number of hops. In some embodiments, the identity linking module 110 may include a neighbor limiting component, which may restrict expansion of the search to a predetermined number of neighbors for high-volume nodes including shared IP addresses and / or shipping addresses.

[0122] In some embodiments, the identity linking module 110 may include a temporal filtering component, which may only traverse IP node relationships where the timestamp difference between incoming and / or outgoing relationships is within a predetermined time. In some embodiments, the identity linking module 110 may include a graph structure that connects nodes representing order IDs, login IPs, shipping addresses, payment tokens, login devices, and / or email recipients through relationship attributes including payment, login, shipping, and / or gift card information. In some embodiments, the identity linking module 110 may identify correlated anomalous activity through breadth-first search limited to 3 hops, temporal edge filtering within 60-day windows, weighted relationship scoring based on attribute overlap, and / or community detection. In some embodiments, the correlated anomalous activity identification may include cycle detection algorithms, pruning of high-centrality nodes based on temporal patterns, community detection with configurable density thresholds, and / or graph updates with atomic edge modification. In some embodiments, the identity linking module 110 may connect customer accounts by analyzing shipping address, payment information, login IP, login device, and / or email attributes in the graph-based representation.124489-5054-WO 35

[0123] Nodes of graphs used by the identity linking module 110 (which may be stored in the graph database 120) may be associated with one or more of the following values: order ID, login IP, shipping address, payment token, login device, gift card email recipient, node type (indicates type of node), and / or order, gc email, login ip, login device, address, and payment. Edges may connect pairs of nodes, indicating relationship attributes, for example. The relationship attributes may include, for example, information for paid with, login ip, login device, ship to, gc email, anomaly label - whether relationship was known to be involved in an anomaly, and / or timestamp when the relationship was established. The relationships may include properties that describe whether an anomaly has occurred. For example, ship to and paid with relationships may have anomaly label property. In addition, the customer node can also have anomaly label property to indicate if it was an anomalous account.

[0124] The natural language processing module 112 may generate (906) communication-based features for the plurality of feature vectors using (i) Term Frequency-Inverse Document Frequency (TF-IDF) analysis, (ii) sequence mapping, and / or (iii) probability scoring, based on unstructured communications. In some embodiments, the NLP module 112 may perform (i) text vectorization using TF-IDF computation, (ii) sequential pattern mining with configurable thresholds, (iii) anomalous activity indicator extraction using domain vocabularies, and / or (iv) real-time probability scoring with feature normalization.

[0125] The real-time scoring module 114 may generate (908) anomalous activity risk scores by processing return transactions at initiation and / or real-time scans at logistics scan points, using a machine learning pipeline, based on the graph-based linking features and / or the communication-based features of the plurality of feature vectors. In some embodiments, the real-time scoring module 114 may include a machine learning pipeline, which may combine scores from a batch-scored anomalous activity model and / or a real-time anomaly detection model. In some embodiments, the batch-scored anomalous activity model may maintain a 90:10 ratio of non-anomaly labeled data to anomaly labeled data through under-sampling during training. In some embodiments, the real-time scoring module 114 may further include a feature normalization module, which may maintain consistent scoring across the models.124489-5054-WO 36

[0126] In some embodiments, the real-time scoring module 114 may further transmit a generated anomalous activity risk score for a return transaction to a computing device for review, and / or receive a decision from the computing device on whether to process or hold the return transaction. In some embodiments, the real-time scoring module 114 may evaluate distances between shipping addresses and / or chat IP locations and / or monitor international IP usage patterns. In some embodiments, the real-time scoring module 114 may analyze return patterns for specific product categories and / or value thresholds using data from the data storage module.

[0127] The decision module 116 may compare the generated anomalous activity risk scores / results to a predetermined threshold or criteria, and / or determine whether to process the return transactions. In some embodiments, the decision module 116 may apply predefined rules in combination with the generated anomalous activity risk scores and / or detected correlated anomalous activities, and / or determine the appropriate level of friction to apply to each return transaction. Some embodiments may apply one or more levels of friction within the returns process. A first level of friction may include requiring the package to be returned to fulfillment center(s). A second level of friction may include holding the return payment until the package is received by the fulfillment center(s) and contents of the package are validated. Validation may include, for example, determining if the product was the intended product to receive back, if the quantity is correct, if an item is a counterfeit item, and / or if the item is heavily used.

[0128] In some embodiments, the shared data layer server 122 (sometimes referred to as the shared data layer) may include a centralized profile database that serves as a source of truth by storing real-time scoring updates. In some embodiments, the shared data layer server 122 may further include a versioning system that manages updates to the database by checking for newer real-time updates within a predetermined time window before applying any batch changes. In some embodiments, the shared data layer server 122 may further include a batch processing component, which may process incremental changes through the versioning system to update the centralized database without overwriting decisions.

[0129] In some embodiments, the system 100 may provide monitoring of latency metrics across processing stages, measurement of graph traversal performance, feature importance analysis (e.g., using SHAP), and / or accuracy metrics with false positive / negative124489-5054-WO 37tracking. In some embodiments, the system 100 may provide optimization through caching with invalidation, feature computation pipelines, batched database operations, and / or query optimization for graph traversals.

[0130] In some embodiments, the system 100 further includes the refund hold mechanism 130, which activates (e.g., based on one or more results of the decision module 116) when a returned product reaches a logistics provider for a return, applies additional risk assessment based on the pickup location, account history, and / or return details, and / or determines, based on the additional risk assessment, whether to hold the refund for further review. In some embodiments, the refund hold mechanism 130 may be activated and / or may determine whether to hold the refund for further review based on the one or more results of the decision module 116.

[0131] Figure 10 is a flowchart of another example method 1000 for detecting and / or preventing anomalous activity in e-commerce transactions, according to some embodiments. The method may be performed by one or more modules and / or components of the system 100. The system 100 may include a computing architecture, such as the one shown in Figure 1. The computing architecture may include one or more servers (e.g., the servers 106, 118), a graph database (e.g., the graph database server 120), and / or a shared data layer (e.g., the shared data layer server 122). One or more processors may be configured to perform real-time scoring (e.g., perform operations of the application servers 106 and / or the modules therein). The graph database may be configured to store graph traversal paths. The shared data layer may be configured to maintain consistency between real-time and batch processing.

[0132] The data storage module 108 may store (1002) historical transaction data, customer behavior data, and / or identity information. In some embodiments, the data storage module 108 may store a plurality of features related to customer transactions, device fingerprinting, geographic data, and / or interaction patterns with customer service.

[0133] The identity linking module 110 may analyze (1004) the stored data to identify relationships between customer accounts and / or entities, generate (1006) a graph-based representation of the identified relationships, and / or detect (1008) potential correlated anomalous activities based on the graph-based representation. In some embodiments, the identity linking module 110 may further use attributes including shipping address, payment124489-5054-WO 38information, login IP, login device, and / or email to connect customer accounts, use graph algorithms, including breadth-first search and / or cycle detection, to identify potential correlated anomalous activities, and / or represent the identified relationships in a graph-based data structure.

[0134] In some embodiments, the graph-based identity linking module 110 may store node types including order ID, login IP, shipping address, payment token, login device, and / or gift card email recipient, create edges between nodes with relationship attributes including paid with, login IP, login device, ship to, and / or email, and / or apply limited traversal depth and / or neighbor expansion to optimize performance for real-time anomaly detection. In some embodiments, the graph-based identity linking module 110 may use community detection algorithms to identify correlated anomalous activities. In some embodiments, the graph-based identity linking module 110 may limit traversals from IP nodes based on timestamp difference between incoming and / or outgoing relationships to eliminate noise from mobile IPs.

[0135] The real-time scoring module 114 may receive (1010) a return request (e.g., the return request 102) from a customer, access (1012) the stored data to evaluate the risk of anomalies associated with the return request, and / or generate (1014) an anomalous activity risk score for the return request based on the evaluation. In some embodiments, the real-time scoring module 114 may further access a batch-scored anomalous activity model and / or an order anomaly detection model, and / or combine the scores from the two models to generate the anomalous activity risk score for the return request. The batch-scored anomalous activity model and / or order anomaly detection model may be trained on historical transaction data using decision tree algorithms. In some embodiments, the real-time scoring module 114 may provide the generated anomalous activity risk score to a customer service representative for review, and / or receive the customer service representative’s decision on whether to process or hold the return request.

[0136] The decision module 116 may compare (1016) the generated anomalous activity risk score to a predetermined threshold, and / or determine (1018), based on the comparison, whether to process the return request or hold the request for further review. In some embodiments, the decision module 116 may further apply predefined rules in combination with the generated anomalous activity risk scores and / or detected correlated anomalous activities,124489-5054-WO 39and / or determine the appropriate level of friction to apply to each return request, balancing customer satisfaction and / or anomaly prevention. In some embodiments, the decision tree algorithm used in the machine learning models may be configured to handle class imbalance inherent in anomaly detection by under-sampling non-anomaly labeled data to achieve a 90: 10 ratio between non-anomalies to anomalies. In some embodiments, the decision module 116 may apply a combination of the generated risk scores, detected correlated anomalous activities, and / or predefined rules to determine the appropriate level of friction for each return request, and / or ensure that legitimate customers receive a satisfactory experience while minimizing exposure to anomalous activity.

[0137] The system 100 may use (1020) the generated anomalous activity risk scores, the detected correlated anomalous activities, and / or predefined rules to selectively apply friction to return requests, ensuring a balance between response times and / or anomaly prevention. In some embodiments, the system 100 may include the refund hold mechanism 130 to activate when the returned product reaches a logistics provider for the return process, apply additional risk assessment based on the pickup location, account history, and / or return details, and / or determine, based on the additional risk assessment, whether to hold the refund for further review.

[0138] In some embodiments, the system 100 may further include the natural language processing (NLP) module 112 to analyze customer communications, including chat, email, and / or call interactions, identify patterns and / or language indicative of potential anomalous activity, generate a probability score predicting the likelihood of anomalous activity, and / or provide the probability score as an input feature to the real-time scoring module.

[0139] In some embodiments, the data consistency between the real-time scoring module 114 and / or the batch processing components (e.g., the batch processing 126) may be maintained through the shared data layer server 122 (sometimes referred to as the shared data layer) with the centralized customer profile database 124, incremental updates to the batch jobs (via the batch processing 126) to incorporate real-time changes, and / or the versioning system 128 (sometimes referred to as the version control) to resolve any data conflicts.124489-5054-WO 40

[0140] The terminology used in the description of the invention herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used in the description of the invention and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof.

[0141] The foregoing description, for purpose of explanation, has been described with reference to specific embodiments. However, the illustrative discussions above are not intended to be exhaustive or to limit the invention to the precise forms disclosed. Many modifications and variations are possible in view of the above teachings. The embodiments were chosen and described to best explain the principles of the invention and its practical applications, to thereby enable others skilled in the art to best utilize the invention and various embodiments with various modifications as are suited to the particular use contemplated.124489-5054-WO 41

Claims

1. What is claimed is:

1. A computer-implemented system for detecting and preventing anomalous activity in e-commerce transactions, comprising:a computing architecture comprising:one or more servers configured to perform real-time scoring,a graph database configured to store graph traversal paths, and a shared data layer configured to maintain consistency between real-time and batch processing;a data storage module configured to store and maintain: (i) a graph database schema defining node types including return origin identification, login network identification, physical address of origin, authorization token, and device identification, for return transactions, (ii) relationship attributes including timestamps and anomaly labels for the return transactions, and (iii) a plurality of feature vectors related to the return transactions; an identity linking module configured to identify correlated anomalous activities by obtaining graph-based linking features for the plurality of feature vectors by performing: (i) breadth-first search, based on the graph database schema, with predetermined maximum depth and configurable neighbor expansion, (ii) temporal filtering of network address-based relationships within predetermined time windows, and (iii) community detection algorithms;a natural language processing module configured to generate a communication-based features for the plurality of feature vectors using (i) Term Frequency-Inverse Document Frequency (TF-IDF) analysis, (ii) sequence mapping, and (iii) probability scoring, based on unstructured communications;a real-time scoring module configured to generate anomalous activity risk scores by processing return transactions at initiation and real-time scans at logistics scan points, using a machine learning pipeline, based on the graph-based linking features and the communicationbased features of the plurality of feature vectors; anda decision module configured to compare the generated anomalous activity risk scores to a predetermined threshold, and determine whether to process, allow, or return the return transactions based on the comparison of the generated anomalous activity scores to the predetermined threshold.124489-5054-WO 422. The system of claim 1, wherein the real-time scoring module comprises a machine learning pipeline configured to combine scores from a batch-scored anomalous activity model and a real-time anomaly detection model.

3. The system of claim 2, wherein the batch-scored anomalous activity model is configured to maintain a 90: 10 ratio of non-anomaly labeled data to anomaly labeled data through under-sampling during training.

4. The system of claim 2, wherein the real-time scoring module further comprises a feature normalization module configured to maintain consistent scoring across the models.

5. The system of claim 2, wherein the shared data layer comprises a centralized profile database that serves as a source of truth by storing real-time scoring updates.

6. The system of claim 5, wherein the shared data layer further comprises a versioning system that manages updates to the database by checking for newer real-time updates within a predetermined time window before applying any batch changes.

7. The system of claim 5, wherein the shared data layer further comprises a batch processing component that processes incremental changes through the versioning system to update the centralized database without overwriting decisions.

8. The system of claim 5, wherein the identity linking module comprises a graph processing engine that performs breadth-first search with a depth limit of a predetermined number of hops.

9. The system of claim 8, wherein the identity linking module comprises a neighbor limiting component that restricts expansion of the search to a predetermined number of neighbors for high-volume nodes including shared IP addresses and shipping addresses.

10. The system of claim 9, wherein the identity linking module comprises a temporal filtering component that only traverses IP node relationships where the timestamp difference between incoming and outgoing relationships is within a predetermined time period.124489-5054-WO 4311. The system of claim 10, wherein the identity linking module comprises a graph structure that connects nodes representing order IDs, login IPs, shipping addresses, payment tokens, login devices, and email recipients through relationship attributes including payment, login, shipping, and gift card information.

12. The system of claim 11, wherein the identity linking module is configured to identify correlated anomalous activities through breadth-first search limited to 3 hops, temporal edge filtering within 60-day windows, weighted relationship scoring based on attribute overlap, and community detection (e.g., using Louvain algorithm variants).

13. The system of claim 12, wherein the NLP module performs (i) text vectorization using TF-IDF computation, (ii) sequential pattern mining with configurable thresholds, (iii) anomalous activity indicator extraction using domain vocabularies, and (iv) real-time probability scoring with feature normalization.

14. The system of claim 13, wherein the system provides monitoring of latency metrics across processing stages, measurement of graph traversal performance, feature importance analysis, and accuracy metrics with false positive / negative tracking.

15. The system of claim 14, wherein the correlated anomalous activity identification comprises cycle detection algorithms, pruning of high-centrality nodes based on temporal patterns, community detection with configurable density thresholds, and graph updates with atomic edge modification.

16. The system of claim 15, wherein the system provides optimization through caching with invalidation, feature computation pipelines, batched database operations, and query optimization for graph traversals.

17. The system of claim 16, wherein the decision module is configured to:apply predefined rules in combination with the generated anomalous activity risk scores and detected correlated anomalous activities, anddetermine the appropriate level of friction to apply to each return transaction.124489-5054-WO 4418. The system of claim 17, wherein the real-time scoring module is further configured to:transmit a generated anomalous activity risk score for a return transaction to a computing device for review, andreceive a decision from the computing device on whether to process or hold the return transaction.

19. The system of claim 18, further comprising a refund hold mechanism configured to:activate when a returned product reaches a logistics provider for a return,apply additional risk assessment based on the pickup location, account history, and return details, anddetermine, based on the additional risk assessment, whether to hold the refund for further review.

20. The system of claim 19, wherein the data storage module is configured to store features including at least one of: a total ratio of processed refunds, replacements and concessions, an expected count of refunds, replacements and concessions to be received, a maximum ratio of category refunds, replacements and concessions value over 365 days, and an international chat IP address indicator, and a count of chat IP addresses linked to a network identifier within 7 days.

21. The system of claim 20, wherein the identity linking module is configured to connect customer accounts by analyzing shipping address, payment information, login IP, login device, and email attributes in the graph-based representation.

22. The system of claim 21, wherein the real-time scoring module is configured to evaluate distances between shipping addresses and chat IP locations and monitors international IP usage patterns.

23. The system of claim 22, wherein the real-time scoring module is configured to analyze return patterns for specific product categories and value thresholds using data from the data storage module.124489-5054-WO 4524. A computer-implemented system for detecting and preventing anomalous activity in e-commerce transactions, the system comprising:a data storage module configured to store historical transaction data, customer behavior data, and identity information;an identity linking module configured to:analyze the stored data to identify relationships between customer accounts and entities,generate a graph-based representation of the identified relationships, and detect potential correlated anomalous activities based on the graph-based representation,a real-time scoring module configured to:receive a return request from a customer,access the stored data to evaluate the risk of anomalies associated with the return request, andgenerate an anomalous activity risk score for the return request based on the evaluation;a decision module configured to:compare the generated anomalous activity risk score to a predetermined threshold, anddetermine, based on the comparison, whether to process the return request or hold the request for further review.

25. The computer-implemented system of claim 24, wherein the real-time scoring module is further configured to:access a batch-scored anomalous activity model and an order anomaly detection model, andcombine the scores from the two models to generate the anomalous activity risk score for the return request,wherein the batch-scored anomalous activity model and order anomaly detection model are trained on historical transaction data using decision tree algorithms.124489-5054-WO 4626. The computer-implemented system of claim 25, wherein the real-time scoring module is further configured to:provide the generated anomalous activity risk score to a customer service representative for review, andreceive the customer service representative’s decision on whether to process or hold the return request.

27. The computer-implemented system of claim 25, wherein the refund hold mechanism is configured to:activate when the returned product reaches a logistics provider for the return process, apply additional risk assessment based on the pickup location, account history, and return details, anddetermine, based on the additional risk assessment, whether to hold the refund for further review.

28. The computer-implemented system of claim 27, wherein the identity linking module is further configured to:use attributes including shipping address, payment information, login IP, login device, and email to connect customer accounts,use graph algorithms, including breadth-first search and cycle detection, to identify potential correlated anomalous activities, andrepresent the identified relationships in a graph-based data structure.

29. The computer-implemented system of claim 28, further comprising a natural language processing (NLP) module configured to:analyze customer communications, including chat, email, and call interactions, identify patterns and language indicative of potential anomalous activity, generate a probability score predicting the likelihood of anomalous activity, and provide the probability score as an input feature to the real-time scoring module.

30. The computer-implemented system of claim 29, wherein the decision module is further configured to:124489-5054-WO 47apply predefined rules in combination with the generated anomalous activity risk scores and detected correlated anomalous activities, anddetermine the appropriate level of friction to apply to each return request, balancing customer satisfaction and anomaly prevention.

31. The computer-implemented system of claim 30, wherein the data storage module is configured to store a plurality of features related to customer transactions, device fingerprinting, geographic data, and interaction patterns with customer service.

32. The computer-implemented system of claim 31, wherein the data consistency between the real-time scoring module and batch processing components is maintained through a shared data layer with a centralized customer profile database, incremental updates to the batch jobs to incorporate real-time changes, and a versioning system to resolve any data conflicts.

33. The computer-implemented system of claim 32, wherein the decision tree algorithm used in the machine learning models is configured to handle class imbalance inherent in anomaly detection by under-sampling non-anomaly labeled data to achieve a 90: 10 ratio between non-anomalies to anomalies.

34. The computer-implemented system of claim 33, wherein the graph-based identity linking module is configured to:store node types including order ID, login IP, shipping address, payment token, login device, and gift card email recipient;create edges between nodes with relationship attributes including paid with, login IP, login device, ship to, and email; andapply limited traversal depth and neighbor expansion to optimize performance for real-time anomaly detection.

35. The computer-implemented system of claim 34, wherein the graph-based identity linking module is configured to use community detection algorithms to identify correlated anomalous activities.124489-5054-WO 4836. The computer-implemented system of claim 35, wherein the graph-based identity linking module is configured to limit traversals from IP nodes based on timestamp difference between incoming and outgoing relationships to eliminate noise from mobile IPs.

37. The computer-implemented system of claim 36, wherein the decision module is configured to:apply a combination of the generated risk scores, detected correlated anomalous activities, and predefined rules to determine the appropriate level of friction for each return request, andensure that legitimate customers receive a satisfactory experience while minimizing exposure to anomalous activity.124489-5054-WO 49