Method, system and computer program product for detection of recurring, duplicate, or similar unwanted user profiles in online communities

A multi-modal system using metadata, content, activity, and network analysis with machine learning identifies and flags duplicate or recurring user profiles, improving detection and management in online communities.

WO2026159113A1PCT designated stage Publication Date: 2026-07-30AIBA AS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
AIBA AS
Filing Date
2026-01-21
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Existing methods for detecting duplicate or recurring user profiles in online communities are inadequate, particularly due to their reliance on exact matches and are easily circumvented, leading to inefficiencies and imprecise sanctions, and manual moderation is labor-intensive and unsustainable.

Method used

A multi-modal approach combining metadata analysis, content analysis, activity pattern analysis, network analysis, and machine learning techniques to identify and flag profiles with a probability score, enabling automated or manual sanctions.

Benefits of technology

Enhances the detection and management of unwanted user profiles by providing accurate, automated, and scalable solutions, reducing false positives and improving community security and integrity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2026051386_30072026_PF_FP_ABST
    Figure EP2026051386_30072026_PF_FP_ABST
Patent Text Reader

Abstract

A system and method for detecting and managing recurring, duplicate, or similar unwanted user profiles in online communities which combines multiple analytical modalities, including account metadata analysis, content analysis, activity pattern analysis, and network analysis, to identify problematic user profiles. Behavioral clustering analysis is employed to group profiles with shared behavioral characteristics, enhancing detection precision. The system integrates these analytical results into a configurable rules engine that assigns probability scores to flagged profiles, enabling automated or manual sanctions. A user-friendly interface visualizes flagged profiles, clusters, and associated risk scores for moderator review. By leveraging advanced machine learning techniques, including graph neural networks and clustering algorithms, a robust, scalable solution to enhance community security and integrity is provided.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Method, system and computer program product for detection of recurring, duplicate, or similar unwanted user profiles in online communities Technical field

[0002] The present patent application relates to systems and methods aimed at identifying and managing recurring, duplicate, or similar unwanted user profiles in online communities. By leveraging advanced analytics and a multi-modal approach, this invention facilitates effective community moderation and enhances information security by addressing the challenges posed by malicious actors creating multiple user accounts.

[0003] Background

[0004] With the increasing use of social media and online gaming among children, there has been a corresponding rise in their exposure to unwanted behavior and threats from malicious users. Incidents of grooming, toxic behavior, and the use of fake profiles targeting children have become widespread, representing a significant threat to their safety. Traditionally, online communities have attempted to mitigate these risks by employing human moderators to manually review user-generated content and sanction users based on their findings. Moderators inspect various pieces of content to determine the appropriate actions, if any, to sanction users. In recent years, online communities have increasingly faced challenges related to users creating multiple accounts to bypass detection mechanisms or avoid sanctions. These duplicate or recurring accounts are often used for malicious purposes, such as spamming, trolling, or coordinated manipulations, thereby compromising the integrity of online platforms. Traditional approaches, such as IP address detection, device fingerprinting, and password hash matching, have proven inadequate in effectively tackling this issue. These methods require exact matches and can often be circumvented through readily available tools, including IP masking or the use of alternative devices.

[0005] Known solutions often involve the use of automated systems that sanction users based on the detection of specific words or phrases. However, these methods are limited by their reliance on exact matches, making them less effective in identifying nuanced or evolving threats. Furthermore, manual moderation requires a substantial amount of labor, making it unsustainable and inefficient in the long term. The vast volume of data generated by users exceeds human capacity to review and evaluate effectively, leading to imprecise sanctions and the potential for falsely accusing innocent users.

[0006] To solve the above-mentioned, an additional problem arises when recurring user profiles or duplicate user profiles where the same actor creates multiple accounts toalternate between them to avoid being detected or to be able to continue certain behaviors in online communities from other user profiles if one or more of the existing user profiles are blocked or sanctioned because of breach of community guidelines and rules.

[0007] Know solutions to this is to do IP-detection or device-detection to identify user accounts from the identical IP-address or identical device signature, for instance MAC-address. The problem with this is that there has to be an exact match and that the information can be difficult to obtain for all devices in online communities. Instead of depending on exact matches there is a need for analyzing user profiles based on several factors, combine these and group them together based on how many of these factors either match or are close to each other. The similar profiles are grouped based on a certainty score / confidence score weighted from match or closeness in the different factors.

[0008] Summary

[0009] In view of the above, an object according to embodiments of the present application is to overcome or at least mitigate drawbacks of prior art.

[0010] In a first aspect, a computer implemented method for detecting recurring, duplicate, or similar unwanted user profiles in online communities is provided, wherein the method comprising receiving metadata associated with user accounts, the metadata including at least one of IP addresses, email addresses, usernames, or device identifiers, analyzing the metadate to identify matches or close similarities, analyzing user content to identify patterns or phrases indicative of malicious activities, monitoring user activity patterns, including login frequencies, interaction timings, and response intervals, to calculate a behavioral risk score, performing network analysis to construct a graph of user interactions and applying machine learning techniques to identify anomalous clusters or central nodes, aggregating results from the metadata analysis, the content analysis, the activity pattern analysis, and network analysis into a rules engine to generate a probability score for each profile being duplicate, recurring, or similar and presenting flagged profiles, along with their probability scores and analysis details, for manual or automated sanctioning. A system corresponding to the above described method is also provided.

[0011] Brief description of the drawings

[0012] Figure 1 illustrates the integration of multiple analytical modalities into the system’s rules engine, showcasing how data from different analyses contribute toidentifying unwanted profiles according to an example embodiment of the present application.

[0013] Figure 2 depicts a behavioral clustering analysis according to an example embodiment of the present application.

[0014] Figure 3 illustrates a section of the clustering analysis illustrated in Figure 2.

[0015] Figure 4 illustrates a section of the clustering analysis illustrated in Figure 2.

[0016] Detailed description

[0017] Embodiments of the present application address and mitigate the disadvantages associated with prior art solutions, particularly the issue of recurring, duplicate or similar user profiles. These profiles are often created by the same actor using multiple accounts to circumvent detection and continue certain behaviors in online communities even after some profiles are blocked or sanctioned for breaching community guidelines and rules. Detection of similar user profiles may also be utilizing the same embodiments of the present application. These similar profiles may then also be banned because of similar behavior.

[0018] A first aspect according to embodiments of the present application is account metadata analysis. This analytical modality focuses on metadata associated with user accounts. Attributes such as IP addresses, email addresses, usernames, and device identifiers are compared against a database of previously flagged accounts. Matches or close similarities across these attributes generate a list of potentially problematic accounts. For instance, variations in usernames or email addresses that are designed to evade detection can still be identified through pattern recognition techniques. The system provides moderators with a detailed breakdown of matched attributes, aiding in the efficient identification of unwanted profiles. Figure 1 demonstrates how these metadata attributes contribute to the rules engine for detecting recurring or duplicate profiles.

[0019] A second aspect according to embodiments of the present application is content analysis. This content modality involves the evaluation of content produced by user profiles. By comparing this content with a repository of historical data from flagged accounts or a predefined watchlist, the system identifies problematic patterns or specific phrases linked to malicious activities. This analysis enables the detection of accounts that attempt to evade traditional content moderation methods. The flagged content is highlighted for review, along with the associated user profiles. Refer to Figure 1 for the integration of this content analysis process into the system’s workflow.A third aspect according to embodiments of the present application is activity pattern analysis. This activity modality captures and analyzes user behavior over time. This includes monitoring login patterns, interaction frequencies, and response times. The system calculates a risk score based on deviations from typical user behaviors. For example, accounts that exhibit burst activities or prolonged periods of anomalous interactions are flagged as high-risk. This modality provides valuable insights into user behavior, enhancing the detection of problematic accounts. Figure 1 provides a visual representation of how activity pattern analysis feeds into the rules engine.

[0020] A fourth aspect according to embodiments of the present application is network analysis. This fourth modality utilizes advanced network analysis techniques to evaluate relationships and interactions between user profiles. By constructing a network graph, the system applies Graph Machine Learning, GML, and Graph Neural Networks, GNNs, to detect anomalies in reciprocity, clustering, and centrality measures. GML, as used here, refers to the application of machine learning techniques specifically designed for graph-structured data, such as identifying clusters or influential nodes. GNNs extend this by employing neural networks to process and learn from the complex structures and relationships inherent in graph data. These technologies are crucial in identifying coordinated activities, echo chambers, or bridge accounts connecting disparate groups. Figure 1 visually integrates this network analysis into the overall detection system.

[0021] According to certain embodies of the present application, the results from the aforementioned modalities are fed into a rules engine, which aggregates the data and calculates a probability score for each profile. If there is not relevant data present from the client’s integration so one or more of the analytics processes cannot be conducted, the rules engine will use the analysis that is present. If data is present that makes it possible to conduct matches on account meta data or historic content of previously sanctioned content, the rules engine can predict the probability of recurring or duplicate profiles. If analysis of activity patterns or network patterns are available, the rules engine can also predict similar profiles in addition to duplicate or recurring profiles. The profiles identified by one or more of the analytics processes will be displayed in a dedicated view in the application with a probability score of the profile either being a duplicate, recurring or similar profile with the main reasons for why they are considered one of the three.

[0022] Based on the probability score it is also possible to perform automatic sanctions if the customer chooses to set rules for which probability score is enough to confidently sanction the profiles. If they choose not to set up automatic sanctions,they can manually mark the profiles they want to sanction and ban those profiles from their platform.

[0023] This score reflects the likelihood of a profile being a duplicate, recurring, or similar account. The rules engine is highly configurable, allowing administrators to define thresholds for automatic sanctions or manual review. The system’s flexibility ensures that it can adapt to varying levels of data availability and platform -specific requirements. Figure 1 encapsulates the interaction between all modalities and the rules engine.

[0024] Embodiments of the present application also include a behavioral clustering analysis, as depicted in Figures 2, 3 and 4. This analytical modality focuses on identifying patterns of user behavior by grouping user profiles based on shared behavioral characteristics. Unlike activity pattern analysis, which evaluates individual user actions over time, behavioral clustering leverages collective patterns observed across multiple accounts to identify groups of profiles exhibiting similar behaviors.

[0025] The system employs unsupervised machine learning algorithms, such as k-means clustering and hierarchical clustering, to analyze behavioral data. These algorithms evaluate multiple dimensions of user activity, including frequency of logins, content posting intervals, interaction styles, and engagement levels with other profiles. By clustering profiles based on these dimensions, the system identifies anomalous groups that deviate from typical behavioral norms.

[0026] For instance, profiles that display synchronized login activities, repetitive posting patterns, or coordinated interactions across a network are flagged as potentially problematic clusters. These clusters are assigned confidence scores based on the consistency and severity of their anomalous behaviors. Profiles within high-confidence clusters are prioritized for review by moderators or subjected to automated sanctions based on predefined rules.

[0027] Figures 2, 3 and 4 illustrates the clustering process, showcasing how the system visualizes clusters of similar user profiles in a two-dimensional representation. Each point represents a user profile, with proximity indicating behavioral similarity. The visualization aids moderators in understanding the scope and nature of identified clusters, facilitating informed decision-making.

[0028] To enhance accuracy, the system integrates data from other modalities, such as account metadata and activity patterns, into the clustering analysis. This integration ensures that the clustering process accounts for a broader range of attributes, reducing the likelihood of false positives and improving detection precision.By incorporating behavioral clustering analysis, the embodiments of the present application provide an additional layer of detection capability. This modality complements the existing analytical approaches, offering a more comprehensive solution for identifying and managing recurring, duplicate, or similar unwanted user profiles in online communities.

[0029] The system provides a dedicated interface for moderators, enabling them to review flagged profiles with ease. The interface displays probability scores, matched attributes, and relevant analysis details, allowing for informed decision-making. Moderators can choose to sanction flagged profiles, request further investigation, or dismiss alerts as needed. Automated sanctions can also be implemented based on predefined rules, streamlining the moderation process.

[0030] By combining these modalities and functionalities, the embodiments according to the present application offer a comprehensive solution for detecting and managing unwanted user profiles, thereby enhancing the security and integrity of online communities.

Claims

7CLAIMS1. A computer implemented method for detecting recurring, duplicate, or similar unwanted user profiles in online communities, the method comprising:receiving metadata associated with user accounts, the metadata including at least one of IP addresses, email addresses, usernames, or device identifiers; analyzing the metadate to identify matches or close similarities; analyzing user content to identify patterns or phrases indicative of malicious activities;monitoring user activity patterns, including login frequencies, interaction timings, and response intervals, to calculate a behavioral risk score; performing network analysis to construct a graph of user interactions and applying machine learning techniques to identify anomalous clusters or central nodes;aggregating results from the metadata analysis, the content analysis, the activity pattern analysis, and network analysis into a rules engine to generate a probability score for each profile being duplicate, recurring, or similar; and presenting flagged profiles, along with their probability scores and analysis details, for manual or automated sanctioning.

2. The method of claim 1, wherein the metadata analysis includes identifying variations in usernames or email addresses using pattern recognition algorithms.

3. The method of claim 1, wherein the content analysis involves comparing usergenerated content with a repository of historical flagged data.

4. The method of claim 1, wherein the network analysis employs graph neural networks to detect anomalous reciprocity, clustering, or centrality measures in user interactions.

5. The method of claim 1, wherein the rules engine assigns configurable thresholds for automatic sanctions based on the probability score.

6. The method of claim 1, further comprising visualizing user clusters in a two-dimensional graph for moderator review.

7. The method of claim 1, wherein the system integrates behavioral clustering analysis using unsupervised machine learning techniques to group user profiles with similar behavioral patterns.

88. A system for detecting recurring, duplicate, or similar unwanted user profiles in online communities, the system comprising:a metadata analysis module configured to evaluate user account metadata, including IP addresses, email addresses, usernames, or device identifiers, to identify matches or close similarities;a content analysis module configured to analyze user-generated content for patterns or phrases linked to malicious activities;an activity analysis module configured to monitor user behavior over time, calculating risk scores based on deviations from typical activity;a network analysis module configured to construct a graph of user interactions and analyze it using machine learning techniques to detect anomalous clusters or central nodes;a rules engine configured to aggregate analysis results, calculate a probability score for each profile, and apply configurable thresholds for sanctions; andan interface configured to display flagged profiles, probability scores, and analysis details to moderators for manual review or automated sanctions.

9. The system of claim 8, wherein the metadata analysis module includes algorithms for detecting variations in usernames or email addresses.

10. The system of claim 8, wherein the content analysis module compares flagged content with a predefined watchlist or repository of historical data.

11. The system of claim 8, wherein the network analysis module utilizes graph neural networks to evaluate clustering and centrality measures.

12. The system of claim 8, wherein the rules engine is configured to automatically impose sanctions based on predefined probability score thresholds.

13. The system of claim 8, wherein the interface provides a two-dimensional visualization of user profile clusters for enhanced moderator decision-making.

14. The system of claim 8, comprising a behavioral clustering analysis module employing unsupervised machine learning to group user profiles based on shared behavioral characteristics.