SYSTEM AND METHOD FOR MAPPING A NETWORK ENVIRONMENT USING CROSS-ACCOUNT CLUSTERING FOR THE PURPOSE OF MONITORING AND / OR DETECTING ABUSE OF NETWORKS OF UNAUTHORIZED ENTITIES - Patent application
Cross-account clustering in network environments allows for efficient detection and removal of unauthorized entity networks by mapping and graphing customer accounts, addressing the challenges of identifying and remediating harmful digital content across data channels.
Patent Information
- Application Number
- JP2023580363
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-06-30
- Filing Date
- 2022-06-30
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2042-06-30
AI Technical Summary
Identifying, tracking, and remediating harmful digital content on the Internet is challenging due to its dynamic nature and perpetrators' ability to hide identities and create pseudonyms, making it difficult to access the extent of harmful content and entities that perpetuate it across various data channels.
A system and method using cross-account clustering to map network environments, identify unauthorized entity networks, and generate a network graph to detect and remove fraudulent entities by creating edges between nodes in different customer accounts' databases, while protecting confidential information.
Facilitates coordinated and efficient tracking and remediation of harmful content by identifying networks of fraudulent entities across multiple data channels, enabling targeted actions against entire networks rather than individual sites, and preserving customer confidentiality.
Smart Images

Figure 0007763873000005 
Figure 0007763873000006 
Figure 0007763873000007
Abstract
Description
[Technical Field]
[0001] [Related Applications] This application claims the benefit of and priority to U.S. Provisional Application No. 63 / 216,878, filed June 30, 2021, which is incorporated by reference in its entirety.
[0002] The present invention relates to a system and method for mapping a network environment using cross-account clustering for the purpose of monitoring and / or detecting rogue entity networks. [Background technology]
[0003] An overwhelming amount of digital content is accessible via networked environments such as the Internet. The content spans multiple data channels and / or sources, and the amount of available content continues to grow daily. While most content is legitimate / benign, some content is harmful (e.g., fraudulent, counterfeit, infringing, malicious, etc.).
[0004] Given the fluidity with which digital content can be added or removed from one or more data channels or sources, and the ability of perpetrators of harmful content to hide their identities and / or create pseudonyms or subsidiaries, identifying, tracking, and remediating harmful content on the Internet often resembles aiming at a moving target. The dynamic nature of digital content on the Internet can make it difficult to access the extent of harmful content and / or the extent of the entities that perpetuate harmful content at any given time and / or across various data channels, thereby making it difficult to appropriately track and remediate harmful content in a coordinated, effective, and efficient manner. Summary of the Invention [Problem to be solved by the invention]
[0005] Embodiments of the present disclosure provide for the detection, monitoring, and / or removal of unauthorized entity networks in a network environment. Cross-account clustering can be used to map the network environment and identify nodes associated with one or more entity networks in the network environment. Additionally, it can be determined whether one or more entity networks are unauthorized entity networks based on a determination that one or more nodes in the one or more entity networks are a source of harmful content (e.g., fraudulent, counterfeit, infringing, malicious content, etc.). Upon detecting an unauthorized entity network, embodiments of the present disclosure can alert parties potentially affected by the one or more unauthorized entity networks and / or initiate one or more actions against the unauthorized entity networks.
[0006] Cross-account clustering within a network map / network graph can be used to create edges between nodes in the network graph by determining explicit and implicit connections or links between documents in different customer accounts' databases based on data in those customer accounts' databases. This allows for the generation of a robust network map / network graph of networks of entities in a network environment while protecting confidential and / or personal customer information from other customers. This approach can determine that seemingly unrelated and / or distinct networks of entities detected by different customer accounts are the same network of entities. And / or, it can determine that a network of entities that appears legitimate based on database records associated with one customer account is part of one or more networks of entities determined to be fraudulent based on database records of one or more other customer accounts. This approach also allows embodiments of the present disclosure to determine that a particular network of fraudulent entities is targeting a particular industry, product, and / or type of brand. This can alert a customer account to the presence of a fraudulent network of entities even if the customer account is not currently determined to be a target of the fraudulent network.
[0007] Embodiments of the present disclosure may address the challenges associated with identifying, tracking, and remediating harmful content on the Internet, where digital content may be added or removed from one or more data channels or sources, allowing perpetrators of harmful content to hide their identities and / or create pseudonyms or subsidiaries. Embodiments of the present disclosure may provide customers with easy access to the scope of harmful content and / or the scope of entities that perpetuate harmful content over any given time and / or across various data channels, enabling coordinated, effective, and efficient tracking and remediation of harmful content.
[0008] In an embodiment of the present disclosure, a system, method, and non-transitory computer-readable medium are provided for detecting and monitoring a network of unauthorized entities in a network environment. The system includes a computing system communicatively coupled to a data source in the network environment, the data source including one or more servers configured to host digital content, and one or more processors disposed within the computing system. The non-transitory computer-readable medium stores instructions for performing the method, the instructions being executed by the one or more processors. The one or more processors are configured to establish separate customer accounts and search for content hosted on one or more remote servers in the network environment for each customer account, generating separate aggregated data sets for each customer account. The one or more processors are also configured to tag each search result in the aggregated data sets as legitimate or malicious based on an analysis of each search result in the aggregated data sets, and generate a network graph by combining data from each search result in the aggregated data sets for the customer accounts. The one or more processors are also configured to generate, at the remote server, a plurality of clusters including a cross-account cluster including data from two or more customer accounts, identify one or more networks of fraudulent entities based on the plurality of clusters in the network graph, and initiate removal actions against the identified one or more networks of fraudulent entities.
[0009] Any combination and / or permutation of the embodiments is contemplated. Other objects and features will become apparent from the following detailed description considered in conjunction with the accompanying drawings. It is to be understood, however, that the drawings are designed solely as illustrations and are not intended as a definition of the limits of the present disclosure. [Brief explanation of the drawings]
[0010] In the drawings, like numerals refer to like parts throughout the various views of non-limiting and non-exhaustive embodiments. [Figure 1] FIG. 1 is a block diagram of an exemplary unauthorized content monitoring / detection engine according to an embodiment of the present disclosure. [Figure 2] FIG. 2 is a block diagram of an exemplary computing device according to an embodiment of the present disclosure. [Figure 3] FIG. 3 illustrates an exemplary network environment for facilitating the collection, analysis, analysis, and removal of abusive content on the Internet, in accordance with an embodiment of the present disclosure. [Figure 4] FIG. 4 is a flowchart illustrating an example process for creating and / or updating a graph database and generating graphs and subgraphs to detect networks of unauthorized entities in accordance with an embodiment of the present disclosure. [Figure 5] FIG. 5 is a block diagram illustrating a visualized embodiment of a graph data model defined for a graph database according to embodiments of the present disclosure. [Figure 6] FIG. 6 illustrates a simple example of a graph generated for data in a graph database based on a graph data model defined for the graph database, in accordance with an embodiment of the present disclosure. [Figure 7] FIG. 7 illustrates a graphical user interface showing a list of documents in a graph database and their corresponding keys for each customer account in accordance with an embodiment of the present disclosure. [Figure 8] FIG. 8 illustrates the graphical user interface of FIG. 7 with an area showing the total number of clusters in accordance with an embodiment of the present disclosure. [Figure 9] FIG. 9 illustrates a graphical user interface showing a cluster browser in which a rendered graph is visualized in accordance with an embodiment of the present disclosure. [Figure 10] FIG. 10 illustrates a graphical user interface showing a grid view in an embodiment of the present disclosure. [Figure 11] FIG. 11 is a flowchart illustrating an example process for detecting, monitoring, and / or removing a network of unauthorized entities in accordance with an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0011] Embodiments of the present disclosure relate to systems, methods, and non-transitory computer-readable media for detecting, monitoring, and / or removing unauthorized entity networks in a network environment. Cross-account clustering can be used to map the network environment and identify nodes associated with one or more entity networks in the network environment. Embodiments of the present disclosure can also detect whether one or more entity networks are unauthorized entity networks based on a determination that one or more nodes in the one or more entity networks are a source of harmful content (e.g., fraudulent, counterfeit, infringing, malicious content, etc.). Upon detecting an unauthorized entity network, embodiments of the present disclosure can alert parties potentially affected by the one or more unauthorized entity networks and / or initiate one or more actions against the unauthorized entity networks.
[0012] In a non-limiting example application, embodiments of the present disclosure can be implemented for brand protection in a network environment. Embodiments of the present disclosure provide the ability to collect digital content (e.g., web pages) from data sources in a network environment. In the network environment, for example, one or more collection engines search for digital content in the data sources based on searches (e.g., keyword searches, Uniform Resource Locators (URLs), etc.), and an extraction engine extracts attributes from the digital content. Upon detecting harmful content in the collected digital content, embodiments of the present disclosure can create one or more tags used to define the type of harmful content detected (e.g., fraudulent, counterfeit, infringing, malicious, etc.). As an example, harmful content includes web pages offering to sell counterfeit, infringing, fraudulent, or malicious goods.
[0013] Exemplary embodiments of the present disclosure can be utilized (e.g., for brand protection) to create separate customer accounts for each customer. This allows the searches used to aggregate digital content, the aggregation results, and the tagging of the results to be unique to each customer account and typically not shared between customer accounts, thereby maintaining customer confidentiality and privacy. The aggregation results for each customer account can be stored as separate records in a database. In the database, content and information extracted from the results can form data fields of the records. Additionally, a customer identifier can be added to each record in the database to associate each record with the corresponding customer for which the record was generated, an industry identifier can be added, and / or a product category identifier can be added. The customer identifier can be unique to a customer. Meanwhile, the industry identifier and product category identifier can be shared between customer accounts in the same industry or between customer accounts selling products in the same category. In some embodiments, a customer may belong to multiple industries and / or sell products in multiple product categories. In such embodiments, a customer account can be associated with an industry identifier corresponding to each industry associated with the customer account and / or with a product category identifier corresponding to each product category associated with the customer account. The tags generated for each result / record can also be added to each record as data fields in the database.
[0014] By utilizing tags from collected digital content and customer account database records, embodiments of the present disclosure can be configured to create a cumulative or aggregated network map or network graph by combining the digital content and tags collected from customer accounts, identifying networks of fraudulent entities, and assessing the scope and nature of harmful content across multiple data channels, industries, and brands. Embodiments of the present disclosure can utilize cross-account clustering within the network map / network graph to create edges between nodes in the network graph by determining explicit and implicit connections or links between records in different customer accounts based on data in data fields in those records. This can generate a robust network map / network graph of networks of entities in a network environment. This approach can determine that seemingly unrelated and / or distinct networks of entities detected by different customer accounts are part of the same entity network. Additionally, a network of entities that appears legitimate based on database records associated with one customer account can be determined to be part of a network of one or more entities determined to be fraudulent based on database records for one or more other customer accounts. This approach also allows embodiments of the present disclosure to determine that a particular network of fraudulent entities is targeting a particular industry, product, and / or type of brand, thereby alerting a customer account to the presence of a network of fraudulent entities even if the customer account is not currently determined to be a target of the network of fraudulent entities.
[0015] In utilizing cross-account clustering, confidential or personal information associated with one customer may be disclosed to another customer via the network map / network graph. To prevent disclosure of confidential and / or personal information between customer accounts, embodiments of the present disclosure may anonymize and / or obfuscate information in the network map / network graph. Furthermore, nodes and / or edges of the graph may be omitted or modified to protect confidential or personal customer information.
[0016] Generating a network map / network graph that identifies a network of fraudulent entities using cross-account clustering facilitates targeted and widespread remediation actions, taking down a network of fraudulent entities on a larger scale than would otherwise be possible. Thus, embodiments of the present disclosure facilitate coordinated action against an entire (or large portion of) a network of fraudulent entities, rather than having to target individual websites and / or e-commerce platforms. Furthermore, embodiments of the present disclosure may be used as evidence in legal proceedings.
[0017] FIG. 1 is a block diagram of an exemplary malicious entity network detection and / or monitoring engine 100 according to an embodiment of the present disclosure. The engine 100 allows users to establish customer accounts 102. In the customer accounts 102, data associated with each customer account may be confidential and / or personal data that is not typically shared with other customer accounts 102. In an example application, the engine 100 is implemented for brand protection in a network environment. In this case, the customer accounts 102 are associated with organizations or businesses seeking to detect, monitor, and remove harmful content associated with their products, trademarks, and / or brands from the Internet. The engine 100 may include a user interface 110, a collection engine 115, an extraction engine 120, a tagging engine 125, an analysis engine 130, a database transformation engine 140, a clustering engine 150 including an entity resolution engine 152 and a probabilistic modeling engine 154, a network graphing engine 155, and a removal engine 160.
[0018] The engine 100 collects, extracts, and analyzes digital content (collected data sets 117) for each account from heterogeneous data sources 104 associated with nodes in the network environment. Additionally, the data sources 104 and / or collected data sets 117 from the heterogeneous sources 104 can be different for each customer account 102 (e.g., each customer account 102 can have different search criteria). As an example, a first account can use the engine 100 to collect, extract, and analyze a first set of content / information (e.g., a first collected data set) from a first set of data sources 104. A second account can use the engine 100 to collect, extract, and analyze a second set of content / information (e.g., a second collected data set) from a second set of data sources 104. In this case, the first collected data set and the second collected data set can have common elements (e.g., have some of the same results) or can be mutually exclusive (e.g., have no common results). The engine 100 also receives additional data that supplements the data extracted from the collected digital content. As one example, a user of a customer account can input data (e.g., merchant information, domains, contacts, merchant tracking) to be included in the collection dataset 117. As another example, additional data (e.g., merchant information) can be received from an online marketplace and included in the collection dataset 117. The engine 100 generates a network graph of the network environment by combining data extracted from results in the collection datasets (e.g., the first collection dataset and the second collection dataset) associated with different customer accounts 102. The network graph is used to identify and detect one or more networks of fraudulent entities in the network environment, associate one or more networks of fraudulent entities with one or more customer accounts 102, and / or determine the scope and pseudonyms of entities in the network of fraudulent entities in the network environment.
[0019] The disparate data sources 104 can be associated with various data channels on the Internet or in other network environments. For example, the disparate sources 104 can include servers and / or databases hosting Internet / digital content, such as websites, social media, e-commerce marketplaces, online marketplaces, the dark web, databases for identifying Internet resource information (e.g., registrant names, registrar names, physical addresses, phone numbers, email addresses, seller names, domain name owners, IP address blocks, etc.).
[0020] In one embodiment, the collection engine 115 is configured to search for harmful online content by crawling the web and / or dark web, collecting search engines and / or APIs to search web pages, mobile application data, and / or other data in the network environment. The collection engine 115 searches for content and information from one or more disparate data sources 104 in the network environment based on item identifiers, keyword strings, or combinations thereof that are unique to each customer account.
[0021] The collection engine 115 generates or constructs one or more queries (e.g., database, API, or web-based queries) based on one or more search terms (e.g., keywords) entered by one or more users 106 of one or more customer accounts 102 via one or more graphical user interfaces 114 in the user interface 110. As an example, the collection engine 115 constructs multiple queries from a single set of search terms, where each query can be specific to a search engine and / or application programming interface (API).
[0022] The collection engine 115 may be implemented to facilitate parallel searches of various data sources 104 for similar content. Queries may be generated or constructed using one or more query languages, such as Structured Query Language (SQL), Contextual Query Language (CQL), proprietary query languages, domain-specific query languages, and / or other suitable query languages. In some embodiments, the collection engine 115 may generate or construct one or more queries using one or more programming languages or scripts, such as Java, C, C++, Perl, Ruby, etc.
[0023] The collection engine 115 may execute each query for each customer account using a search engine and / or API to return Internet content and / or other content in the network environment. By way of example, the collection engine 115 may execute to return one or more web pages from one or more Internet domains hosted by one or more web servers of one or more data sources returned in response to a query using a search term.
[0024] In an exemplary embodiment, results returned via the collection engine 115 are retrieved, downloaded to a storage device, and stored as a collection dataset 117. For example, each result (e.g., each web page, etc.) may be saved as a file or other data structure. In some examples, one or more results may be saved in the same format as they were in the original data source. For example, each web page may be saved in a native text-based markup language (e.g., HTML, XHTML, etc.). In some examples, one or more results may be saved in a format different from the format in which they were saved in the original data source. In an exemplary embodiment, the collection engine 115 may return over hundreds of millions (100 million or more) unique results over time. The frequency with which the collection engine 115 collects content and information may be specific to each account. Thus, for any given customer account, the collection engine 115 may collect content and information from the data sources 104 in the network environment hourly, daily, weekly, monthly, quarterly, yearly, etc. And / or the collection engine 115 may collect content and information on demand (e.g., in response to a request from a user of the customer account). The queries and search terms utilized by the collection engine 115 may be updated and / or modified, for example, based on analysis of results from previous collection operations and / or based on detection and mapping of networks of unauthorized entities.
[0025] The extraction engine 120 extracts content and information from each result (e.g., each web page and associated metadata, etc.) in the collection dataset 117. In an exemplary embodiment, the content and information extracted from the results may include product information (e.g., brand name, company name, logo, product description, product image, product price, GTIN, SKU, UPC, EAN, etc.), seller or user information (e.g., seller / user name, address, phone number, email address, domain name, uniform resource locator, etc.), social media profile information (e.g., information containing product information and / or seller information), website information (e.g., information such as images and / or text contained in the body of a web page and / or information contained in the source code of a web page), network information (e.g., registrant name, registrar name, domain name, IP address, internet resource owner information, uniform resource locator, uniform resource identifier (URI), etc.), etc. The content and information extracted from the results for each customer account may be used to determine whether each record corresponds to legitimate content or harmful content.
[0026] As the extraction engine 120 extracts content and information from each result for each account, it builds and / or updates a database 135 with the content and information extracted from the results. The database 135 may be a relational database. The extraction engine 120 creates a record 137 in the database 135 for each result in the collected dataset 117 and for each customer account 102. The extraction engine 120 also stores the content and information extracted from each result as data in data fields in the corresponding record 137. As an example, each unique result is stored as a record (defined as a row in the database 135). In this case, the content and information extracted from each record may be stored in data fields or columns of the record. In addition to data fields for storing data extracted from the results, the records 137 in the database 135 may include additional data fields based on the associated customer account and / or based on an analysis of the corresponding result.
[0027] Examples of data fields or data strings that may be included in record 137 include, for example, the following data fields: product name, product description, seller name, GTIN, SKU, UPC, EAN, a marketplace-specific identifier (e.g., Amazon® Standard Identification Number (ASIN)), the geographic location of the seller, the geographic location where the seller ships the product, seller reviews, the result title (e.g., the title of the webpage), the product price, the quantity of the product available for purchase, product dimensions, images, product images, logos and / or artwork, video, audio, the registered name of the domain of the webpage, the domain name server hosting the result, the name of the registrar registering the result, the IP address of the domain, the domain name, a tag indicating whether the record is associated with legitimate content or harmful content, a customer account identifier (identifying which customer account the record belongs to), an industry identifier (identifying the customer's industry), one or more product type identifiers (identifying the type of products the customer sells and / or the type of products offered by the result corresponding to the record), the source code of an HTML page, an XML file, and JavaScript.
[0028] To extract content and information from the results in the collection dataset 117, the extraction engine 120 identifies item identifiers within the results using, for example, natural language processing, machine learning, similarity measures, image matching techniques including pixel matching, and / or pattern matching techniques. The extraction engine 120 utilizes one or more entity ontologies to derive and / or identify entities (e.g., vendor names, internet resource owners, etc.) included in the results. The extraction engine 120 can utilize a variety of algorithms and / or techniques. For example, fuzzy text pattern matching algorithms, such as the Baeza-Yates-Gonnet algorithm, can be used for single strings, and fuzzy Aho-Corasick algorithm (AC) can be used for multiple strings. Additionally, supervised or unsupervised document classification algorithms can be implemented after converting text into numerical vectors using algorithms for fuzzy text pattern matching against multiple strings, such as the fuzzy Aho-Corasick algorithm, and topic models, such as Latent Dirichlet Allocation (LDA) and Hierarchical Dirichlet Process (HDP).
[0029] Also, in another embodiment, rather than having the collection engine 115 download the results and create the collection data set 117, the collection engine 115 identifies the results (e.g., web pages, etc.) and the extraction engine 120 analyzes the content and information of the data sources associated with the results. The extraction engine 120 uses the content and information to create the database 135, as described above.
[0030] The tagging engine 125 executes to tag the collected dataset 117, for example, via the database 135. For example, the tagging engine 125 executes to add tags to fields of each record in the database 135 so that the record 137 can be identified. This allows digital content in the results (e.g., web pages) in the collected dataset 117 to be associated with the record 137 as benign or harmful (e.g., fraudulent, invasive, counterfeit, malicious, etc.) content. The user 106 can specify tags for the records 137 in the database 135 by interacting with the tagging engine 125 via the user interface 110. In some embodiments, the tagging engine 125 is configured to automatically tag the records 137 in the attribute database 135. For example, the tagging engine 125 is configured to use one or more machine learning algorithms to specify tags for the records 137 in the database 135. In this case, the machine learning algorithms can be trained using a collection of training data. The collected data sets 117 may be tagged before, during, or after the collected data sets 117 are collected by one or more collection engines 115 .
[0031] Database conversion engine 140 executes to convert, format, and load data from database 135 into database 145. In an exemplary embodiment, database 145 is a graph database utilizing a graph data model and / or a multi-model database utilizing one or more data models (e.g., a graph data model, a document data model, a key-value model, etc.). However, other types of databases and other types of data models may be used in embodiments of the present disclosure. In an embodiment, records 137 from database 135 are converted into documents 147 in database 145 and may be stored, for example, as Java Script Object Notation (JSON) documents. However, documents 147 may also be stored using other data structures, such as eXtensible Mark-up Language (XML) documents. Database conversion engine 140 converts data fields or columns of records 137 in database 135 into keys for documents 147 in database 145 and modifies data such as phone numbers, email addresses, and addresses to display them in a standard form. The database transformation engine 140 transforms portions of the data in data fields using one or more hash functions / algorithms to make the data suitable for use as a key in the database 145. As a non-limiting example, portions of the data may be transformed using MD5 hashes. Even when MD5 hashes are used, data can be cleaned for alignment purposes either directly on the data or during transformation to reduce the effort of post-transformation reconciliation (entity resolution) as an analytical step.
[0032] The engine 100 utilizes an inverted search index to evaluate documents 147 in the database 145 and maintains statistics of documents 147 in the database 145 tagged as harmful and statistics of actions taken against entities involved in harmful content associated with documents 147 in the database 145 tagged as harmful.
[0033] In one embodiment, engine 100 creates database 145 based on records 137 stored in database 135. Alternatively, in multiple embodiments, engine 100 creates database 145 from results collected from data sources 104 by collection engine 115, whereby databases 135 and 145 are created and updated in parallel based on results from collection engine 115, and / or database 145 is created and updated independently of or in place of database 135. Engine 100 periodically updates database 145 based on updates to records 137 in database 135, and / or after database 145 is initially created using database 135, engine 100 updates the data in database 145 based on content and information extracted from results generated by collection engine 115.
[0034] Collections of documents 147 in database 145 can be defined for the vertices / nodes and edges of a graph data model. As non-limiting examples, collections of nodes may be defined for entity / merchant names (and other personally identifiable information), domains, domain name servers, domain registrant information, IP addresses, URLs, and URIs, and collections of document 147 edges may be defined for physical addresses, phone numbers, email addresses, domains, domain name servers, domain registrant information, IP addresses, URLs, URIs, product descriptions, and product listings on websites, etc. Collections of edges define relationships between collections of nodes and include "to" and "from" keys used to define explicit relationships that form edges between two or more nodes.
[0035] In a non-limiting example, the collection of nodes / vertices for a seller node can be represented as follows:
number
[0036] In yet another non-limiting example, the collection of nodes / vertices for a domain name server node can be represented as follows:
number
[0037] In a non-limiting example, the collection of edges for a domain name server edge can be represented as follows:
number
[0038] Properties that reference raw data such as name, physical address, phone number, email address, seller IP address, and domains can follow a pattern, so you use one parametric query to create vertices and one query to create edges from these properties. Below are some non-limiting examples of parametric queries to create vertices and edges from source data properties:
number
[0039] The clustering engine 150 includes an entity resolution engine 152 and a probabilistic modeling engine 154 to utilize the database 145 to discover and identify one or more clusters or subgraphs corresponding to networks of entities in the network graph. As a non-limiting example, the entity resolution engine 152 utilizes one or more entity resolution algorithms to identify clusters / subgraphs. A user can specify the scope within which entity resolution occurs. For example, the user can select whether entity resolution begins against the entire database or begins based on any vertices / nodes associated with a particular merchant name. One or more entity resolution algorithms implemented via the entity resolution engine 152 include distributed iterative graph processing or the Pregel algorithm. As a non-limiting example, the entity resolution engine 152 executes a connected component algorithm to discover and identify clusters / subgraphs corresponding to networks of entities in the network graph. The connected component algorithm is used to identify groups of connected merchant accounts in the merchant graph. Merchant accounts that represent the same entity are connected through identifying information such as phone numbers, email addresses, and physical addresses. The connected component algorithm can find connected groups based on these keys. When Pregel's connected component algorithm is executed by the entity resolution engine 152, properties are added to the vertices in the connected component subgraph. The properties can then be queried via a query language to find, for example, the largest connected component graph (cluster / subgraph corresponding to a network of entities). The largest connected component graph could be, for example, the group containing the largest number of entity aliases that may be used to obfuscate behavior to avoid detection.
[0040] The probabilistic modeling engine 154 of the clustering engine 100 identifies probabilistic connections between nodes based on similarities between parameters (key values) associated with the nodes. For example, multiple seller accounts based on the same entity can be linked through common and / or inferred relationships, each of which can form a separate subgraph. These relationships can be used to identify clusters / subgraphs containing related seller accounts. The inferred relationships are added to the graph data model. The probabilistic modeling engine 154 uses probability measures and / or similarity measures, such as one or more machine learning-based probability measures, that can be assigned to the probabilistic connections. The probabilistic modeling engine 154 retains probabilistic connections (and associated edges) above a specified threshold and removes probabilistic connections (and associated edges) below a specified threshold. One non-limiting example of a machine learning-based probability measure utilized by the probabilistic modeling engine 154 is the Levenshtein distance.
[0041] Based on the output of the entity resolution engine 152 and the probabilistic modeling engine 154, the clustering engine 150 creates clusters / subgraphs corresponding to networks of entities contained within a single customer account and / or networks of entities that include multiple customer accounts (cross-account clusters). By evaluating the graph across customer accounts, the clustering engine 150 identifies connections between nodes that would not otherwise be identified. This expands the size and scope of the entity networks, providing users of the customer accounts with a more robust and accurate view of the network of cross-account clusterities and facilitating single-cluster-based removal actions targeted at the networks and pseudonyms of fraudulent entities detected therein. Once cross-account clusters are created in the network graph, the engine 100 can automatically alert one or more users associated with the customer account that the scope of the fraudulent entity network has expanded. Each node in the graph can be associated with a customer account identifier. Additionally, each node in a cluster / subgraph in the graph can include a cluster identifier. A cross-account cluster is determined to have been created based on the presence of one or more customer account identifiers in the cross-account cluster / subgraph.
[0042] Cross-account clustering performed by clustering engine 150 may expose confidential or personal information of other customer accounts to users of customer accounts viewing the network map / network graph. To prevent disclosure of confidential and / or personal information between customer accounts, engine 150 anonymizes and / or obfuscates information in the network map / network graph and / or omits or modifies graph nodes and / or edges to protect customer confidential or personal information while still providing the benefits of cross-account clustering. For example, for a merchant active in another customer's account, each node representing the associated merchant name or other personally identifiable information, such as a phone number or email address, may be displayed as a distinct colored icon without a text label providing the actual personally identifiable information. In this case, a grayed-out icon may be used as an example.
[0043] The network graph contains an overwhelming number of nodes. Engine 100 reduces the scope of the graph and / or clusters / subgraphs in response to a selection or request from a user. As one example, the scope of a network graph or cluster / subgraph can be specified to include only nodes tagged as associated with harmful content. As another example, the scope of a network graph or cluster / subgraph can be limited by geographic location or region (e.g., United States, North America, Northern Hemisphere, etc.) based on physical address data, IP addresses, telephone numbers, etc. As another example, the scope of a graph can be limited by industry type or product type, such that only nodes associated with a particular industry or product category are included in the graph.
[0044] The graphing engine 155 utilizes a graph data model to generate a graphical map of the documents 147 in the database 145 based on a collection of nodes and edges and / or edges explicitly and / or implicitly defined by the engine 100. The graphical map identifies subgraphs corresponding to clusters representing a network of entities. The nodes / vertices and edges of the graphical map provide a visualization of the network that allows users to track and discover relationships between content and information from data sources in a network environment. Nodes can be rendered to include an icon or other graphical representation of the node type. For example, different types of nodes, such as entity names, domain names, domain name servers, etc., can each be represented using a different icon in the graphical map.
[0045] The removal engine 160 initiates an automated takedown of the detected fraudulent content and / or products. Once a record in the database is tagged or determined to be fraudulent, the removal engine 160 can initiate a takedown request for the fraudulent content. For example, the removal engine 160 can generate a Digital Millennium Copyright Act notice (DMCA notice) by retrieving data from the collection dataset 117, the database 135, and / or the database 145 and generating a structured file or email. After the notice is generated, the removal engine 160 can send the notice to the content host or owner. In another example, the removal engine 160 communicates the takedown notice to the content host or owner via an API.
[0046] In an exemplary embodiment, user interface 110 generates one or more graphical user interfaces (GUIs) 114 that include a list of records or documents from a search, e.g., using views of database 135 and / or database 145. In this case, the records or documents are grouped in one or more graphical user interfaces 114 based on one or more identifiers contained in the records in database 135 or the documents in database 145. As a non-limiting example, documents tagged as harmful content and associated with database 145 and / or a network graph associated with the documents may be displayed in graphical user interface 114. As another non-limiting example, records tagged as harmful content and associated with database 135 may be displayed in graphical user interface 114.
[0047] The user interface 110 includes a presentation / visualization engine 112 and one or more graphical user interfaces 114. The presentation engine 112 is configured to provide an interface between one or more services and / or one or more engines implemented in the engine 100. Upon receiving data, the presentation engine 112 executes and generates one or more graphical user interfaces 114 and renders the data within the one or more graphical user interfaces 114. The one or more graphical user interfaces 114 enable the user 106 to interact with the engine 100. The graphical user interfaces 114 also include data output areas for displaying information to the user 106 and data entry fields for receiving information from the user 106. In some examples, the data output areas can include, but are not limited to, text, graphics (e.g., graphs, maps (geographical or other), images, etc.), and / or other suitable data output areas. In some examples, the data entry fields can include, but are not limited to, text boxes, check boxes, buttons, drop-down menus, and / or other suitable data entry fields.
[0048] User interface 110 is generated by engine 100, in embodiments executed by one or more servers and / or one or more user computing devices. User interface 110 is configured to render data corresponding to content and information extracted from data sources (e.g., internet content, etc.) as described herein. User interface 110 provides an interface that allows user 106 to interact with content, information, and identifiers stored in database 135 and / or database 145. For example, user interface 110 can be configured to structure and arrange content and information extracted from web pages collected via collection engine 115 and extraction engine 120.
[0049] As a non-limiting example, the user interface 110 may provide a list or table containing data from database 135 and / or database 145. As one non-limiting example, the user interface 110 may include a list of web page entries collected via the collection engine 115. For example, rows may be associated with records in database 135 and / or documents in database 145 that correspond to the web pages. As another non-limiting example, the user interface may render an interactive network graph, where the interactive network graph has subgraph clusters that identify a network of entities across multiple customer accounts 102.
[0050] The rows and / or row values may be selectable by the user 106, thereby allowing the user 106 to interact with the list to change the item identifiers and / or perform one or more other actions. For example, if the extraction engine 120 is unable to parse one or more item identifiers from the results, the analyst may review the results and enter one or more item identifiers into the rows. The entered item identifiers are then used by the tagging engine 125 and the analysis engine 130 in determining whether the content is legitimate or harmful.
[0051] As described herein, engine 100 further includes a recollection frequency option that allows an account user 106 to specify how often collection engine 115 re-queries data sources in a network environment. For example, user 106 can specify that collection engine 115 should perform searches every hour, every day, every week, every month, every quarter, etc.
[0052] FIG. 2 is a block diagram of an exemplary computing device according to an embodiment of the present disclosure. In this embodiment, computing device 200 is configured as a server. The server is programmed and / or configured to perform one or more of the operations and / or functions of engine 100 and facilitate network detection of unauthorized entities and removal of harmful content on the Internet or other network environment. Computing device 200 includes a non-transitory computer-readable medium for storing one or more computer-executable instructions or software for implementing the exemplary embodiment. The non-transitory computer-readable medium may include, but is not limited to, one or more types of hardware memory, non-transitory tangible media (e.g., one or more magnetic storage disks, one or more optical disks, one or more flash drives, etc.), etc. For example, memory 206 included in computing device 200 may store computer-readable and computer-executable instructions or software for implementing engine 100, or portions thereof, according to an exemplary embodiment.
[0053] Computing device 200 further includes a configurable and / or programmable processor 202 and associated cores 204 for executing computer-readable and computer-executable instructions or software stored in memory 206 for controlling system hardware and other programs, and optionally includes one or more additional configurable and / or programmable processors 202′ and associated cores 204′ (e.g., in the case of a computer system having multiple processors / cores). Processor 202 and processor 202′ may each be a single-core processor or a multi-core (204 and 204′) processor.
[0054] Virtualization may be implemented in computing device 200 to dynamically share infrastructure and resources within the computing device. One or more virtual machines 214 may handle processes running on multiple processors such that the processes appear to utilize a single computing resource rather than multiple computing resources and / or such that the processes appear to allocate computing resources to perform functions and operations associated with engine 100. Multiple virtual machines may be utilized on a single processor, or multiple virtual machines may be distributed across several processors.
[0055] The memory 206 may include computer system memory or random access memory such as DRAM, SRAM, EDORAM, etc. The memory 206 may also include other types of memory or combinations thereof.
[0056] The computing device 200 may further include one or more storage devices 224, such as a hard drive, CD-ROM, high-capacity flash drive, or other computer-readable medium, for storing data and computer-readable instructions and / or software that can be executed by the processor 202 to implement the engine 100 in the exemplary embodiments herein.
[0057] Computing device 200 may include a network interface 212 configured to interact with one or more networks, such as a local area network (LAN), a wide area network (WAN), or the Internet, via one or more network devices 222 through various connections, including, but not limited to, standard telephone lines, LAN or WAN links (e.g., 802.11, T1, T3, 56kb, X.25, etc.), broadband connections (e.g., ISDN, Frame Relay, ATM, etc.), wireless connections (including via cell sites), Controller Area Networks (CAN), or any combination of any or all of the above. Network interface 212 may include a built-in network adapter, a network interface card, a PCMCIA network card, a Card Bus-compatible network adapter, a wireless network adapter, a USB network adapter, a modem, or any other device suitable for interfacing computing device 200 to any type of communication-enabled network to perform the operations herein. Computing device 200 depicted in FIG. 2 is implemented as a server. However, in the exemplary embodiment, computing device 200 may be any computer system, such as a workstation, desktop computer, or other type of computing or communication device, capable of communicating with other devices via wireless or wired communication and having sufficient processor power and memory capacity to perform the operations herein.
[0058] Computing device 200 may run a server application 216, such as any version of a server application, including any Unix-based server application, Linux-based server application, proprietary server application, or other server application capable of running on computing device 200 and performing the operations herein. An example of a server application that may run on a computing device is the Apache server application.
[0059] FIG. 3 illustrates an exemplary network environment 300 for facilitating network detection and monitoring of unauthorized entities on the Internet and / or other network environments, according to an embodiment of the present disclosure. The environment 300 includes user computing devices 310-312 operably integrated with a remote computing system 320, including one or more (local) servers 321-323, via a communications network 340. The communications network 340 can be any network capable of transmitting information between devices communicatively integrated therewith. For example, the communications network 340 can be the Internet, an intranet, a virtual private network (VPN), a wide area network (WAN), a local area network (LAN), etc. The environment 300 can include a repository or database 330 operably connected to the servers 321-323 and the user computing devices 310-312 via the communications network 340. Those skilled in the art will appreciate that the database 330 can be incorporated into one or more of the servers 321-323, such that one or more of the servers can include the database. In an exemplary embodiment, the engine 100 in this embodiment may be implemented by one or more of the servers 321-323, independently or collectively, by one or more user computing devices (e.g., user computing device 312), and / or may be distributed between the servers 321-323 and the user computing devices.
[0060] User computing devices 310-312 may be operated by users to facilitate interaction with engine 100 implemented by one or more of servers 321-323. In an exemplary embodiment, user computing devices (e.g., user computing devices 310-311, etc.) may include a customer application 315 programmed and / or configured to interact with one or more of servers 321-323. In one embodiment, customer application 315 implemented by user computing devices 310-311 may be a web browser capable of navigating to one or more web pages hosting the GUI of engine 100. In some embodiments, customer application 315 implemented by one or more of user computing devices 310-311 may be an application specific to engine 100 (e.g., an application providing a user interface for interacting with server 321, server 322, and / or server 323, etc.) that enables interaction with engine 100 implemented by one or more servers.
[0061] One or more servers 321-323 (and / or user's computing device 312) can execute engine 100 to search content available on communications network 340. For example, engine 100 can be programmed to facilitate searching data sources 350, 360, and 370, each of which can include one or more (remote) servers 380. One or more servers 380 are also programmed to host content and make the content available on communications network 340. As a non-limiting example, server 380 can be a web server configured to host one or more search engines and / or websites searched via APIs utilizing one or more queries generated by engine 100. For example, at least one of data sources 350, data source 360, and / or data source 370 can provide an online marketplace website.
[0062] Database 330 may store information used by engine 100. For example, as described herein, database 330 may store queries, datasets of item identifiers extracted by engine 100, tags associated with engine 100, and / or other suitable information / data used by engine 100 in the present embodiment. Database 330 may further store collected datasets (i.e., collected datasets 117) and / or include database 135 and / or database 145.
[0063] FIG. 4 illustrates an exemplary process 400 for creating and / or updating database 145 and generating graphs and subgraphs to detect networks of unauthorized entities in one embodiment shown in FIG. 1 . In step 402, data in database 135 to be transferred to database 145 is staged in an intermediate data storage repository. In step 404, the data is ingested into database 145. In step 406, the data is transformed (e.g., using regular expressions, hash functions, and / or other data transformations) to conform to the graph data model of database 145, and the graph data model of database 145 is updated. In step 408, entity resolution is performed on the graph data model using a community detection algorithm, such as Pregel's connected component algorithm, to disambiguate and group entities in subgraphs in database 145. An entity identifier unique to each disambiguated entity can be added to documents associated with that entity to link documents related to each entity contained in database 145. In step 410, a report can be generated based on the graph identifying networks of unauthorized entities.
[0064] 5 illustrates a visualization of a graph data model 500 in one embodiment that may be defined for database 145. As shown in FIG. 5, the exemplary graph data model 500 may define nodes of a graph that include a “seller” node 502, a “site” node 504, a “phone” node 506, an “e-mail” node 508, an “address” node 510, an “entity” node 512, a “listing” node 514, a “domain” node 504, an “IP address” node 506, and a “domain name server” node 518. The graph data model 500 can define edges between nodes and can include keys associated with the nodes, such as a “sellersite” key 520, a “sellerPhone” key 522, a “sellerEmail” key 524, a “sellerName” key 526, a “sellerListing” key 528, a “registrantPhone / adminPhone / techPhone” key 532, a “registrantEmail / adminEmail / techEmail” key 534, a “registrantAddress” key 536, a “registrar / registrant” key 538, a “domainIPaddress” key 540, and a “domainNameServer” key 542.
[0065] FIG. 6 illustrates a simple example of a graph 600 generated by engine 100 for data in database 145 based on a graph data model defined for database 145 and execution of a connected component algorithm. As shown in FIG. 6, graph 600 includes subgraphs 610 and 650. Subgraphs 610 and 650 define separate networks for each network of entities discovered by execution of engine 100 in a network environment. Subgraph 610 may include three entities, "Seller1," "Seller2," and "Seller3," represented by nodes 612, 614, and 616. Nodes 612, 614, and 616 are interconnected within subgraph 610 based on intervening nodes 618 and 620, which represent telephone numbers and physical addresses. This connection indicates that Seller1 and Seller3 contain a shared key representing a telephone number, and that Seller1 and Seller2 contain a shared key representing a physical address. Using these common keys, engine 100 resolves that Seller1, Seller2, and Seller3 represent the same entity, and that the nodes in the network associated with Seller1, Seller2, and Seller3 are linked to each other. Subgraph 650, on the other hand, represents a network of entities that includes a single entity, "Seller5," represented by node 652. As shown in FIG. 6, node 652 is not connected to any other "seller" nodes in graph 600.
[0066] FIG. 7 illustrates a graphical user interface 700 showing a list 702 of documents and their corresponding keys in database 145, by customer account, in one embodiment. List 702 may correspond to all lists available for daily review, for example, based on collection. As shown in FIG. 7 , the list includes a “title” column 704, a “cluster connection” column 706, a “URL” column 708, a “Domain” column 710, a “Hosting status” column 712, a “Registrar” column 714, a “Registrant Name” column 716, and a “First Detected” column 718. The “title” column 704 includes the title of the network content collected from the data source. In an exemplary embodiment, the title may be extracted from the web page or the source code of the web page. The “cluster connection” column 706 identifies the number of connections present in the subgraph containing the document’s key. The “URL” column 708 may correspond to the URL address of the network content collected from the data source. The “Domain” column 710 may correspond to the URL address of the network content collected from the data source. The "Hosting Status" column 712 can specify whether the domain is active or inactive. The "Registrar" column 714 can specify the registrar that registers the domain. The "Registrant Name" column 716 can specify the entity that registered the domain with the registrar. The "First Detected" column 718 can specify the date the engine first identified the content from a data source.
[0067] In an example interaction with the graphical user interface 700, a user may select a number from the "cluster connection" column 706 to view a cluster summary. As an example, the user may select the number 720 associated with the fourth row of the list. In response to selecting the number, the graphical user interface 700 may render an area that displays the cluster summary.
[0068] The graphical user interface 700 may also include selectable options to facilitate one or more actions or functions. The options may include a "Detect" option 730, a "Review" option 732, an "Enforce" option 734, a "Report" option 736, and a "Cluster Browser" option 738. When the Detect option 730 is selected if the list 702 is not already rendered in the graphical user interface, the engine 100 may navigate to the list 702 of documents to view or review. In response to a selection of the Review option 732, the engine 100 may navigate to a graphical user interface that allows the user to review the documents and files associated with information contained in the collected data that form the list. The list also allows the user to tag or re-tag documents as legitimate or fraudulent. In response to a selection of the Enforce option 734, the engine 100 may initiate enforcement action against one or more of the sellers, listings, or domains identified as fraudulent, and / or simultaneously initiate action against the fraudulent entity's network. In response to a selection of Report option 736, engine 100 may generate one or more reports associated with the document including statistics and / or one or more statuses associated with fraudulent activity or removal actions against fraudulent entities and / or networks of fraudulent entities. In response to a selection of Cluster Browser option 738, engine 100 may navigate to a graphical user interface that provides a cluster browser that can render a network graph or subgraph containing clusters of related or connected nodes, where the related or connected nodes form a network of one or more entities.
[0069] 8 illustrates a graphical user interface 700 that includes an area 800 showing a cluster summary in response to a user's selection of the number of clusters 720 in FIG. 7. As shown in FIG. 8, area 800 includes information about the selected cluster, such as the amount of entities included in the cluster 802, the defect rate (or fraud rate) 804, the types of entities included in the cluster 806, the quantity 808, the types of content included in the cluster 810, the defect rate (or fraud rate) of the content 812, and the amount of that type of content in the cluster 814. Area 800 may also include a "view cluster" option 816. A user can select the "view cluster" option 816 to render a cluster browser that facilitates visualization of the cluster in a map or graph.
[0070] FIG. 9 is a graphical user interface 900 illustrating a cluster browser that can render a visualization of a graph 910 in response to selection of the "View Clusters" option 816 in area 800 shown in FIG. 8. As shown in FIG. 9, graph 902 has nodes 904 and edges 906, and nodes 904 include icons that visually indicate the node's type, e.g., phone number, email address, physical address, domain name, etc. The cluster browser can also include an area 910 containing selectable options for entity type 912 and item type 914. From within the cluster browser, a user can search for a node (e.g., seller name, personally identifiable information, registrant information, etc.). The user can select a cluster in the cluster browser graphical user interface (URL, list, or post) 914 to navigate to a grid view of the selected option and display a grid view of items in the cluster corresponding to the selected item type.
[0071] FIG. 10 is a graphical user interface 1000 illustrating a grid view in response to selection of one of options 912 and / or 914. As shown in FIG. 10, the graphical user interface 1000 includes a list 1002 of documents and their respective keys in one embodiment of database 145 for customer accounts. The list 1002 can correspond to a refined selection of lists associated with a particular cluster / subgraph detected in the network graph using both deterministic and probabilistic processing as described herein (as compared to FIG. 7 ). As shown in FIG. 10 , the list includes a “Title” column 704, a “Cluster Connection” column 706, a “URL” column 708, a “Domain” column 710, a “Hosting Status” column 712, a “Registrar” column 714, a “Registrant Name” column 716, and a “First Found” column 718, as described herein.
[0072] From the graphical user interface 1000, the user can initiate a delete action on some or all of the items in the list as a cluster to delete these items at once in a single action initiated by the engine 100.
[0073] 11 illustrates an exemplary method 1100 for parsing and classifying item identifiers utilizing an unauthorized content detection engine implemented in accordance with an embodiment of the present disclosure. At operation 1102, the unauthorized content detection engine (i.e., engine 100) performs a content search against one or more data sources for each account. At operation 1104, search results are obtained in response to the content search. At operation 1106, the unauthorized content detection engine extracts content and information from the search results.
[0074] At operation 1108, a record is created in a database (e.g., database 135) for each unique search result, and content and information extracted from the search result are stored as data in the record's data fields. At operation 1110, engine 100 adds tags to the record to identify whether the record (and corresponding search result) is legitimate or malicious. At operation 1112, a graph data model is defined for a graph database (e.g., database 145) (including node and edge collections), and records in the database are copied to the graph database, where the records are converted into documents and data fields in the records are converted into keys as described herein. At operation 1114, one or more entity resolution algorithms are run on the graph including the documents from the customer accounts to identify edges between nodes and / or define subgraphs / clusters corresponding to networks of distinct entities. As an example, the Pregel algorithm of the connection component can be run. In some cases, distinct subgraphs / clusters contained within a single customer account can be identified. However, in some cases, distinct clusters / subgraphs can be identified across boundaries from one customer account to another. This extends the size and scope of the network of entities beyond a single customer account. In step 1116, the identified network of entities is utilized to identify networks of unauthorized entities. In step 1118, for networks of unauthorized entities that cross boundaries between customer accounts, an engine is executed to anonymize and / or obfuscate information in the network map / graph, and / or nodes and / or edges of the graph may be omitted or modified to preserve sensitive or private customer information. In step 1120, one or more graphical user interfaces may be rendered to users of the client accounts, as described herein.
[0075] The exemplary flowcharts are provided herein for illustrative purposes and are non-limiting examples of methods. Those skilled in the art will recognize that the exemplary methods may include more or fewer steps than those shown in the exemplary flowcharts, and that the steps of the exemplary flowcharts may be performed in a different order than that shown in the exemplary flowcharts.
[0076] The foregoing descriptions of specific embodiments of the subject matter disclosed herein have been presented for purposes of illustration and description and are not intended to limit the scope of the subject matter described herein. It is fully contemplated that various other embodiments, modifications, and applications will become apparent to those skilled in the art from the foregoing description and the accompanying drawings. Accordingly, such other embodiments, modifications, and applications are intended to be included within the scope of the appended claims. Furthermore, those skilled in the art will appreciate that the embodiments, variations, and applications described herein relate to particular environments and that the subject matter described herein is not limited thereto and may be beneficially applied in a variety of environments. There are numerous other methods, environments, and objectives. Accordingly, the claims set forth below should be construed in light of the full scope and spirit of the novel features and techniques disclosed herein.
Claims
1. 1. A system for detecting and monitoring a network of unauthorized entities in a networked environment, comprising: The system is a computing system communicatively coupled to a data source in the network environment, the data source including one or more remote servers configured to host digital content; One or more processors located within the computing system are programmed to: Generate individual customer accounts, for each customer account, searching for said digital content hosted by one or more remote servers within said network environment to generate a separate aggregated data set for each customer account; tagging each search result in the aggregated dataset as either legitimate or harmful based on an analysis of each search result, where a tagging as harmful indicates that the digital content is fraudulent, counterfeit, infringing, or malicious; combining data from each search result in a dataset collected for said customer account to generate a network graph; generating clusters within the network graph, including cross-account clusters that include data from two or more customer accounts; Identifying one or more networks of unauthorized entities based on the clusters in the network graph; and Initiating a takedown action against the identified one or more unauthorized entity networks, wherein the takedown action includes automatically initiating a takedown request against the identified one or more unauthorized entity networks by a takedown engine.
2. 2. The system of claim 1, wherein the customer account includes confidential or private data utilized to generate the network graph, and the one or more processors are further programmed to: Preventing confidential or private data from each customer account from being disclosed to other customer accounts.
3. 3. The system of claim 2, wherein the one or more processors are programmed to prevent disclosure of the sensitive or private data by modifying cross-account clusters in the graph to at least one of obfuscate or remove the sensitive or private data.
4. 2. The system of claim 1, wherein the one or more processors are programmed to: For each search result in the collected data set for each customer account, analyze whether the search result corresponds to legitimate content or harmful content; and Based on the analysis, each search result in the collected data set for each customer is tagged as either good or bad.
5. The system of claim 1 , wherein the one or more processors are programmed to: creating multiple records in a relational database for each unique search result in the collected data set for each customer account; and Data extracted from each result of the collected data set for each customer account is stored in a data field of a corresponding record of the plurality of records.
6. The system of claim 5 , wherein the one or more processors are further programmed to: Create a graph database, defining a graph data model for said graph database; Copying the plurality of records from the relational database to documents in the graph database; copying the data fields from the plurality of records in the relational database to keys of documents in the graph database; and At least one of a node collection or an edge collection is generated in the graph database based on the document and the key of the document.
7. The system of claim 6 , wherein the data forming the key is converted to a canonical form or converted by a hashing algorithm.
8. The system of claim 1 , wherein the one or more remote servers in the network environment are web servers, and the digital content hosted by the one or more remote servers is a website including web pages.
9. The system of claim 1 , wherein the one or more processors are further programmed to generate at least one of the network graph or the clusters in response to execution of an entity resolution algorithm.
10. The system of claim 9 , wherein the entity resolution algorithm is a connected components algorithm.
11. The system of claim 1 , wherein the one or more processors are programmed to: Detecting the formation of one of the cross-account clusters; and A user of one of the customer accounts associated with one of the cross-account clusters is alerted.
12. A method for detecting and monitoring a network of unauthorized entities in a network environment, the method being implemented via a computing system communicatively coupled to a data source in the network environment, the data source including one or more remote servers configured to host digital content, the method comprising one or more processors disposed within the computing system; The method includes, by at least one of the one or more processors: Create individual customer accounts for each customer account, searching for said digital content hosted by one or more remote servers within said network environment to generate a separate aggregated data set for each customer account; tagging each search result in said aggregated data set as either legitimate or harmful based on an analysis of each search result, wherein a tagging as harmful indicates that the digital content is fraudulent, counterfeit, infringing, or malicious; combining data from each search result in a dataset collected for said customer account to generate a network graph; generating clusters within the network graph, including cross-account clusters that include data from two or more customer accounts; Identifying one or more networks of unauthorized entities based on the clusters in the network graph; and Initiating a takedown action against the identified one or more unauthorized entity networks, wherein the takedown action includes automatically initiating a takedown request against the identified one or more unauthorized entity networks by a takedown engine.
13. 13. The method of claim 12, the customer account includes confidential or private data utilized to generate the network graph; The method further includes preventing, by at least one of the one or more processors, confidential or private data of each customer account from being disclosed to other customer accounts.
14. 14. The method of claim 13, further comprising: by at least one of the one or more processors; Modifying cross-account clusters in the graph to at least one of obfuscate or remove the sensitive or private data, thereby preventing disclosure of the sensitive or private data.
15. 13. The method of claim 12, further comprising: by at least one of the one or more processors; For each search result in the collected data set for each customer account, analyze whether the search result corresponds to legitimate content or harmful content; and Based on the analysis, each search result in the collected data set for each customer is tagged as either good or bad.
16. 13. The method of claim 12, further comprising: creating, by at least one of the one or more processors, a plurality of records in a relational database for each unique search result in the collected dataset for each customer account; and At least one of the one or more processors stores data extracted from each result of the collected data set for each customer account in a data field of a corresponding record of the plurality of records.
17. 17. The method of claim 16, further comprising: by at least one of the one or more processors; Create a graph database, defining a graph data model for said graph database; Copying the plurality of records from a relational database to a document in the graph database; copying the data fields from the plurality of records in the relational database to keys of documents in the graph database; and At least one of a node collection or an edge collection is generated in the graph database based on the document and the key of the document.
18. 18. The method of claim 17, wherein the data forming the key is at least one of transformed into a canonical form and transformed by a hashing algorithm.
19. The method of claim 12 , wherein the one or more remote servers in the network environment are web servers, and the content hosted by the one or more remote servers is a website including web pages.
20. The method of claim 12, further comprising a step of generating at least one of the network graph or the clusters by at least one of the one or more processors in response to execution of an entity resolution algorithm.
21. The method of claim 20 , wherein the entity resolution algorithm is a connected components algorithm.
22. 13. The method of claim 12, further comprising: Detecting, by at least one of the one or more processors, the formation of one of the cross-account clusters; and At least one of the one or more processors alerts a user of one of the customer accounts associated with one of the cross-account clusters.
23. 1. A non-transitory computer-readable medium having stored thereon instructions for detecting and monitoring a network for unauthorized entities in a network environment, the instructions, when executed by one or more processors, causing the one or more processors to: Generate individual customer accounts, for each customer account, searching for digital content hosted by one or more remote servers within said network environment to generate a separate aggregated data set for each customer account; tagging each search result in the aggregated dataset as either legitimate or harmful based on an analysis of each search result, where a tagging as harmful indicates that the digital content is fraudulent, counterfeit, infringing, or malicious; combining data from each search result in a dataset collected for said customer account to generate a network graph; generating clusters within the network graph, including cross-account clusters that include data from two or more customer accounts; Identifying one or more networks of unauthorized entities based on the clusters in the network graph; and Initiating a takedown action against the identified one or more unauthorized entity networks, wherein the takedown action includes automatically initiating a takedown request against the identified one or more unauthorized entity networks by a takedown engine.
24. 24. The medium of claim 23, wherein the customer accounts include sensitive or private data utilized to generate the network graph, and execution of the instructions causes the one or more processors to: Preventing confidential or private data from each customer account from being disclosed to other customer accounts.
25. 25. The medium of claim 24, wherein execution of the instructions causes the one or more processors to modify cross-account clusters in the graph to at least one of obfuscate or remove the sensitive or private data to prevent disclosure of the sensitive or private data.
26. 24. The medium of claim 23, wherein execution of the instructions causes the one or more processors to: For each search result in the collected data set for each customer account, analyze whether the search result corresponds to legitimate content or harmful content; and Based on the analysis, each search result in the collected data set for each customer is tagged as either good or bad.
27. 24. The medium of claim 23, wherein execution of the instructions causes the one or more processors to: creating multiple records in a relational database for each unique search result in the collected data set for each customer account; and Data extracted from each result of the collected data set for each customer account is stored in a data field of a corresponding record of the plurality of records.
28. 30. The medium of claim 27, wherein execution of the instructions causes the one or more processors to: Create a graph database, defining a graph data model for said graph database; Copying the plurality of records from the relational database to documents in the graph database; copying the data fields from the plurality of records in the relational database to keys of documents in the graph database; and At least one of a node collection or an edge collection is generated in the graph database based on the document and the key of the document.
29. 30. The medium of claim 28, wherein the data forming the key is at least one of converted to a canonical form and converted by a hashing algorithm.
30. 24. The medium of claim 23, wherein the one or more remote servers in the network environment are web servers, and the digital content hosted by the one or more remote servers is a website including web pages.
31. 24. The medium of claim 23, wherein execution of the instructions causes the one or more processors to generate at least one of the network graph or the clusters in response to execution of an entity resolution algorithm.
32. 32. The medium of claim 31, wherein the entity resolution algorithm is a connected components algorithm.
33. 24. The medium of claim 23, wherein execution of the instructions causes the one or more processors to: Detecting the formation of one of the cross-account clusters; and A user of one of the customer accounts associated with one of the cross-account clusters is alerted.
Citation Information
Patent Citations
Electronic commerce system
JP2002203143A
Network-Oriented Product Rollout in Online Social Networks
JP2016530586A
System and method for limiting user access to suspicious objects in social network
JP2018190370A
Systems and methods for identifying matching content
JP2019509577A
Social graph aggregation systems and methods
US10991010B1